S

Site Reliability Engineer

Job Description - Site Reliability Engineer

Roles & Responsibilities:



  • Ensure production reliability and performance by monitoring service availability, latency, capacity, and overall system health while managing SLOs, SLIs, and error budgets.

  • Automate operational processes and reduce toil by developing tools and scripts for deployment, infrastructure management, incident response, capacity planning, and other repetitive operational tasks.

  • Manage incident response and root-cause analysis, participate in on-call rotations, troubleshoot production issues, and conduct blameless postmortems to implement long-term corrective actions.

  • Design and improve scalable infrastructure by collaborating with software engineering teams on system architecture, distributed systems, reliability, scalability, and performance requirements.

  • Manage cloud infrastructure and deployment processes using technologies such as Kubernetes, Terraform, Ansible, and CI/CD pipelines, including canary releases and automated deployment practices.

  • Implement observability, capacity planning, and performance optimization across production systems, using monitoring and logging platforms such as Datadog or Splunk and optimizing SQL/NoSQL databases and distributed services.


Qualifications:



  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent practical experience.

  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Infrastructure, or a closely related role, with strong exposure to production environments.

  • Strong programming or scripting skills in at least one language such as Python, Go, Java, or C++, with the ability to develop automation tools and solve complex technical problems.

  • Strong knowledge of Linux/Unix systems, networking, and distributed systems, including TCP/IP, DNS, load balancing, service reliability, and troubleshooting across distributed environments.

  • Hands-on experience with cloud-native infrastructure and Infrastructure as Code, including Kubernetes, Terraform, Ansible, CI/CD, and automated deployment practices.

  • Experience with observability and database technologies, including tools such as Datadog/Splunk and both relational and NoSQL databases, along with strong analytical, debugging, incident-management, and performance-tuning skills.

Original job Site Reliability Engineer posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Site Reliability Engineer Jobs in India

GrabJobs is the no1 job portal in India, connecting you to thousands of jobs fast! Find the best jobs in India, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.