The Site Reliability Engineer (SRE) is responsible for ensuring the availability, performance, scalability, and reliability of enterprise platforms and applications. The role focuses on monitoring, automation, incident management, and continuous improvement, working closely with engineering and DevOps teams to build resilient and highly available systems.
2. Key Responsibilities
Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
Infrastructure as Code: Terraform, Ansible, CloudFormation
Responsibilities
2. Key Responsibilities
Ensure high availability and reliability of production systems and services Monitor system health using observability tools (logs, metrics, traces) Define and track SLIs, SLOs, and SLAs to measure system performance Lead/support incident management, root cause analysis (RCA), and post-incident reviews Automate operational tasks and implement Infrastructure as Code (IaC) practices Support and improve CI/CD pipelines for stable and efficient releases 3. Skills & Competencies
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in India.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip