K

Site Reliability Engineer II

Job Description - Site Reliability Engineer II

Description
  • 5+ years of experience in site reliability engineering, DevOps, infrastructure, or production operations roles.

  • Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers

  • Experience operating in shift-based, on-call, or follow-the-sun coverage models

  • Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)

  • Proficiency in at least one scripting or programming language sufficient to review and guide automation work

  • Understanding of SLI/SLO frameworks and reliability engineering fundamentals

  • Strong written and verbal English communication skills for cross-region collaboration with US and India teams



Responsibilities
  • 5+ years in site reliability engineering, DevOps, infrastructure, or production operations roles

  • 1+ years of people management experience, or 2+ years as a senior technical lead with demonstrated coaching and delivery ownership

  • Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers

  • Experience operating in shift-based, on-call, or follow-the-sun coverage models

  • Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)

  • Proficiency in at least one scripting or programming language sufficient to review and guide automation work

  • Understanding of SLI/SLO frameworks and reliability engineering fundamentals

  • Strong written and verbal English communication skills for cross-region collaboration with US and India teams



Qualifications
  • Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)

  • Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns

  • Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps

  • Experience in multi-region or globally distributed team models

  • Relevant certifications (AWS, CKA, or similar)



Original job Site Reliability Engineer II posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Site Reliability Engineer Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.