B

Site Reliability Engineer

Job Description - Site Reliability Engineer

Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 4–8 years of experience in SRE, Platform Engineering, or DevOps, with a strong senior IC track record of owning production systems.
  • Production-grade expertise with Kubernetes, containers, and cloud platforms (AWS preferred) in distributed-systems environments.
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets, along with leading incident response and blameless postmortems.
  • Strong experience with observability tools covering metrics, logging, tracing, dashboards, and alerting, such as Datadog, Prometheus/Grafana, or equivalent platforms.
  • Proficient in automation and infrastructure-as-code, using technologies such as Python, Go, Shell, and Terraform.
  • Comfortable working with GitOps and CI/CD practices for reliable software and infrastructure delivery.
  • Strong understanding of Linux/Unix internals, networking, and cloud-native security fundamentals.
  • Strong operational rigor, ownership mindset, and ability to communicate clearly in both written and verbal formats, including during high-pressure incidents.


Requirements

Responsibilities

  • Own the end-to-end reliability of Syfe’s production platform, ensuring high availability, performance, and operational stability.
  • Work with a Kubernetes-native, multi-region platform across Singapore, Hong Kong, and Sydney, supporting a regulated digital wealth-management product.
  • Operate as a senior individual contributor, defining measurable reliability standards and building systems, automation, and processes to maintain them.
  • Define and drive SLIs, SLOs, and error budgets across critical services, partnering with product and engineering teams to balance velocity and stability.
  • Own the on-call, escalation, and incident-response program, including incident command, blameless postmortems, RCA tracking, and reducing MTTD and MTTR.
  • Own the reliability of the AWS EKS-based deployment platform, including GitOps with ArgoCD, Helm-based release configuration, and Infrastructure as Code using Terraform/OpenTofu.
  • Ensure deployments are safe, progressive, and reversible, with strong rollout and rollback mechanisms.
  • Build and continuously improve the observability stack using tools such as Datadog, Grafana, VictoriaMetrics, and ClickHouse.
  • Develop meaningful dashboards, actionable alerts, and monitoring practices while reducing unnecessary alert noise.
  • Lead capacity planning, scalability analysis, failure-mode analysis, disaster recovery, and business continuity planning across regions.
  • Plan and conduct game days and chaos engineering exercises to validate system resilience and recovery processes.
  • Identify operational toil and eliminate it through automation, self-service tooling, and improved engineering practices.
  • Strengthen production safety, including deployment guardrails, rollback processes, and secrets management using HashiCorp Vault.
  • Partner with engineering teams to improve production readiness and service reliability before systems are deployed to production.
  • Drive reliability through hands-on engineering, design reviews, production-readiness reviews, runbooks, and technical documentation.
  • Mentor engineers and influence engineering teams to adopt SRE best practices and build a strong culture of production ownership.



Original job Site Reliability Engineer posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Site Reliability Engineer Jobs in India

GrabJobs is the no1 job portal in India, connecting you to thousands of jobs fast! Find the best jobs in India, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.