Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
4–8 years of experience in SRE, Platform Engineering, or DevOps, with a strong senior IC track record of owning production systems.
Production-grade expertise with Kubernetes, containers, and cloud platforms (AWS preferred) in distributed-systems environments.
Hands-on experience defining and operating against SLIs, SLOs, and error budgets, along with leading incident response and blameless postmortems.
Strong experience with observability tools covering metrics, logging, tracing, dashboards, and alerting, such as Datadog, Prometheus/Grafana, or equivalent platforms.
Proficient in automation and infrastructure-as-code, using technologies such as Python, Go, Shell, and Terraform.
Comfortable working with GitOps and CI/CD practices for reliable software and infrastructure delivery.
Strong understanding of Linux/Unix internals, networking, and cloud-native security fundamentals.
Strong operational rigor, ownership mindset, and ability to communicate clearly in both written and verbal formats, including during high-pressure incidents.
Requirements
Responsibilities
Own the end-to-end reliability of Syfe’s production platform, ensuring high availability, performance, and operational stability.
Work with a Kubernetes-native, multi-region platform across Singapore, Hong Kong, and Sydney, supporting a regulated digital wealth-management product.
Operate as a senior individual contributor, defining measurable reliability standards and building systems, automation, and processes to maintain them.
Define and drive SLIs, SLOs, and error budgets across critical services, partnering with product and engineering teams to balance velocity and stability.
Own the on-call, escalation, and incident-response program, including incident command, blameless postmortems, RCA tracking, and reducing MTTD and MTTR.
Own the reliability of the AWS EKS-based deployment platform, including GitOps with ArgoCD, Helm-based release configuration, and Infrastructure as Code using Terraform/OpenTofu.
Ensure deployments are safe, progressive, and reversible, with strong rollout and rollback mechanisms.
Build and continuously improve the observability stack using tools such as Datadog, Grafana, VictoriaMetrics, and ClickHouse.
Develop meaningful dashboards, actionable alerts, and monitoring practices while reducing unnecessary alert noise.
Lead capacity planning, scalability analysis, failure-mode analysis, disaster recovery, and business continuity planning across regions.
Plan and conduct game days and chaos engineering exercises to validate system resilience and recovery processes.
Identify operational toil and eliminate it through automation, self-service tooling, and improved engineering practices.
Strengthen production safety, including deployment guardrails, rollback processes, and secrets management using HashiCorp Vault.
Partner with engineering teams to improve production readiness and service reliability before systems are deployed to production.
Drive reliability through hands-on engineering, design reviews, production-readiness reviews, runbooks, and technical documentation.
Mentor engineers and influence engineering teams to adopt SRE best practices and build a strong culture of production ownership.
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in India.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip