Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
Job responsibilities
Required qualifications, capabilities, and skills
Formal training or certification on software engineering concepts and 5+ years applied experience.
Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
Ability to identify and solve problems related to complex data structures and algorithms
Deep expertise in public cloud platforms (AWS or equivalent), infrastructure automation tools (CloudFormation, Terraform), and capacity planning for large-scale environments, with a track record of driving DevOps and SRE adoption across teams.
Expertise in distributed systems building resilient systems
Ability to expand and collaborate across different levels and stakeholder groups
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.