We're looking for a Site Reliability Engineer, focused on building and operating the data and AI/ML infrastructure platform that powers NetApp's cloud-native data services. You'll work at the intersection of software engineering and infrastructure operations — designing systems for reliability, driving automation, and ensuring our platforms meet the highest availability standards for customers worldwide.
This is an infrastructure-focused SRE role, you'll own the reliability of large-scale Kubernetes clusters (including GPU workloads), streaming data pipelines (Kafka), and analytical compute infrastructure (Spark, Dremio) across hybrid-cloud and multi-cloud environments.
Nice to Have
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.