N

Software Engineer SRE (Site Reliability Engineer)

Job Description - Software Engineer SRE (Site Reliability Engineer)


Job Summary

We're looking for a Site Reliability Engineer, focused on building and operating the data and AI/ML infrastructure platform that powers NetApp's cloud-native data services. You'll work at the intersection of software engineering and infrastructure operations — designing systems for reliability, driving automation, and ensuring our platforms meet the highest availability standards for customers worldwide.


This is an infrastructure-focused SRE role, you'll own the reliability of large-scale Kubernetes clusters (including GPU workloads), streaming data pipelines (Kafka), and analytical compute infrastructure (Spark, Dremio) across hybrid-cloud and multi-cloud environments.

Job Requirements


  • 5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering roles

  • Extensive experience with Linux (RHEL/CentOS), including shells, filesystems, kernel tuning, networking, and performance optimization

  • Deep expertise with Kubernetes at scale, including cluster administration, troubleshooting, networking, storage, RBAC, and lifecycle management (on-premises and Rancher Kubernetes)

  • Hands-on experience operating GPU workloads on Kubernetes, including NVIDIA GPU Operator, device plugins, scheduling, and resource management

  • Strong experience managing Confluent Kafka in production, including operations, monitoring, performance tuning, and disaster recovery

  • Experience operating Apache Spark and/or Dremio, including cluster management, job scheduling, scaling, and performance optimization

  • Proficiency in Infrastructure as Code using Terraform, Helm, and GitOps workflows with ArgoCD/FluxCD

  • Proficiency in scripting and automation using Shell, Ansible, and Python, with a strong automation-first mindset

  • Experience with scheduling and orchestration tools such as cron jobs and Apache Airflow

  • Deep familiarity with monitoring and observability tools, including Dynatrace, Grafana, and Prometheus

  • Solid understanding of SQL and NoSQL databases, including operations, backup, and monitoring

  • Experience designing and maintaining CI/CD pipelines and release processes

  • Expertise in AWS cloud platforms and hybrid-cloud integration

  • Strong systems thinking, with an understanding of how infrastructure design choices impact failure modes, scalability, and recovery

  • Strong incident management skills and post-mortem facilitation experience

  • Excellent written communication skills for design documents, runbooks, post-mortems, and operational documentation



Nice to Have



  • Knowledge of Generative AI tools and frameworks, including the application of AI-based predictive analytics and automation in infrastructure operations

  • Familiarity with ML platforms such as Kubeflow, MLflow, and Ray, as well as AI/ML training infrastructure

  • Experience with Kafka Streams, ksqlDB, or Apache Flink


Education


  • 5-8 years of relevant experience.

  • Bachelor of Science Degree in Computer Science, Electrical Engineering, or a related field; a Master’s Degree is preferred. 


Original job Software Engineer SRE (Site Reliability Engineer) posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Software Engineer SRE Jobs in India

GrabJobs is the no1 job portal in India, connecting you to thousands of jobs fast! Find the best jobs in India, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.