W

Site Reliability Engineer

Job Description - Site Reliability Engineer




We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-performance applications running in a hybrid environment (on-premises and cloud). The ideal candidate will have strong experience in cloud infrastructure, Kubernetes, observability, automation, and production operations.


Requirements



  • Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).

  • Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.

  • Experience working with Programming languages such as Go, Python, Java, Rust etc.

  • Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any  time-series databases

  • Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher

  • Experience maintaining containerized app in GKE/RKE/AKE environments.

  • Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.

  • Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).

  • Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).

  • Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.

  • Experience working with Programming languages such as Go, Python, Java, Rust etc.

  • Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any  time-series databases

  • Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher

  • Experience maintaining containerized app in GKE/RKE/AKE environments.

  • Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.

  • Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).

  • Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.




Original job Site Reliability Engineer posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Site Reliability Engineer Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.