We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-performance applications running in a hybrid environment (on-premises and cloud). The ideal candidate will have strong experience in cloud infrastructure, Kubernetes, observability, automation, and production operations.
Requirements
Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
Experience working with Programming languages such as Go, Python, Java, Rust etc.
Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
Experience in transitioning platforms to the cloud and Containerization â GCPand Rancher
Experience maintaining containerized app in GKE/RKE/AKE environments.
Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
Experience working with Programming languages such as Go, Python, Java, Rust etc.
Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
Experience in transitioning platforms to the cloud and Containerization â GCPand Rancher
Experience maintaining containerized app in GKE/RKE/AKE environments.
Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in the US.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip
GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast!
Find the best jobs in the US, apply in 1 click and get a job today!