Logo-of-Upscaleai-hiring-for-jobs-in-US-on-GrabJobs

Senior Principal DevOps Engineer Orchestration

Job Description - Senior Principal DevOps Engineer Orchestration

Why join Upscale AI


Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.


We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.


If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.



About the role


Own the reliability, deployment, and operational infrastructure behind Orchestrator and the AI Fabric environments it manages.


You will build and maintain Kubernetes clusters across on-prem and cloud, design CI/CD pipelines for continuous delivery, manage Terraform-driven infrastructure-as-code, and handle secret and certificate rotations.


You will stand up and operate the full observability stack — Prometheus, Grafana, Loki, Splunk, Datadog, and Timestream — ensuring end-to-end visibility across customer deployments. When things break, you are the person who troubleshoots and debugs infrastructure issues at the platform and customer site level.


Beyond keeping things running, you will build internal tooling and analytics that improve operational efficiency, reduce incident response time, and scale our infrastructure as deployments grow. We are looking for someone who has been through production pain, knows what good SRE looks like, and can bring that discipline to a fast-moving team.


What you'll work on



  • Kubernetes cluster lifecycle across hybrid environments: provision, upgrade, scale, and harden clusters running on bare-metal (on-prem customer datacenters) and cloud (AWS/GCP) using kubeadm, Rancher, or equivalent tooling

  • CI/CD pipeline design and ownership: build and maintain pipelines (GitHub Actions or equivalent) that deliver Go microservices, React UI, Helm charts, and edge appliance images from commit to production with automated testing gates

  • Infrastructure-as-code: manage all cloud and on-prem infrastructure through Terraform modules with proper state management, drift detection, and PR-based review workflows

  • Observability stack operations: deploy, tune, and maintain Prometheus (metrics), Grafana (dashboards), Loki (logs), Splunk and Datadog (enterprise monitoring), and Timestream (time-series analytics) — build the dashboards and alerts that give the team real-time visibility into platform and customer-site health

  • Secret and certificate management: automate mTLS certificate rotation across hub-to-edge communication channels, manage Vault or equivalent secret stores, and handle credential lifecycle for multi-tenant deployments

  • Incident response and debugging: own the runbooks, triage production issues across the distributed hub/edge architecture, perform root cause analysis, and drive post-incident reviews that result in real fixes — not just documents

  • Customer site operations: support deployment, upgrade, and troubleshooting of edge appliances running in customer datacenters with varying network constraints and access patterns

  • Internal tooling: build CLI tools, deployment automation, environment provisioners, and operational dashboards that reduce toil and make the engineering team faster

  • Capacity planning and cost optimization: monitor resource utilization across clusters and cloud accounts, right-size workloads, and forecast infrastructure needs as site count grows


What you bring



  • 6–10 years in SRE, DevOps, or infrastructure engineering roles supporting production distributed systems

  • Deep Kubernetes expertise: cluster administration, networking (CNI, ingress, service mesh), storage (PV/PVC, CSI drivers), RBAC, and troubleshooting pod/node-level issues in both cloud and bare-metal environments

  • Strong Terraform skills with experience managing multi-environment, multi-provider infrastructure at scale

  • Hands-on experience building and operating CI/CD pipelines end-to-end — not just configuring someone else's templates

  • Production experience with at least three of: Prometheus, Grafana, Loki, Splunk, Datadog, Timestream, or comparable observability tools

  • Solid scripting and automation skills in Python, Bash, or Go

  • Working knowledge of Linux systems internals: networking (iptables, DNS, TCP debugging), storage, process management, and performance analysis

  • Experience managing TLS/mTLS certificates, secret rotation, and Vault or equivalent in production

  • Comfort working across cloud (AWS/GCP) and on-prem environments with different constraints and access models

  • Strong debugging instincts — you can follow a problem from a user report through load balancers, ingress, service mesh, application logs, and database queries to root cause


Nice to have



  • Experience supporting network infrastructure or datacenter automation platforms

  • Helm chart authoring and management for complex multi-service applications

  • Bare-metal Kubernetes provisioning (not just managed EKS/GKE)

  • eBPF-based observability tools (Cilium, Pixie, Hubble)

  • Experience operating Kafka, ClickHouse, ArangoDB, or Redis in production

  • Familiarity with SONiC, network switch management, or ZTP workflows

  • On-call experience with a structured incident management process (PagerDuty, Opsgenie)

  • SOC 2, FedRAMP, or equivalent compliance experience for infrastructure


 

$263,000 - $284,000 a year

Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.


Equal Opportunity


Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.


Accessibility & Accommodations


We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at [email protected]—we’re happy to help. Note: This inbox is only for accommodation requests.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Original job Senior Principal DevOps Engineer Orchestration posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Senior Principal DevOps Engineer Orchestration Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.