Logo-of-Upscaleai-hiring-for-jobs-in-US-on-GrabJobs

Senior Principal Engineer Orchestration

Job Description - Senior Principal Engineer Orchestration

Why join Upscale AI


Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.


We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.


If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.



The role


You will define the cross-cutting architecture that the engineering team builds within. You own platform-scale initiatives from vision through delivery — authoring architecture decision documents, building POCs to validate direction, and partnering with Product to shape the roadmap. You write code, you ship, and you are accountable for the quality of what reaches stakeholders.


Responsibilities



  • Distributed control plane spanning cloud services and on-prem edge appliances connected via mTLS gRPC streams

  • Observability and telemetry at scale: OpenTelemetry collection pipelines, stream processing (Kafka), time-series storage (ClickHouse/VictoriaMetrics), real-time fabric state views, packet-event analysis, and fleet-wide health aggregation

  • Agentic AI operations: design and build autonomous infrastructure agents that evaluate prerequisites, orchestrate multi-step workflows (onboarding, upgrades, drift remediation), handle failure recovery, and interact with the control plane through tool-use patterns (LangGraph, MCP)

  • Fleet orchestration: aggregate health/compliance/drift APIs, cross-site campaign execution, template promotion workflows, parallel site onboarding

  • Data architecture across Postgres, ArangoDB, ClickHouse, Redis, Kafka, and Git-backed content stores

  • Multi-tenant SaaS with site-scoped RBAC, session lifecycle, and enterprise IdP integration

  • Device lifecycle management: zero-touch provisioning, enrollment protocols, config push via edge relay, drift detection and remediation

  • Intent compilation engine: workspace management, merge request lifecycle, change request execution with gate-based verification and rollback


Qualifications



  • 15–18 years building and operating distributed systems serving enterprise customers across cloud and on-prem environments

  • Deep proficiency in Go (or comparable systems language) with strong distributed systems fundamentals

  • Significant experience designing observability and telemetry platforms: collection agents, stream processing, time-series databases, alerting pipelines, and real-time dashboards at scale

  • Production experience with microservices architecture: gRPC, Protocol Buffers, spec-first REST APIs (OpenAPI)

  • Hands-on with multiple storage paradigms: relational, graph, time-series, key-value, and streaming

  • Track record building multi-tenant platforms with tenant isolation, RBAC, and identity federation

  • Experience with Kubernetes, Helm, and hybrid cloud/on-prem deployment models

  • Strong API design sense: versioning, backward compatibility, contract-first development

  • Demonstrated ability to lead platform-scale technical initiatives across multiple teams and deliver on time

  • Ability to communicate architectural decisions clearly through writing and diagrams 



How we expect you to work



  • Cross-functional by default — you partner with Product to shape the roadmap, work with QA to define quality gates, and collaborate with Customer Engineering to ground decisions in real deployment reality

  • Solution-oriented — you don't just identify architectural problems, you build POCs to prove the path forward and help the team move

  • Accountable for delivery — you own outcomes across the initiatives you lead, not just the designs

  • Hands-on always — you write code, review the most consequential PRs, and prove your architecture works by building it


 


Nice to have



  • AI/ML agent architectures for infrastructure operations: LangGraph, AutoGen, MCP tool-use, human-in-the-loop gating, and autonomous workflow orchestration

  • Network automation or infrastructure management platforms

  • Datacenter networking experience: OpenConfig, gNMI, or fabric management at scale

  • Hub-and-spoke / edge computing / control-plane-data-plane separation architectures

  • OpenTelemetry contributor experience or deep familiarity with the collector ecosystem

  • Device enrollment or zero-touch provisioning systems

  • Open-source contributions or published work in distributed systems


 

$280,000 - $306,000 a year

Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.


Equal Opportunity


Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.


Accessibility & Accommodations


We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at [email protected]—we’re happy to help. Note: This inbox is only for accommodation requests.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Original job Senior Principal Engineer Orchestration posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Senior Principal Engineer Orchestration Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.