N

Principal Deployment Engineer

icon building Company : Nscale
icon briefcase Job Type : Full Time

Number of Applicants

 : 

000+

Click to reveal the number of candidates who applied for this job.
icon loader
icon loader

Let AI Supercharge Your Job Hunt!

JobCopilot scans 500,000+ company career sites daily to find jobs for you

Never miss an opportunity Save hours by auto-filling applications forms Land more interviews with tailored applications
happy man
thunder iconActivate JobCopilot

Job Description - Principal Deployment Engineer


.


Principal Deployment Engineer – GPU Supercluster Bringup


About Us


We are building AI infrastructure for frontier-scale workloads. Our platform is designed for high-density, high-performance GPU clusters that push the limits of power, networking, and distributed compute.


As a startup, we move fast, operate with ownership, and expect technical leaders to define standards—not just follow them.


The Role


We are hiring a Principal Deployment Engineer to architect and lead the bringup of large-scale GPU clusters (hundreds to thousands of GPUs). This is a technical leadership role responsible for defining how we deploy, validate, and scale AI superclusters across sites.


You will own the full lifecycle of deployment—from rack design and fabric architecture to cluster validation frameworks and production readiness standards. You will set the bar for performance, reliability, and operational excellence.


This role combines deep hands-on expertise with system-level thinking and cross-functional leadership.


What You’ll Do


End-to-End Supercluster Bringup Ownership



  • Define the technical standards for node, rack, and full-cluster bringup.


  • Lead large-scale GPU cluster deployments (multi-rack, multi-pod environments).


  • Architect high-performance network fabrics (IB, RoCE, Ethernet) optimized for AI workloads.


  • Establish cluster-level acceptance criteria and validation frameworks.



Performance & Fabric Architecture



  • Tune and validate NCCL, RDMA, GPUDirect, and collective operations at scale.


  • Identify and eliminate performance bottlenecks across hardware, topology, and firmware layers.


  • Drive congestion control and fabric optimization strategies.


  • Define performance benchmarking methodology for AI training workloads.



Deployment Strategy & Scalability



  • Design repeatable deployment models for multi-site expansion.


  • Build automation frameworks for provisioning and cluster validation.


  • Establish deployment SLAs, quality gates, and operational readiness standards.


  • Reduce time-to-capacity while increasing reliability.



Technical Leadership



  • Serve as the escalation point for complex bringup and performance issues.


  • Mentor senior engineers and shape infrastructure best practices.


  • Influence hardware selection, rack topology, and data center design decisions.


  • Partner with executive leadership on infrastructure scaling strategy.



What We’re Looking For


Required



  • 10+ years of experience in large-scale infrastructure or HPC environments.


  • Proven experience bringing up large GPU clusters (hundreds+ GPUs).


  • Deep expertise in high-speed networking (InfiniBand, RoCE, Ethernet fabrics).


  • Strong understanding of server architecture (PCIe, NUMA, memory hierarchy).


  • Experience debugging performance issues across compute and network layers.


  • Strong automation and systems-level thinking.



Strongly Preferred



  • Experience scaling AI training clusters for frontier models.


  • Experience with liquid cooling or ultra-high-density deployments.


  • Knowledge of distributed storage systems (Lustre, Ceph, NVMe-oF).


  • Experience defining infrastructure standards in a fast-growing organization.



What Success Looks Like



  • Superclusters are brought online quickly, predictably, and at peak performance.


  • Deployment processes scale from first cluster to multi-site expansion.


  • Infrastructure becomes a competitive advantage.


  • You define the technical blueprint for how we scale AI infrastructure.


For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.


Original job Principal Deployment Engineer posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Auto-Apply to Deployment Engineer Jobs with your AI JobCopilot

thunder icon Auto-Apply with AI

Similar Deployment Engineer Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.