Job Description - ML Infrastructure Engineer

ML Infrastructure Engineer

Company: Dyna Robotics
Location: Redwood City, CA (in office 5 days per week)
Compensation: $220,000 - $350,000 + competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers (OPT, H-1B transfer)

About Dyna Robotics

Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model that generalizes and self-improves across environments with commercial-grade performance. Its affordable, intelligent robotic arms are already deployed at customer sites in hospitality and restaurants, automating repetitive, stationary tasks.

Founded in 2024 by repeat founders who previously built and sold Kaper AI to Instacart, with a team from Google DeepMind, Meta and Cruise, Dyna has about 130 people and has raised $143.5M from investors including NVentures, Samsung NEXT, Salesforce Ventures, First Round Capital and CRV.

The Role

Dyna Robotics is hiring an ML Infrastructure Engineer to own training infrastructure end to end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You will be the connective tissue between researchers and compute, and your work directly speeds the path from model to deployed robot.

What You Will Do

  • Architect and scale distributed training across large GPU clusters, implementing sharding, activation checkpointing and memory optimization (ZeRO, FSDP).
  • Build researcher-friendly tooling and job scheduling (Kubernetes, SLURM) with fast iteration, automated retries and failure recovery.
  • Design high-throughput pipelines that ingest terabytes of multimodal robot data (video, proprioception, 3D signals) so GPUs never starve.
  • Build low-latency inference pipelines for real-time robot control using quantization, distillation and compilation (TensorRT, Triton).
  • Profile GPU utilization, I/O bottlenecks and memory fragmentation to maximize fleet performance.

What You Bring

  • 5-7+ years as an infrastructure engineer, including leading technical projects in HPC or ML infrastructure
  • Built and maintained ML or data infrastructure on a team with a high talent bar
  • Deep PyTorch experience and hands-on large-scale distributed training (FSDP, ZeRO, failure recovery)
  • GPU performance optimization and profiling (CUDA, NCCL, Triton)
  • Genuine interest in robotics and physical AI
  • Ability to work in the Redwood City office 5 days a week

Nice to Have

  • Robotics experience at startups or enterprise teams
  • Early-stage or founding infrastructure hire
  • Multimodal systems (video, audio, multimedia models)
  • DeepSpeed or Accelerate; model serving optimization and monitoring

Interview Process

Recruiter screen (30 min), system design (45 min), two coding rounds (45 and 30 min), behavioral.

Tech Stack

PyTorch, DeepSpeed, Accelerate, FSDP, Kubernetes, SLURM, GCP, AWS, TensorRT, Triton, NCCL, Docker, Python, CUDA


REVENUE: 21.25% of first-year salary. Est. fee per hire $47K-$74K; 4 seat(s) = up to $242K if all filled.

TARGET COMPANIES (suggested): Google DeepMind, Tesla (Optimus), Physical Intelligence, Figure AI, Cruise, Waymo, NVIDIA.

BEST-FIT CANDIDATE: 5-7+ yrs; large-scale PyTorch distributed training infra (FSDP/ZeRO); multimodal data pipelines at TB scale; GPU perf profiling; visa: transfers only; location: Redwood City 5 days. Avoid pure ML modelers and inference/DevOps-only profiles; robotics passion strongly preferred.

Original job ML Infrastructure Engineer posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar ML Infrastructure Engineer Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.