Logo-of-Upscaleai-hiring-for-jobs-in-US-on-GrabJobs

Senior Staff DevOps Engineer Scaleup

salary Salary :

$222,000 - 250,000 yearly

Job Description - Senior Staff DevOps Engineer Scaleup

Why join Upscale AI


Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.


We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.


If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.



About the role


AI training and inference clusters live or die on the network fabric. ScaleUp builds the software that programs our high-performance Ethernet switch silicon — SDK, SAI, simulation, and the CI that keeps that stack shippable.


We’re looking for a DevOps engineer at the intersection of AI infrastructure, data-center networking, and silicon-aware software: reliable pipelines for an ASIC SDK, fast feedback for developers, and release paths worthy of cloud- and AI-scale deployment.


You’ll partner with SDK, SAI, and QA engineers, and work in GitHub + Jira day to day. Success means CI is trusted, simulation and software gates are clear, and infrastructure never blocks the next generation of AI networking features.


Responsibility



  • Own CI/CD for the ScaleUp stack that powers AI/HPC Ethernet switching (build, unit/integration, gating, artifacts, release promotion).

  • Scale pipelines across multi-repo dependencies: switch SDK, SAI adapters, shared test frameworks, and switch simulation / model targets.

  • Improve build caching, shared runners, and nightlies so large C/C++ and Python SDK builds stay fast and predictable.

  • Own release / promote workflows (branch policies, artifact publish, pre-gate vs post-gate steps) so SDK drops are repeatable.

  • Shorten commit → green for teams building programmable data-plane, QoS, ACL, RoCE-era fabrics, and AI-workload traffic patterns.

  • Add and maintain quality gates: lint, unit tests, model smoke, coverage where useful, overnight soak — high signal, low flake.

  • Manage artifacts and caches (CI artifacts, object storage as needed) with clear retention and failure handling.

  • Operate Linux build/test farms for C/C++ SDKs and Python harnesses (toolchains, containers, lab/sim hosts).

  • Improve pipeline observability (dashboards, alerts, runbooks, postmortems).

  • Secure secrets and access for GitHub Actions, artifact stores, and internal services.

  • Partner with QA on PR / nightly / soak jobs that validate end-to-end switch behavior before customer and cloud deployments.

  • Publish reusable pipeline templates so new ScaleUp / AI-networking repos onboard quickly.

  • Jira: keep engineering work traceable — epics/stories/bugs for CI and infra, link PRs and releases to tickets, support sprint and release planning with SDK/SAI/QA, tighten ticket hygiene (status, components, labels) so blockers and CI debt are visible.


Qualification



  • Strong CI/CD experience (GitHub Actions and/or Jenkins; GitLab CI also fine) in multi-repo environments.

  • Deep Linux fluency: shells, packaging, toolchains, debugging native builds and shared-library / SDK load paths.

  • Automation in Python and bash; enough CMake/Make to unblock switch-SDK CI.

  • Containers (Docker) and self-hosted or cloud runners at meaningful scale.

  • Proven work reducing flaky tests and improving CI signal-to-noise.

  • Hands-on Jira (or equivalent): workflows, boards, linking commits/PRs to issues, release/version fields — comfortable driving process with developers, not only “keeping the lights on.”

  • Clear communication with hardware-adjacent software teams.

  • Excitement about AI data centers, Ethernet fabrics, and silicon-software co-design — not only generic cloud DevOps. 


    Nice to have



    • Background in networking ASICs, switch SDKs, SAI, DPDK, or NIC/SmartNIC software.

    • Exposure to AI/ML cluster networking (GPU fabrics, RoCE/RDMA, congestion control, telemetry).

    • Simulation / hardware-in-the-loop CI, pytest at scale, junit/Allure-style reporting.

    • Build-cache and large-monorepo / multi-repo release patterns; backport or branch-gating workflows.

    • Static analysis / lint gates in CI (e.g. cppcheck, language linters).

    • IaC (Terraform/Ansible) and cloud runners (AWS/GCP).

    • Atlassian suite beyond Jira (Confluence runbooks, Jira + GitHub automation).

    • Release hygiene: versioning, artifacts, SBOM/signing, branch protection / merge gates.


    Why ScaleUp



    • Your pipelines sit under real AI networking product software — not a side CRUD service.

    • You influence how fast we ship switch SDK + SAI features that AI clusters depend on.

    • Small team, high ownership: changes land and developers feel them the next day.




$222,000 - $250,000 a year

Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.


Equal Opportunity


Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.


Accessibility & Accommodations


We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at [email protected]—we’re happy to help. Note: This inbox is only for accommodation requests.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Original job Senior Staff DevOps Engineer Scaleup posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Senior Staff DevOps Engineer Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.