Pay: $140,000.00 - $240,000.00 per year
Why This Is a Great Opportunity
- Own the technical architecture behind next-generation GPU and AI compute deployments.
- Design sophisticated GPU clusters from client requirements through production-ready architecture.
- Work directly with NVIDIA Reference Architecture, high-speed networking, storage, connectivity, and availability strategy.
- Solve challenging infrastructure problems where performance, reliability, power, cooling, and hardware constraints all matter.
- Have significant technical ownership over designs supporting enterprise and neocloud deployments.
- Collaborate closely with deployment, data center, supply chain, and program leadership to turn architecture into real-world infrastructure.
- Join a fast-growing AI infrastructure environment where your technical decisions directly impact customer outcomes.
- Competitive bonus and equity opportunity in addition to base compensation.
Location: Remote nationwide, with preference for candidates based in the US. Travel to data center partner sites is required.
Note: Must have 7+ years of directly relevant experience in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure. Candidates must have hands-on GPU cluster design experience plus strong InfiniBand, RoCE, or high-speed Ethernet networking expertise. Generic cloud architecture or enterprise networking experience without meaningful GPU/HPC infrastructure exposure will not meet the requirements.
About Us
We are building next-generation AI infrastructure that gives enterprises access to high-performance GPU compute with speed, flexibility, and reliability. Our technical teams design and deploy sophisticated GPU clusters across data center environments, and we are looking for an architect who can turn demanding customer requirements into robust, buildable infrastructure. Confidential Employer.
Job Description
- Own end-to-end technical architecture for GPU cluster deployments from customer requirements through deployment-ready design.
- Design GPU cluster configurations spanning compute, storage, networking, software requirements, and supporting infrastructure.
- Translate client technical requirements into complete bills of design covering all required compute, storage, networking, and connectivity components.
- Apply NVIDIA Reference Architecture principles, including HGX and NVL72-based GPU cluster designs.
- Design high-performance network fabrics using InfiniBand, RoCE, and high-speed Ethernet based on workload and performance requirements.
- Incorporate internet, VPN, firewall, dedicated circuit, protected optical, and other connectivity requirements into cluster architectures.
- Develop hot and cold sparing strategies designed to meet contracted availability and SLA commitments.
- Adapt cluster designs to site-specific power, cooling, space, hardware, and deployment constraints.
- Partner with data center teams to account for real-world facility limitations when finalizing technical architecture.
- Work with Supply Chain to ensure architecture decisions align with realistic hardware availability and lead times.
- Partner with deployment leadership and program management to translate designs into executable build plans.
- Support acceptance test planning and define technical criteria that validate the deployed architecture against the approved design.
- Evaluate and incorporate high-speed shared storage solutions such as Weka, VAST Data, and DDN where appropriate.
- Maintain technical ownership of architecture decisions while balancing performance, availability, cost, schedule, and operational supportability.
Qualifications
- 7+ years of experience in solutions architecture, network engineering, systems engineering, or similar roles supporting GPU, HPC, or large-scale compute infrastructure.
- Deep working knowledge of NVIDIA Reference Architecture and GPU cluster design principles.
- Hands-on experience designing InfiniBand, RoCE, and/or high-speed Ethernet fabrics.
- Proven experience designing GPU or HPC clusters rather than solely consuming cloud infrastructure.
- Experience developing sparing and spares strategies for mission-critical infrastructure.
- Experience integrating firewalls, VPNs, dedicated circuits, protected optical connectivity, and related networking requirements into infrastructure designs.
- Experience with high-speed shared storage technologies such as Weka, VAST Data, or DDN.
- Strong understanding of compute, storage, networking, and data center infrastructure dependencies.
- Ability to translate complex customer requirements into complete, practical, buildable technical architectures.
- Strong cross-functional communication and documentation skills.
- Experience supporting enterprise customers or neocloud deployments is preferred.
- NVIDIA NCP program or certification experience is a plus.
- Experience with capacity planning or sparing modeling tools is a plus.
Why You Will Love Working Here
- Work on technically challenging GPU infrastructure projects at the center of the AI compute market.
- Own architecture decisions that directly influence performance, reliability, scalability, and customer success.
- Gain exposure to cutting-edge NVIDIA GPU architectures and high-speed networking technologies.
- Collaborate with experienced infrastructure, deployment, data center, supply chain, and executive teams.
- Work remotely while remaining closely connected to real-world data center deployments.
- Opportunity to help establish repeatable architecture standards as the business scales.
- Competitive base compensation plus bonus and equity.
- Make a visible impact in a high-growth environment where strong technical judgment is valued.
JPC-1766
Benefits:
- Dental insurance
- Life insurance
- Paid time off
- Retirement plan
- Vision insurance
Requirements: 7+ yrs solutions architecture, network eng, or systems eng for GPU, HPC, or large-scale compute
- Required Years as Associate, 7+
- Additional context: Relocation- [None; remote nationwide]. Packages - [no]
- Salary Bands (Only if attorney role):
Submission Email: Kyle and Jeffrey - [email protected];[email protected]
Quick Recruiter Reference
Recruiters Submission: To submit, cancel - Lead GPU Cluster Solutions Architect - Axe Compute - JPC- 1766 - source
**New Job Order Alert**
- Client job title: Lead GPU Cluster Solutions Architect
- Location: Nationwide
- On-site, hybrid, remote: Remote
- Experience: 7+
- Good fit job titles/keywords for candidates: GPU Solutions Architect, GPU Cluster Architect, HPC Solutions Architect, NVIDIA Solutions Architect, GPU Infrastructure Architect, Network Architect, HPC Network Engineer, GPU Systems Engineer, InfiniBand, RoCE, HGX, NVL72, Neocloud
- # of hires needed: 1