Senior AI Infrastructure Engineer, Physical Infrastructure

Anduril
United States
Workplace: OnsiteFull timeUSD 166,000 - 220,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Ownership","Automation","Fault-tolerant operations","Troubleshooting","Collaboration"]

Lead the long-term stability and execution of GPU cluster infrastructure for large-scale model training at Anduril. Own hardware and network robustness, build self-healing mechanisms, and replace manual triage with automated deployment tooling and deep observability. Rack, cable, and validate GPU systems, tune high-speed interconnects and parallel storage, and operate Kubernetes/Run:AI/Ray environments to enable resilient, multi-tenant scheduling for research and engineering teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anduril
Anduril
2 days ago

Senior AI Infrastructure Engineer, Physical Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Lead the long-term stability and execution of GPU cluster infrastructure for large-scale model training at Anduril. Own hardware and network robustness, build self-healing mechanisms, and replace manual triage with automated deployment tooling and deep observability. Rack, cable, and validate GPU systems, tune high-speed interconnects and parallel storage, and operate Kubernetes/Run:AI/Ray environments to enable resilient, multi-tenant scheduling for research and engineering teams.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Rack, stack, cable, and bring up GPU compute systems, including validating physical topology, power, cooling, firmware/BIOS, and burn-in.
  • •Build and tune interconnect fabrics connecting hundreds of GPUs for low-latency training and inference clusters.
  • •Integrate high-performance parallel storage to meet throughput needs for distributed training and terabyte-scale datasets.
  • •Automate end-to-end cluster deployment and configuration (including infrastructure as code) to bring new capacity online with minimal manual work.
  • •Operate and extend the Kubernetes/Run:AI environment for GPU scheduling, quota management, and multi-tenant workload isolation while monitoring, alerting, and rapidly triaging hardware and network faults.

Pay and Benefits

Salary: USD 166,000 - 220,000 annually

Key Requirements

  • •10+ years in hands-on infrastructure, HPC, or datacenter engineering supporting GPU compute at scale.
  • •Hands-on experience bringing up H200/B200/B300 (or comparable) GPU systems, including cabling and firmware/driver management.
  • •Experience with high-performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of hundreds of GPUs.
  • •Experience with high performance parallel storage (VAST, DDN, Weka, Lustre, or similar) for distributed training and multi-terabyte datasets.
  • •Kubernetes required; Run:ai (or similar GPU scheduling/orchestration) strongly preferred, with an automation-first mindset and ability to do physical datacenter work (50+ lbs).
Experience:GPU computeHPCDatacenterAI/MLDistributed training
Skills:OwnershipAutomationFault-tolerant operationsTroubleshootingCollaboration
Languages:En
Tech Stack:GPUsH200B200B300NVL72NCCLHigh-speed networkingKubernetesRun:AIRayNVLinkInfiniBandRoCESpectrum-XVASTDDNWekaLustreInfrastructure as codeFirmware/BIOS

Eligibility

Security Clearance:U.S. Top Secret

Company Brief

Anduril
Designs and builds advanced defense systems combining autonomous aircraft, sensors, and AI-driven software for military and national security applications, focused on modernizing battlefield capabilities and distributed sensing.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Costa Mesa, United States
Founded: 2017
WebsiteLinkedIn