Senior AI Infrastructure Engineer, Physical Infrastructure
United States
Workplace: OnsiteFull timeUSD 166,000 - 220,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Ownership","Automation","Fault-tolerant operations","Troubleshooting","Collaboration"]Lead the long-term stability and execution of GPU cluster infrastructure for large-scale model training at Anduril. Own hardware and network robustness, build self-healing mechanisms, and replace manual triage with automated deployment tooling and deep observability. Rack, cable, and validate GPU systems, tune high-speed interconnects and parallel storage, and operate Kubernetes/Run:AI/Ray environments to enable resilient, multi-tenant scheduling for research and engineering teams.
Loading
Loading job details...
Preparing the role view and application actions.

