Staff AI Infrastructure Engineer
Redwood City
Workplace: HybridFull timeUSD 235,000 - 353,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Technical leadership","Judgment","Problem-solving","Debugging","Automation"]Own the reliability of Luma’s 10k+ GPU fleet by architecting and operating large, heterogeneous GPU environments under extreme demand. Lead root-cause and remediation across hardware, Linux, runtimes, and orchestration to eliminate instability and improve utilization, performance, and latency. Partner with research to scale inference for new model capabilities, and set reliability standards by hiring and developing systems and reliability engineers.
Loading
Loading job details...
Preparing the role view and application actions.

