Staff AI Infrastructure Engineer

Luma
Redwood City
Workplace: HybridFull timeUSD 235,000 - 353,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Technical leadership","Judgment","Problem-solving","Debugging","Automation"]

Own the reliability of Luma’s 10k+ GPU fleet by architecting and operating large, heterogeneous GPU environments under extreme demand. Lead root-cause and remediation across hardware, Linux, runtimes, and orchestration to eliminate instability and improve utilization, performance, and latency. Partner with research to scale inference for new model capabilities, and set reliability standards by hiring and developing systems and reliability engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Luma
Luma
2 weeks ago

Staff AI Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Own the reliability of Luma’s 10k+ GPU fleet by architecting and operating large, heterogeneous GPU environments under extreme demand. Lead root-cause and remediation across hardware, Linux, runtimes, and orchestration to eliminate instability and improve utilization, performance, and latency. Partner with research to scale inference for new model capabilities, and set reliability standards by hiring and developing systems and reliability engineers.
Location: Redwood City
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Architect and operate large, heterogeneous GPU environments under extreme demand to improve utilization and performance.
  • •Resolve failures across hardware, OS, runtimes, and orchestration, eliminating recurring instability classes.
  • •Define scheduling, placement, and resource management as clusters scale in size and concurrency.
  • •Partner with research to build systems for new model capabilities and scale inference without sacrificing reliability or latency.
  • •Hire and develop systems and reliability engineers; set reliability direction and production ownership standards.

Pay and Benefits

Salary: USD 235,000 - 353,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Deep expertise in Linux and distributed systems.
  • •Experience operating GPU or accelerator clusters in real production environments.
  • •Fluency in Kubernetes and modern open-source infrastructure.
  • •Strong debugging across hardware, kernel, runtime, and orchestration under contention at scale.
  • •Ability to write code and build automation with judgment engineers trust when systems fail.
Experience:GPU infrastructureDistributed systemsKubernetesReliability engineering
Skills:Technical leadershipJudgmentProblem-solvingDebuggingAutomation
Tech Stack:LinuxDistributed systemsGPUKubernetesContainersKernelsNetworkingStorageOrchestrationSchedulersHardware

Company Brief

Luma
Develops AI-powered tools for capturing, editing, and rendering high-quality 3D scenes from photos and videos, enabling creators to generate photorealistic 3D assets and spatial experiences.
Industry: AR/VR & Spatial Computing
Website