Tech Lead Manager, Inference

Luma
Redwood City
Workplace: HybridFull timeFunction: Administration & Executive AssistanceSkills: ["Technical leadership","Coaching","Incident response","Setting technical direction","Capacity planning"]

Lead the team responsible for Luma’s end-to-end inference serving stack, including routing, scheduling, and fleet-wide orchestration across thousands of GPUs. Spend at least half your time hands-on architecting and debugging core platform components while hiring, coaching, and setting the technical roadmap. Own serving SLOs and economics (latency, availability, utilization, cost), and partner with research to deploy new model architectures into production.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Luma
Luma
2 weeks ago

Tech Lead Manager, Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Lead the team responsible for Luma’s end-to-end inference serving stack, including routing, scheduling, and fleet-wide orchestration across thousands of GPUs. Spend at least half your time hands-on architecting and debugging core platform components while hiring, coaching, and setting the technical roadmap. Own serving SLOs and economics (latency, availability, utilization, cost), and partner with research to deploy new model architectures into production.
Location: Redwood City
Workplace: Hybrid
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Manager level

Key Responsibilities

  • •Lead the inference serving stack ownership (routing, scheduling, and fleet-wide orchestration) and spend at least half your time hands-on architecting, building, and debugging core platform components.
  • •Hire, grow, and develop the inference engineering team, including coaching, on-call coverage, incident response, capacity planning, and postmortems.
  • •Set the technical roadmap for serving engines, routing, scheduling, autoscaling, caching, observability, and deployment.
  • •Own platform SLOs and economics, including latency, availability, GPU utilization, and cost per generation.
  • •Partner with research to ship new architectures to production and integrate serving into online RL and evaluation loops.

Key Requirements

  • •8+ years in large-scale distributed systems or ML infrastructure, with several years building and operating model-serving or inference platforms in production.
  • •Experience running inference at the thousands-of-GPUs scale across multiple clusters or clouds, including strong troubleshooting knowledge.
  • •Technical leadership experience with a genuine desire to stay at least half hands-on.
  • •Deep expertise in LLM and foundation-model serving engines (vLLM, SGLang, TensorRT-LLM), ideally with experience modifying engine internals.
  • •Strong command of continuous batching, KV-cache management, quantization, speculative decoding, and parallelism strategies (TP/EP/pipeline).
Experience:Distributed systemsML infrastructureModel servingInference platformsLLM serving
Skills:Technical leadershipCoachingIncident responseSetting technical directionCapacity planning
Tech Stack:PythonPyTorchKubernetesVLLMSGLangTensorRT-LLMContinuous batchingKV-cache managementQuantizationSpeculative decodingTPEPPipelineAutoscalingCachingObservabilityQueueingSchedulingTraffic controlFleet management

Company Brief

Luma
Develops AI-powered tools for capturing, editing, and rendering high-quality 3D scenes from photos and videos, enabling creators to generate photorealistic 3D assets and spatial experiences.
Industry: AR/VR & Spatial Computing
Website