Software Engineer, Inference

Luma
Redwood City
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["System architecture","Collaboration","Reliability focus","Automation"]

Own how models get served by integrating new architectures into the inference engine, scaling deployments across thousands of machines, and ensuring reliability against internal SLOs. Work on scheduling, fleet management, deployment pipelines, and CI/CD for model checkpoints and SDKs. Partner with research, engineering, and infrastructure to optimize model efficiency, build tooling to track inference job lifetimes, and automate inference services for maximum uptime.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Luma
Luma
2 weeks ago

Software Engineer, Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Own how models get served by integrating new architectures into the inference engine, scaling deployments across thousands of machines, and ensuring reliability against internal SLOs. Work on scheduling, fleet management, deployment pipelines, and CI/CD for model checkpoints and SDKs. Partner with research, engineering, and infrastructure to optimize model efficiency, build tooling to track inference job lifetimes, and automate inference services for maximum uptime.
Location: Redwood City
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Ship new model architectures by integrating them into the inference engine.
  • •Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
  • •Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
  • •Automate, test, and maintain inference services for maximum uptime and reliability.
  • •Manage and optimize inference workloads across clusters and hardware providers, including GPU-aware scheduling and CI/CD for model checkpoints and SDKs.

Key Requirements

  • •Strong Python and system-architecture skills.
  • •Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar.
  • •Experience with queues, scheduling, traffic control, and fleet management at scale.
  • •Experience with Linux, Docker, and Kubernetes, including orchestration, deployment, and scheduling.
  • •Familiarity with Redis and S3-compatible storage.
Experience:Model servingInference systemsLarge-scale MLGPU fleets
Skills:System architectureCollaborationReliability focusAutomation
Tech Stack:PythonPyTorchHugging FaceVLLMSGLangTensorRT-LLMLinuxDockerKubernetesRedisS3CI/CDQueuesSchedulingRDMARoCEInfiniBandNVLinkCUDAFFmpeg

Company Brief

Luma
Develops AI-powered tools for capturing, editing, and rendering high-quality 3D scenes from photos and videos, enabling creators to generate photorealistic 3D assets and spatial experiences.
Industry: AR/VR & Spatial Computing
Website