Cloud Inference Engineer

Luminal
San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 350,000 annuallyFunction: Software EngineeringSkills: ["CUDA","PyTorch","Torch","KV caching","Paged attention","Batching","Token streaming","TensorRT","VLLM","SGLang","GPU inference","Distributed compute"]

Join Luminal to build and optimize a high-performance AI inference stack for production models. You will deploy and tune models on Luminal Cloud, focusing on CUDA GPU inference, KV caching, paged attention, and low-latency serving. Collaborate on scheduler and autoscaling, profile latency and cost, and occasionally write kernels. On-site in downtown SF with a founding engineering role.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Luminal
Luminal
9 months ago

Cloud Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Join Luminal to build and optimize a high-performance AI inference stack for production models. You will deploy and tune models on Luminal Cloud, focusing on CUDA GPU inference, KV caching, paged attention, and low-latency serving. Collaborate on scheduler and autoscaling, profile latency and cost, and occasionally write kernels. On-site in downtown SF with a founding engineering role.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
  • •Conducting model performance reviews
  • •Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
  • •Sometimes write kernels and, yes, occasional tasteful shitposting
  • •Ship low latency, high throughput model serving on Luminal Cloud

Pay and Benefits

Salary: USD 150,000 - 350,000 annually

Key Requirements

  • •CUDA + GPU inference optimization (no degree required).
  • •vLLM, SGLang, or TensorRT-LLM experience.
  • •KV caching, paged attention, batching, token streaming, etc.
Experience:Machine learning
Skills:CUDAPyTorchTorchKV cachingPaged attentionBatchingToken streamingTensorRTVLLMSGLangGPU inferenceDistributed compute
Tech Stack:CUDAPyTorchTorchKV cachingPaged attentionBatchingToken streamingTensorRTVLLMSGLangGPU inferenceDistributed compute

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Luminal
Builds AI-driven video analytics software that analyzes surveillance and camera feeds to detect incidents, enhance safety, and optimize operations for enterprise and public sector environments.
Industry: Cybersecurity
Website