Cloud Inference Engineer
San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 350,000 annuallyFunction: Software EngineeringSkills: ["CUDA","PyTorch","Torch","KV caching","Paged attention","Batching","Token streaming","TensorRT","VLLM","SGLang","GPU inference","Distributed compute"]Join Luminal to build and optimize a high-performance AI inference stack for production models. You will deploy and tune models on Luminal Cloud, focusing on CUDA GPU inference, KV caching, paged attention, and low-latency serving. Collaborate on scheduler and autoscaling, profile latency and cost, and occasionally write kernels. On-site in downtown SF with a founding engineering role.
Loading
Loading job details...
Preparing the role view and application actions.

