Inference Performance Engineer
San Francisco, Singapore, London, Mexico City, Paris, India, Dublin, Berlin, Netherlands, São Paulo, Argentina, Toronto
Workplace: RemoteFull timeFunction: Software EngineeringSkills: ["Cost optimization","Performance optimization","Profiling","Measurement"]Own the cost and performance of the inference stack, driving improvements in throughput and latency as workloads, traffic, and hardware evolve. Partner with engineers operating the serving fleet while managing key performance levers such as KV-cache, continuous batching, speculative decoding, and quantization. Optimize prefill/decode for long-context workloads, tune routing across infrastructure and providers, and build profiling systems. Work directly in serving engines like vLLM, SGLang, and TensorRT-LLM.
Loading
Loading job details...
Preparing the role view and application actions.

