Inference Performance Engineer

Adaption Labs
San Francisco, Singapore, London, Mexico City, Paris, India, Dublin, Berlin, Netherlands, São Paulo, Argentina, Toronto
Workplace: RemoteFull timeFunction: Software EngineeringSkills: ["Cost optimization","Performance optimization","Profiling","Measurement"]

Own the cost and performance of the inference stack, driving improvements in throughput and latency as workloads, traffic, and hardware evolve. Partner with engineers operating the serving fleet while managing key performance levers such as KV-cache, continuous batching, speculative decoding, and quantization. Optimize prefill/decode for long-context workloads, tune routing across infrastructure and providers, and build profiling systems. Work directly in serving engines like vLLM, SGLang, and TensorRT-LLM.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Adaption Labs
Adaption Labs
4 days ago

Inference Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live

Job Summary

Own the cost and performance of the inference stack, driving improvements in throughput and latency as workloads, traffic, and hardware evolve. Partner with engineers operating the serving fleet while managing key performance levers such as KV-cache, continuous batching, speculative decoding, and quantization. Optimize prefill/decode for long-context workloads, tune routing across infrastructure and providers, and build profiling systems. Work directly in serving engines like vLLM, SGLang, and TensorRT-LLM.
Location: San Francisco, Singapore, London, Mexico City, Paris, India, Dublin, Berlin, Netherlands, São Paulo, Argentina, Toronto
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Improve throughput, cost, and tail latency using KV-cache management, continuous batching, speculative decoding, and quantization.
  • •Optimize long-context prefill and decode workloads using real production traffic.
  • •Tune routing between infrastructure and external providers based on cost, capacity, and performance.
  • •Work within serving engines (vLLM, SGLang, TensorRT-LLM), going below the framework when needed.
  • •Build profiling and measurement systems to identify where time, memory, and compute are spent.

Pay and Benefits

Perks:Travel AllowanceMeal AllowanceMedical BenefitsPaid Leave

Key Requirements

  • •5+ years in ML systems, inference infrastructure, or performance engineering with measurable cost or latency improvements.
  • •Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • •Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • •Strong Python skills and proficiency in C++, Rust, or another systems language.
  • •Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, and/or quantization.
Experience:ML systemsInference infrastructurePerformance engineering
Skills:Cost optimizationPerformance optimizationProfilingMeasurement
Tech Stack:PythonC++RustVLLMSGLangTensorRT-LLMKV-cacheContinuous batchingSpeculative decodingQuantizationCUDANCCLMixed precision

Company Brief

Adaption Labs
Adaption Labs delivers AI consulting, engineering, and MLOps services to help organizations design, build, and deploy machine learning and generative AI solutions. They focus on productizing models, integrating AI into production systems, and accelerating digital transformation for enterprise clients.
Industry: Consulting
Website