Senior ML Systems Engineer, Inference

RunPod
United States
Workplace: RemoteFull timeUSD 150,000 - 220,000 annuallyFunction: IT Operations (Systems/Network Admin)Skills: ["Ownership","Problem-solving","Technical writing","Benchmarking rigor","Decision-making"]

Own end-to-end LLM inference serving performance for a developer-focused AI infrastructure platform. Define rigorous metrics, profile and diagnose bottlenecks across the serving stack (from scheduling and memory to kernels and interconnect), and improve efficiency for single- and multi-node GPU deployments. Turn insights into production-ready runtimes and configurations while partnering with product and infrastructure teams to shape how inference is offered.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
RunPod
RunPod
2 days ago

Senior ML Systems Engineer, Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Own end-to-end LLM inference serving performance for a developer-focused AI infrastructure platform. Define rigorous metrics, profile and diagnose bottlenecks across the serving stack (from scheduling and memory to kernels and interconnect), and improve efficiency for single- and multi-node GPU deployments. Turn insights into production-ready runtimes and configurations while partnering with product and infrastructure teams to shape how inference is offered.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Define and operationalize how inference performance is measured, including throughput, time to first token, inter-token latency, and cost per token.
  • •Profile and diagnose performance issues across the serving stack, from scheduling/memory management to kernels and interconnect.
  • •Improve serving efficiency for large state-of-the-art models on single-node and multi-node GPU deployments.
  • •Convert learnings into production-ready runtimes, configurations, and defaults customers benefit from automatically.
  • •Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning isn’t enough.

Pay and Benefits

Salary: USD 150,000 - 220,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPaid LeaveEquityHome Office

Key Requirements

  • •5+ years of professional system engineering experience.
  • •Deep, hands-on production experience with vLLM, SGLang (or comparable serving engine) at benchmark scale.
  • •Strong Python software engineering skills in performance-critical codebases.
  • •Solid understanding of LLM inference performance drivers (batching, memory, parallelism, latency/throughput trade-offs).
  • •Experience with inference optimization techniques (quantization, speculative decoding, or distributed serving) and GPU profiling/benchmarking rigor.
Experience:AI infrastructureLLM inferenceGPUDistributed systems
Skills:OwnershipProblem-solvingTechnical writingBenchmarking rigorDecision-making
Tech Stack:PythonVLLMSGLangCUDATritonQuantizationSpeculative decodingGPU profiling toolsKernels

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

RunPod
Provides on-demand GPU cloud and marketplace services for machine learning workloads, offering rentable GPU instances, scalable compute for training and inference, and tools to run ML jobs cost-effectively without long-term commitments.
Industry: Cloud Computing
Website