Research Engineer - Inference

ElevenLabs
London, New York, San Francisco, Warsaw, Bulgaria
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Skills: ["Problem-solving","Profiling","Diagnosis","Engineering judgment","Autonomous ownership"]

Own the systems that turn frontier AI research into fast, reliable, at-scale production serving. Deploy and optimize inference performance across the stack—latency, throughput, and cost—using techniques like quantization, distillation, KV-cache optimization, batching strategies, and custom kernels. Build high-performance serving systems for real-time, streaming workloads and create tooling/infrastructure so researchers can ship new models safely and quickly with clear performance measurements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ElevenLabs
ElevenLabs
4 days ago

Research Engineer - Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own the systems that turn frontier AI research into fast, reliable, at-scale production serving. Deploy and optimize inference performance across the stack—latency, throughput, and cost—using techniques like quantization, distillation, KV-cache optimization, batching strategies, and custom kernels. Build high-performance serving systems for real-time, streaming workloads and create tooling/infrastructure so researchers can ship new models safely and quickly with clear performance measurements.
Location: London, New York, San Francisco, Warsaw, Bulgaria
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Deploy state-of-the-art models to production and own the path from research checkpoint to serving infrastructure.
  • •Optimize inference performance across the stack, improving latency, throughput, and cost.
  • •Build and tune high-performance serving systems for real-time, streaming workloads where millisecond latency matters.
  • •Create tooling and infrastructure that enables researchers to ship new models to production quickly, safely, and confidently.
  • •Measure and improve performance characteristics using profiling and diagnostics across kernels and orchestration.

Pay and Benefits

Perks:Learning BudgetTravel AllowanceAnnual OffsiteCo-working Stipend

Key Requirements

  • •Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications.
  • •Strong GPU programming and inference optimization skills (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang).
  • •Ability to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack (model architecture, kernels, orchestration).
  • •Capacity to build tooling to measure and improve serving performance characteristics.
  • •Showcase solving hard problems via artifacts such as past projects, designs, or GitHub contributions.
Experience:MLAI
Skills:Problem-solvingProfilingDiagnosisEngineering judgmentAutonomous ownership
Tech Stack:CUDATritonTensorRTVLLMSGLangQuantizationDistillationKV-cache optimizationBatching strategiesCustom kernelsGPU programmingInference optimizationServing frameworks

Company Brief

ElevenLabs
Develops advanced AI audio models and tools for realistic text-to-speech, voice cloning, dubbing, music generation, and conversational voice agents for creators and enterprises.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: London, United Kingdom
Founded: 2022
Glassdoor
Glassdoor: 4.2
WebsiteLinkedInGlassdoor