Senior Staff Research Scientist | Voice

DeepL
London
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Hands-on","Collaboration","Communication","Mentoring","Problem-solving"]

Lead scientific innovation for real-time multilingual voice translation, advancing speech and translation models across ASR, MT, TTS, streaming inference, and large speech models. Own the end-to-end lifecycle from prototyping and training through evaluation, optimization, and production deployment. Collaborate closely with engineering to integrate models reliably at ultra-low latency, improve cascaded and end-to-end systems, and raise quality through mentoring and production monitoring.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
DeepL
DeepL
22 hours ago

Senior Staff Research Scientist | Voice

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live

Job Summary

Lead scientific innovation for real-time multilingual voice translation, advancing speech and translation models across ASR, MT, TTS, streaming inference, and large speech models. Own the end-to-end lifecycle from prototyping and training through evaluation, optimization, and production deployment. Collaborate closely with engineering to integrate models reliably at ultra-low latency, improve cascaded and end-to-end systems, and raise quality through mentoring and production monitoring.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Lead hands-on research and development across ASR, MT, TTS, and speech-to-speech translation for real-time voice products.
  • •Design, train, and optimize large-scale ASR models for multilingual accuracy, robustness, and ultra-low-latency streaming.
  • •Improve cascaded translation pipelines end to end, including segmentation, ASR→MT interfaces, streaming MT inference, and incremental decoding.
  • •Develop and refine real-time TTS models with natural prosody, stable speaker characteristics, and fast inference.
  • •Own the full model delivery lifecycle—prototyping, ablations, training, evaluation, optimization, and production deployment—and partner with engineering for integration and monitoring.

Key Requirements

  • •Deep expertise in speech/audio or multilingual ML, especially ASR, MT, TTS, end-to-end ST, or large speech models.
  • •Hands-on experience training models, running experiments, debugging ML pipelines, and integrating ML systems into production.
  • •Strong understanding of real-time streaming constraints and designing models for ultra-low latency.
  • •Experience shipping ML models to production and maintaining them at scale with deployment, monitoring, and serving.
  • •Strong coding and experimentation skills (Python, PyTorch/JAX, audio processing libraries).
Skills:Hands-onCollaborationCommunicationMentoringProblem-solving
Tech Stack:PythonPyTorchJAXASRMTTTSStreaming inferenceLarge speech modelsLLMAudio processing

Company Brief

DeepL
DeepL builds Language AI products (DeepL Translator, DeepL Write, APIs and enterprise solutions) that provide high-accuracy translations and writing assistance to businesses and individuals, focusing on privacy, security and enterprise deployment.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Cologne, Germany
Founded: 2017
Glassdoor
Glassdoor: 3.5
WebsiteLinkedInGlassdoor