Senior Research Scientist | Model Scaling

DeepL
London, Cologne, Munich
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Communication","Collaboration","Experimentation","Debugging"]

Own foundational modeling decisions for next-generation translation models. Drive selection and evaluation of open foundation/open-weight models, lead architecture and scaling choices (including Mixture-of-Experts), and design multi-capability adaptation using LoRA/PEFT. Manage the full modeling lifecycle—prototyping, ablations, scaling experiments, evaluation, and production delivery—while partnering with post-training, RL/RLHF, and instruction-following experts.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
DeepL
DeepL
1 month ago

Senior Research Scientist | Model Scaling

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Own foundational modeling decisions for next-generation translation models. Drive selection and evaluation of open foundation/open-weight models, lead architecture and scaling choices (including Mixture-of-Experts), and design multi-capability adaptation using LoRA/PEFT. Manage the full modeling lifecycle—prototyping, ablations, scaling experiments, evaluation, and production delivery—while partnering with post-training, RL/RLHF, and instruction-following experts.
Location: London, Cologne, Munich
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Drive selection and evaluation of open foundation/open-weight models for next-generation translation systems.
  • •Lead model selection and architecture decisions for scaling to hundreds of billions of parameters, including Mixture-of-Experts and sparse/efficient designs.
  • •Design multi-capability adaptation strategies using LoRA, PEFT, and related methods.
  • •Own the modeling lifecycle—prototyping, ablations, scaling experiments, evaluation, and delivery into production—with rigorous, reproducible evaluation.
  • •Partner with post-training, RL/RLHF, and instruction-following specialists to integrate alignment and capability work into the base model.

Pay and Benefits

Perks:Paid Leave

Key Requirements

  • •Strong hands-on experience adapting and scaling large language models via fine-tuning, instruction-tuning, or post-training on multi-billion-parameter models.
  • •Sound judgment about architecture trade-offs at scale (e.g., dense vs. MoE) and about which open-weight foundation models to build on.
  • •Working knowledge of parameter-efficient and multi-capability adaptation (LoRA/PEFT and related methods).
  • •Hands-on ability to train models, run experiments, debug pipelines, and carry research results into production with engineering.
  • •Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow).
Skills:CommunicationCollaborationExperimentationDebugging
Tech Stack:PythonPyTorchJAXTensorflowLoRAPEFTRLHFMixture-of-ExpertsMixture-of-LoRA-Experts

Company Brief

DeepL
DeepL builds Language AI products (DeepL Translator, DeepL Write, APIs and enterprise solutions) that provide high-accuracy translations and writing assistance to businesses and individuals, focusing on privacy, security and enterprise deployment.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Cologne, Germany
Founded: 2017
Glassdoor
Glassdoor: 3.5
WebsiteLinkedInGlassdoor