Senior Research Scientist | Model Steering

DeepL
London, Cologne, Munich
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Hands-on execution","Experimentation","Debugging","Mentoring","Communication"]

Lead fine-tuning, post-training, model stearability, and reinforcement learning to advance next-generation LLM-based translation models. Fuse human expert and synthetic data to make translation models follow custom user instructions, rules, and context. Build reward and evaluator models, mitigate reward hacking, and expand toward multimodal context. Own the full model lifecycle from prototyping and experiments to production deployment and continuous improvement, mentoring researchers and engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
DeepL
DeepL
1 month ago

Senior Research Scientist | Model Steering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Lead fine-tuning, post-training, model stearability, and reinforcement learning to advance next-generation LLM-based translation models. Fuse human expert and synthetic data to make translation models follow custom user instructions, rules, and context. Build reward and evaluator models, mitigate reward hacking, and expand toward multimodal context. Own the full model lifecycle from prototyping and experiments to production deployment and continuous improvement, mentoring researchers and engineers.
Location: London, Cologne, Munich
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Drive development of steerable translation models conditioned on user preferences, rules, and context.
  • •Conduct hands-on research and development for post-training, including supervised fine-tuning, knowledge distillation, preference optimization, and reinforcement learning tuned to translation quality.
  • •Build reward models and evaluator models for translation, including rubric/reference-based grading, and investigate and mitigate reward hacking and quality-estimation failure modes.
  • •Advance work toward translation models that ingest multimodal content and context to improve translation quality.
  • •Own the full model lifecycle—prototyping, ablations, training, evaluation, optimization, and production deployment—partnering with engineering; establish evaluation/reproducibility/monitoring practices and mentor others.

Pay and Benefits

Equity and Bonus:Equity
Perks:Hybrid WorkAnnual LeaveHealth Insurance

Key Requirements

  • •Proven experience making large models steerable and instruction-following using methods such as instruction tuning, latent space approaches, steering vectors, and/or constrained encoding/decoding.
  • •Deep hands-on expertise in LLM post-training, including SFT and DPO, knowledge distillation, and/or reinforcement learning (e.g., RLHF/RLAIF, PPO/GSPO, reward modeling).
  • •Strong data-centric skills for building synthetic-data and preference-data pipelines, including LLM-as-judge generation, curation/filtering, and reasoning about data mixtures and ablations.
  • •Experience designing evaluation and reward signals using automatic metrics, LLM-as-judge evaluation, non-verifiable rewards, and human-in-the-loop evaluation.
  • •Strong coding and experimentation skills (Python, PyTorch/JAX/TensorFlow) and the ability to communicate clearly and align research with product and engineering priorities.
Experience:LLMMachine learningReinforcement learning
Skills:Hands-on executionExperimentationDebuggingMentoringCommunication
Tech Stack:PythonPyTorchJAXTensorflowFSDPDeepSpeedMegatronVLLMSGLangTensorRT-LLMLLMReinforcement learningReward modeling

Company Brief

DeepL
DeepL builds Language AI products (DeepL Translator, DeepL Write, APIs and enterprise solutions) that provide high-accuracy translations and writing assistance to businesses and individuals, focusing on privacy, security and enterprise deployment.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Cologne, Germany
Founded: 2017
Glassdoor
Glassdoor: 3.5
WebsiteLinkedInGlassdoor