Research Engineer, RL Scaling Science

Anthropic
London
Workplace: OnsiteFull timeGBP 375,000 - 640,000 annuallyFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication","Collaboration","Problem-solving"]

Join the RL Scaling Science team to design and run large-scale reinforcement learning experiments, build benchmarks for long-horizon RL, and ship validated findings directly into production training. Work at the research/engineering boundary on frontier-scale problems, collaborating with cross-functional RL teams to advance training across model sizes, horizons, and compute, with impact on production systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
2 months ago

Research Engineer, RL Scaling Science

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Join the RL Scaling Science team to design and run large-scale reinforcement learning experiments, build benchmarks for long-horizon RL, and ship validated findings directly into production training. Work at the research/engineering boundary on frontier-scale problems, collaborating with cross-functional RL teams to advance training across model sizes, horizons, and compute, with impact on production systems.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Design, run, and interpret large-scale RL experiments, reasoning rigorously about what the data does and doesn’t show.
  • •Investigate how RL improves as horizon, compute, and model size grow.
  • •Build and maintain benchmarks for long-horizon RL so progress is measurable and reproducible.
  • •Translate validated findings into production training recipes, exercising judgment about when a result is robust enough to ship.
  • •Debug complex issues at the seam where research meets infrastructure - failures that only appear at scale.

Pay and Benefits

Salary: GBP 375,000 - 640,000 annually
Perks:Paid LeaveParental LeaveFlexible Hours

Key Requirements

  • •Strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area.
  • •Demonstrated ability to own large experiments end-to-end, from design through interpretation.
  • •Proficiency in Python and experience working with large-scale or distributed ML systems.
  • •Comfort operating at the research/systems boundary, including debugging where the two meet.
  • •Care about the societal impacts of AI and responsible scaling.
Experience:AiMachine learningReinforcement learning
Education:Bachelor's
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:Python

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn