Research Engineer, Pretraining Scaling - London

Anthropic
London
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication","Collaboration","Problem-solving","On-call","Time management"]

Join Anthropic's ML Performance and Scaling team as a Research Engineer overseeing production pretraining pipelines. You will optimize model operations, debug across the full stack, design experiments to boost training efficiency, and support on-call incidents during model launches, while collaborating with teams in SF and London to push scalable, reliable AI systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
11 months ago

Research Engineer, Pretraining Scaling - London

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Join Anthropic's ML Performance and Scaling team as a Research Engineer overseeing production pretraining pipelines. You will optimize model operations, debug across the full stack, design experiments to boost training efficiency, and support on-call incidents during model launches, while collaborating with teams in SF and London to push scalable, reliable AI systems.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Own critical aspects of our production pretraining pipeline, including model operations, performance optimization, observability, and reliability.
  • •Debug and resolve complex issues across the full stack—from hardware errors and networking to training dynamics and evaluation infrastructure.
  • •Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance.
  • •Respond to on-call incidents during model launches, diagnosing problems quickly and coordinating solutions across teams.
  • •Build and maintain production logging, monitoring dashboards, and evaluation infrastructure.

Key Requirements

  • •Hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems.
  • •Ability to balance research and engineering work (approximately 50/50) and thrive under production pressures during launches.
  • •Comfort with being on-call for production systems and solving hard problems under pressure, including extended hours when needed.
  • •Strong debugging skills across multiple layers of the stack and excellent communication for cross-time-zone collaboration.
  • •Passion for societal impacts of AI and responsible scaling.
Experience:AIMachine learningDistributed systems
Education:Bachelor's
Skills:CommunicationCollaborationProblem-solvingOn-callTime management
Languages:English
Tech Stack:JAXTPUPyTorchDistributed systemsObservabilityLoggingTraining pipelines

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn