Research Engineer, Code RL (Reinforcement Learning)

Anthropic
San Francisco, New York
Workplace: OnsiteFull-timeUSD 500,000 - 850,000 annuallyFunction: Education & TrainingEducation: bachelorsSkills: ["Communication","Problem-solving","Collaboration","Attention to detail","Adaptability"]

Hybrid of research and software engineering focused on building end-to-end RL code for real-world coding tasks. Design RL environments, implement reward signals, run large-scale training on frontier models, diagnose performance, and optimize pipelines to accelerate iteration while ensuring safety and quality.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
3 months ago

Research Engineer, Code RL (Reinforcement Learning)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Hybrid of research and software engineering focused on building end-to-end RL code for real-world coding tasks. Design RL environments, implement reward signals, run large-scale training on frontier models, diagnose performance, and optimize pipelines to accelerate iteration while ensuring safety and quality.
Location: San Francisco, New York
Workplace: Onsite
Job Function: Education & Training

Key Responsibilities

  • •Design RL environments and coding tasks, build reward signals and verifiers that capture what good code means.
  • •Run training experiments on frontier models and diagnose why a model improves or fails on a class of software-engineering tasks.
  • •Improve the speed and reliability of pipelines that iterate research ideas into deployed models.
  • •Balance research exploration with engineering implementation and engage in rigorous experimental design and interpretation of results.
  • •Collaborate with teams across alignment, red teams, and production training to ensure safe and scalable systems.

Pay and Benefits

Salary: USD 500,000 - 850,000 annually
Perks:Paid LeaveParental LeaveFlexible Hours

Key Requirements

  • •Have strong software-engineering skills and deep Python expertise, including async/concurrent programming.
  • •Are comfortable owning systems end to end and debugging across the stack.
  • •Can balance research exploration with engineering implementation, and engage rigorously in shaping experimental design and interpreting results.
  • •Care about code quality, testing, and performance.
  • •Are passionate about the potential impact of AI and are committed to developing safe and beneficial systems.
Experience:AIReinforcement learningML engineering
Education:Bachelor's
Skills:CommunicationProblem-solvingCollaborationAttention to detailAdaptability
Languages:English
Tech Stack:PythonPyTorchDistributed trainingCUDAGPUTPUAsyncConcurrent programmingAnacondaCode sandboxesEngineering

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn