Staff Software Engineer, Code RL

Anthropic
San Francisco, New York, Seattle
Workplace: HybridFull timeUSD 405,000 - 625,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Communication","Mentorship","Collaboration","Systems thinking","Problem-solving"]

Build and scale reinforcement learning infrastructure behind Claude’s coding capabilities. You’ll design widely-used Python APIs, frameworks, and abstractions, embed with rotating research teams to transfer ownership, and work inside research codebases to improve reliability and structure. The role also covers the health of production RL runs—monitoring, regression detection, and triage tooling—while helping define engineering standards and mentor engineers adopting them.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
1 month ago

Staff Software Engineer, Code RL

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build and scale reinforcement learning infrastructure behind Claude’s coding capabilities. You’ll design widely-used Python APIs, frameworks, and abstractions, embed with rotating research teams to transfer ownership, and work inside research codebases to improve reliability and structure. The role also covers the health of production RL runs—monitoring, regression detection, and triage tooling—while helping define engineering standards and mentor engineers adopting them.
Location: San Francisco, New York, Seattle
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design widely-used APIs, frameworks, and abstractions with attention to interface legibility and principled defaults.
  • •Embed with research teams on a rotational basis to understand needs, build systems and APIs, and transfer ownership so teams maintain them.
  • •Work in research codebases to improve reliability and structure without slowing research.
  • •Prevent silent failure modes structurally using type safety, invariants, targeted testing, and refactors.
  • •Contribute to production RL reliability through monitoring, regression detection, triage tooling, and engineering standards/mentorship.

Pay and Benefits

Salary: USD 405,000 - 625,000 annually
Perks:Parental Leave

Key Requirements

  • •Deep expertise in Python, including static typing, safe async and concurrency patterns, and writing performant Python code.
  • •Track record designing intuitive, safe APIs or frameworks that other engineers or teams adopted and built on.
  • •Experience working productively in large, evolving, or research-style codebases not originally written by you.
  • •Demonstrated ability to anticipate silent failure modes and prevent them through system design, type safety, and testing.
  • •Strong written and verbal communication skills to explain system designs to collaborators with varied engineering backgrounds.
Experience:Machine learning researchReinforcement learningAgentic systemsLLM training pipelinesDistributed systemsData processing
Education:Bachelor's in A field relevant to the role as demonstrated through coursework, training, or professional experience
Skills:CommunicationMentorshipCollaborationSystems thinkingProblem-solving
Languages:English
Tech Stack:PythonStatic typingAsyncConcurrencyType safety

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn