Full-Stack Software Engineer, Reinforcement Learning

Anthropic
San Francisco, New York
Workplace: OnsiteFull timeUSD 300,000 - 405,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Communication","Problem-solving","Collaboration","Ownership","Ux","User empathy","Teamwork"]

We’re seeking a full-stack software engineer in reinforcement learning to build platforms, tools, and UIs that power environment creation, data collection, and training observability. Own end-to-end product surfaces—from backend services and APIs to web interfaces used by researchers, vendors, and data labelers—shipping polished, reliable features in a fast-moving RL environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
5 months ago

Full-Stack Software Engineer, Reinforcement Learning

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

We’re seeking a full-stack software engineer in reinforcement learning to build platforms, tools, and UIs that power environment creation, data collection, and training observability. Own end-to-end product surfaces—from backend services and APIs to web interfaces used by researchers, vendors, and data labelers—shipping polished, reliable features in a fast-moving RL environment.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build and extend web platforms for RL environment creation, management, and quality review — including environment configuration, versioning, and validation workflows.
  • •Develop vendor-facing interfaces and tooling that let external partners create, submit, and iterate on training environments with minimal friction.
  • •Design and implement platforms for human data collection at scale, including labeling workflows, quality assurance systems, and feedback mechanisms that surface reward signal integrity issues early.
  • •Build evaluation dashboards and observability UIs that give researchers real-time insight into environment quality, training run health, and reward hacking.
  • •Create backend services and APIs that connect environment authoring tools, data collection systems, and RL training infrastructure.

Pay and Benefits

Salary: USD 300,000 - 405,000 annually
Equity and Bonus:Equity
Perks:Paid LeaveParental LeaveEquity

Key Requirements

  • •Strong software engineering fundamentals with real full-stack range from database schema to frontend; able to own a surface end-to-end.
  • •Proficient in Python and a modern web stack (React, TypeScript, or similar).
  • •Track record of shipping systems that solved a hard problem and made the team faster.
  • •Operate with high agency: identify what needs to be done and drive it forward without waiting for a ticket.
  • •Care about UX and build interfaces intuitive for both technical researchers and non-technical labelers.
Experience:AIMachine learningReinforcement learning
Education:Bachelor's
Skills:CommunicationProblem-solvingCollaborationOwnershipUxUser empathyTeamwork
Languages:English
Tech Stack:PythonReactTypeScriptDockerCI/CDGCPAWSCloud

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn