Staff Software Engineer, AI Reliability

Anthropic
San Francisco, New York, Seattle
Workplace: HybridFull timeFunction: Software EngineeringEducation: bachelorsSkills: ["Communication","Collaboration","Ownership","Problem-solving","Teamwork"]

Join the AI Reliability Engineering (AIRE) group to improve reliability across large language model serving paths—from SDK to network, API layers, and cloud infrastructure. You’ll define service level objectives, design monitoring, lead incident response, and help ensure scalable, multi-region serving with strong observability and safety commitments. A cross-functional role requiring collaboration across teams to keep Claude reliable.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
6 months ago

Staff Software Engineer, AI Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 48 minutes agoStatus: Live

Job Summary

Join the AI Reliability Engineering (AIRE) group to improve reliability across large language model serving paths—from SDK to network, API layers, and cloud infrastructure. You’ll define service level objectives, design monitoring, lead incident response, and help ensure scalable, multi-region serving with strong observability and safety commitments. A cross-functional role requiring collaboration across teams to keep Claude reliable.
Location: San Francisco, New York, Seattle
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Develop and monitor Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
  • •Design and implement monitoring and observability across the token path and serving infrastructure.
  • •Assist in creating high-availability serving infrastructure across multiple regions and cloud providers.
  • •Lead incident response for critical AI services, ensuring rapid recovery and thorough post-incident reviews.
  • •Support reliability of safeguard model serving to meet safety commitments.

Pay and Benefits

Perks:Paid LeaveParental LeaveFlexible HoursEquity

Key Requirements

  • •Strong distributed systems, infrastructure, or reliability backgrounds with a software engineering or SRE focus.
  • •Experience operating large-scale model serving or training infrastructure (more than 1000 GPUs).
  • •Experience with ML hardware accelerators (GPUs, TPUs, Trainium).
  • •Knowledge of ML-specific networking optimizations like RDMA and InfiniBand.
  • •Expertise in AI-specific observability tools and frameworks and incident response practices.
Experience:AIML infrastructureDistributed systems
Education:Bachelor's
Skills:CommunicationCollaborationOwnershipProblem-solvingTeamwork
Languages:English
Tech Stack:GPUsTPUsRDMAInfiniBandObservabilityModel servingDistributed systems

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn