Staff Software Engineer, AI Reliability Engineering

Anthropic
London
Workplace: OnsiteFull timeGBP 325,000 - 390,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving"]

Join the AIRE Serving team to build reliable AI infrastructure for model serving and training at scale. You’ll help define Service Level Objectives, design and implement monitoring, enable high-availability serving across regions and clouds, and drive automated failover and incident response. Collaborate with ML and infrastructure teams to improve latency, availability, and cost efficiency for millions of external customers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
7 months ago

Staff Software Engineer, AI Reliability Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Join the AIRE Serving team to build reliable AI infrastructure for model serving and training at scale. You’ll help define Service Level Objectives, design and implement monitoring, enable high-availability serving across regions and clouds, and drive automated failover and incident response. Collaborate with ML and infrastructure teams to improve latency, availability, and cost efficiency for millions of external customers.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Develop appropriate Service Level Objectives for large language model serving and training systems, balancing availability/latency with development velocity.
  • •Design and implement monitoring systems including availability, latency and other salient metrics.
  • •Assist in the design and implementation of high-availability language model serving infrastructure capable of handling the needs of millions of external customers and high-traffic internal workloads.
  • •Develop and manage automated failover and recovery systems for model serving deployments across multiple regions and cloud providers.
  • •Lead incident response for critical AI services, ensuring rapid recovery and systematic improvements from each incident

Pay and Benefits

Salary: GBP 325,000 - 390,000 annually
Perks:Paid LeaveParental LeaveFlexible Hours

Key Requirements

  • •Bachelor's degree in a related field or equivalent experience
  • •Extensive experience with distributed systems observability and monitoring at scale
  • •Proven experience implementing and maintaining SLO/SLA frameworks for business-critical services
  • •Comfortable with both traditional metrics (latency, availability) and AI-specific metrics (model performance, training convergence)
  • •Experience with chaos engineering and systematic resilience testing
  • •Excellent communication skills and ability to bridge ML engineers and infrastructure teams
Experience:AI infrastructureSaaSCloud computing
Education:Bachelor's
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:GPUTPUTrainiumRDMAInfiniBandSLOMonitoringObservability

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn