Staff Software Engineer, AI Reliability Engineering

Anthropic
Dublin
Workplace: OnsiteFull timeEUR 235,000 - 295,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Communication","Problem-solving","Collaboration"]

Join the AI Reliability Engineering team to elevate the reliability of large language model serving and training systems. You’ll implement SLOs, build monitoring and high-availability infrastructure, automate failover across regions and clouds, lead incident response, and optimize cost for large-scale AI infrastructure, enabling reliable service for millions of external customers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
7 months ago

Staff Software Engineer, AI Reliability Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join the AI Reliability Engineering team to elevate the reliability of large language model serving and training systems. You’ll implement SLOs, build monitoring and high-availability infrastructure, automate failover across regions and clouds, lead incident response, and optimize cost for large-scale AI infrastructure, enabling reliable service for millions of external customers.
Location: Dublin
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Develop Service Level Objectives for large language model serving and training systems, balancing availability/latency with development velocity.
  • •Design and implement monitoring systems including availability, latency and other salient metrics.
  • •Assist in the design and implementation of high-availability language model serving infrastructure capable of handling the needs of millions of external customers and high-traffic internal workloads.
  • •Develop and manage automated failover and recovery systems for model serving deployments across multiple regions and cloud providers.
  • •Lead incident response for critical AI services, ensuring rapid recovery and systematic improvements from each incident

Pay and Benefits

Salary: EUR 235,000 - 295,000 annually

Key Requirements

  • •Extensive experience with distributed systems observability and monitoring at scale.
  • •Experience operating AI infrastructure including model serving, batch inference, and training pipelines.
Experience:AI infrastructureDistributed systemsCloud computing
Education:Bachelor's
Skills:CommunicationProblem-solvingCollaboration
Languages:English
Tech Stack:GPUTPUTrainiumRDMAInfiniBandObservabilitySLOSLA

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn