Staff+ Site Reliability Engineer, Safeguards ML Infra

Anthropic
San Francisco, Seattle, New York
Workplace: OnsiteFull timeUSD 320,000 - 485,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Production judgment","Incident response","Automation","Communication"]

Own production change management for Claude’s safety safeguards by launching and deploying safeguards for every new model release. Canaries off-cycle safety classifier deployments, verifies safeguards are correctly live across 1P, AWS Bedrock, and GCP Vertex, and prevents configuration drift. Turn manual launch runbooks and checks into continuous validation pipelines, maintain a safeguards registry with provenance, and participate in on-call and incident response for time-sensitive safety launches.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
1 day ago

Staff+ Site Reliability Engineer, Safeguards ML Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own production change management for Claude’s safety safeguards by launching and deploying safeguards for every new model release. Canaries off-cycle safety classifier deployments, verifies safeguards are correctly live across 1P, AWS Bedrock, and GCP Vertex, and prevents configuration drift. Turn manual launch runbooks and checks into continuous validation pipelines, maintain a safeguards registry with provenance, and participate in on-call and incident response for time-sensitive safety launches.
Location: San Francisco, Seattle, New York
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Launch captain for model releases: stand up, configure, and verify safeguards and act as the safeguards point of contact during release windows.
  • •Own off-cycle deployment of new safety classifiers, including canary rollouts, post-deploy validations, and discrepancy investigations.
  • •Verify safeguards are provably live on the correct models across every deployment platform and detect/eliminate configuration drift.
  • •Automate launch runbooks and manual checks into continuous validation and repeatable deployment pipelines.
  • •Maintain a safeguards registry with full provenance and participate in on-call/operational-duty rotations for service incidents and safety launches.
Travel: Low travel

Pay and Benefits

Salary: USD 320,000 - 485,000 annually
Perks:Parental LeavePaid Leave

Key Requirements

  • •Own production change management at scale, including deploy pipelines, config management systems, and canary analysis.
  • •Have run high-stakes releases (e.g., launch captain, incident commander, or release owner) where bad deploys have real consequences.
  • •Have meaningful on-call experience with incident response and postmortem-driven improvements, including turning action items into tooling.
  • •Have hands-on experience deploying and operating cloud platforms at scale (AWS, GCP).
  • •Be proficient in Python (Rust experience is a plus).
Experience:Site reliability engineeringProduction change management
Education:Bachelor's
Skills:Production judgmentIncident responseAutomationCommunication
Tech Stack:PythonRustAWSGCPBedrockGCP Vertex1PLLM inferenceTransformer architectures

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn