Staff+ Site Reliability Engineer, Safeguards ML Infra

Anthropic
New York, San Francisco, Seattle
Workplace: OnsiteFull timeUSD 405,000 - 485,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Production judgment","Automation mindset","Incident response","Config management","Communication"]

Own production change management for safety systems powering Claude. Stand up and verify safeguards for every model launch, lead off-cycle deployments of safety classifiers, and ensure safeguards are correctly configured across platforms (1P, AWS Bedrock, GCP Vertex) by detecting drift. Turn launch runbooks and checks into automated, repeatable pipelines, maintain a safeguards registry with provenance, and participate in on-call rotations for time-sensitive safety incidents and releases.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
22 hours ago

Staff+ Site Reliability Engineer, Safeguards ML Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Own production change management for safety systems powering Claude. Stand up and verify safeguards for every model launch, lead off-cycle deployments of safety classifiers, and ensure safeguards are correctly configured across platforms (1P, AWS Bedrock, GCP Vertex) by detecting drift. Turn launch runbooks and checks into automated, repeatable pipelines, maintain a safeguards registry with provenance, and participate in on-call rotations for time-sensitive safety incidents and releases.
Location: New York, San Francisco, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Launch captain for model releases: stand up, configure, verify safeguards, and be the safeguards point of contact during release windows.
  • •Own off-cycle deployment of new safety classifiers from research via canarying, post-deploy validations, and discrepancy investigation.
  • •Verify safeguards are provably live on the right models across deployment platforms and detect/eliminate configuration drift.
  • •Automate launch runbooks and hand-built checks into continuous validation and repeatable deployment pipelines.
  • •Maintain a safeguards registry with full provenance and participate in on-call/operational-duty rotations for incidents and time-sensitive launches.

Pay and Benefits

Salary: USD 405,000 - 485,000 annually
Perks:Paid LeaveParental LeaveEquity

Key Requirements

  • •Have owned production change management at scale, including deploy pipelines and config management systems, with strong views on what “verified” means.
  • •Have run high-stakes releases as a launch captain/incident commander/release owner where bad deploys have real consequences.
  • •Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements, and can turn action items into tooling.
  • •Have hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale.
  • •Be proficient in Python (Rust experience is a plus).
Experience:Site reliability engineeringCloudMLProduction operationsOn-call
Education:Bachelor's
Skills:Production judgmentAutomation mindsetIncident responseConfig managementCommunication
Languages:English
Tech Stack:PythonRustAWSGCPAWS BedrockGCP VertexClaudeTransformerLLM inference

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn