Staff+ Software Engineer, ML Inference Path

Anthropic
San Francisco
Workplace: OnsiteFull timeUSD 320,000 - 485,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Reliability focus","Results-oriented","Collaboration","Communication","Bias to impact"]

Build and operate production ML infrastructure for Claude’s safety systems on the token-generation path. Partner with safety researchers to transfer new classifier and defense techniques into reliable, scalable deployments. Own monitoring, observability, automated testing, and deployment/rollback, optimizing inference latency and throughput for real-time safety evaluations across platforms like 1P, Bedrock, and Vertex.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
5 days ago

Staff+ Software Engineer, ML Inference Path

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and operate production ML infrastructure for Claude’s safety systems on the token-generation path. Partner with safety researchers to transfer new classifier and defense techniques into reliable, scalable deployments. Own monitoring, observability, automated testing, and deployment/rollback, optimizing inference latency and throughput for real-time safety evaluations across platforms like 1P, Bedrock, and Vertex.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and build scalable ML infrastructure for real-time safety deployments across classifier and model ecosystems.
  • •Develop monitoring and observability tools to track classifier performance, data quality, and system health.
  • •Collaborate with research teams to productionize safety research, translating experimental techniques into robust systems.
  • •Optimize inference latency and throughput while maintaining high reliability standards for real-time evaluations.
  • •Implement automated testing, deployment, and rollback systems for ML models in production safety applications.

Pay and Benefits

Salary: USD 320,000 - 485,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •Proficient in Python with experience using ML frameworks such as PyTorch, TensorFlow, or JAX.
  • •Experience designing distributed systems handling high-throughput, low-latency workloads.
  • •Built automated/self-service deployment pipelines and evaluation infrastructure for rolling out classifiers and models independently.
  • •Implemented A/B testing and experimentation infrastructure for ML systems.
  • •Results-oriented approach with strong focus on reliability and impact in safety-critical applications.
Experience:5+ yearsAI safetyMachine learning
Education:Bachelor's
Skills:Reliability focusResults-orientedCollaborationCommunicationBias to impact
Languages:English
Tech Stack:PythonPyTorchTensorFlowJAXDistributed systemsA/B testingTransformer architecturesBedrockVertex

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn