Staff+ Software Engineer, Safeguards

Anthropic
San Francisco
Workplace: OnsiteFull timeUSD 320,000 - 425,000 annuallyFunction: Software EngineeringExperience: 5-10 yearsEducation: bachelorsSkills: ["Communication","Teamwork","Problem-solving","Risk assessment"]

Software engineer on the Safeguards team building safety and oversight mechanisms for AI systems. You will monitor models, detect unwanted behaviors, prevent disallowed use, surface abuse patterns to researchers, and implement multi-layered defenses at scale to uphold safety and transparency while enforcing terms of service.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
10 months ago

Staff+ Software Engineer, Safeguards

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Software engineer on the Safeguards team building safety and oversight mechanisms for AI systems. You will monitor models, detect unwanted behaviors, prevent disallowed use, surface abuse patterns to researchers, and implement multi-layered defenses at scale to uphold safety and transparency while enforcing terms of service.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Develop monitoring systems to detect unwanted behaviors from our API partners and potentially take automated enforcement actions; surface these in internal dashboards to analysts for manual review
  • •Build abuse detection mechanisms and infrastructure
  • •Surface abuse patterns to our research teams to harden models at the training stage
  • •Build robust and reliable multi-layered defenses for real-time improvement of safety mechanisms that work at scale

Pay and Benefits

Salary: USD 320,000 - 425,000 annually

Key Requirements

  • •Bachelor’s degree in Computer Science, Software Engineering or comparable experience
  • •5-10+ years of experience in a software engineering position, preferably with a focus on integrity, spam, fraud, or abuse detection and mitigation
  • •Proficiency in Python and Typescript
  • •Ability to work across the stack
  • •Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders
Experience:5-10 yearsAI safetyAbuse detectionSecurityScalability
Education:Bachelor's
Skills:CommunicationTeamworkProblem-solvingRisk assessment
Languages:English
Tech Stack:PythonTypescriptAPIDashboardInfra

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn