Safeguards Enforcement Analyst, User Well-being

Anthropic
San Francisco, New York, Washington
Workplace: HybridFull timeUSD 245,000 - 285,000 annuallyFunction: Solutions Engineering & Sales EngineeringEducation: bachelorsSkills: ["Communication","Sound judgment","Cross-functional collaboration","Escalation judgment"]

Support the User Well-being team by designing and executing mental-health “guardrails” that improve how content is detected, reviewed, and responded to. Define metrics and evaluation datasets, partner with Engineering and Data Science to tune detection models, and monitor performance over time. Review flagged content, build in-product crisis-resource connections with Product and Legal, and keep safeguards policy grounded in real scenarios and emerging AI mental-health research.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
2 days ago

Safeguards Enforcement Analyst, User Well-being

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Support the User Well-being team by designing and executing mental-health “guardrails” that improve how content is detected, reviewed, and responded to. Define metrics and evaluation datasets, partner with Engineering and Data Science to tune detection models, and monitor performance over time. Review flagged content, build in-product crisis-resource connections with Product and Legal, and keep safeguards policy grounded in real scenarios and emerging AI mental-health research.
Location: San Francisco, New York, Washington
Workplace: Hybrid
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and execute interventions by defining key metrics and curating evaluation datasets.
  • •Partner with Engineering and Data Science to build, tune, and validate detection models, including threshold setting and precision/recall tradeoffs.
  • •Monitor how interventions and detection systems perform over time.
  • •Review flagged content to drive enforcement and policy improvements.
  • •Develop in-product features that connect users to crisis resources, working with Product, Legal, and external partners on referral pathways.

Pay and Benefits

Salary: USD 245,000 - 285,000 annually
Perks:EquityPaid LeaveParental LeaveFlexible Hours

Key Requirements

  • •Experience in trust and safety, product policy, content moderation, or related work with exposure to mental health, suicide, or self-harm harms.
  • •Experience designing or running experiments, evaluations, or measurement studies to determine whether an intervention worked.
  • •Translate policy definitions into measurable rubrics, review guidelines, or classification criteria for human or automated review.
  • •Manage or coordinate content review operations, including quality assurance and workflow management.
  • •Proficiency in SQL and/or other data analysis tools to measure intervention efficacy and monitor workflow health.
Experience:Trust & safetyContent moderationMental healthGenerative AI
Education:Bachelor's
Skills:CommunicationSound judgmentCross-functional collaborationEscalation judgment
Languages:English
Tech Stack:SQLData analysisGenerative AILLM-based classificationClaude CodeClaude

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn