Safeguards Enforcement Analyst, Violence & Extremism

Anthropic
San Francisco, New York, Washington
Workplace: HybridFull timeUSD 285,000 - 330,000 annuallyFunction: OtherEducation: bachelorsSkills: ["Communication"]

Build and execute operational workflows to assess model behavior for violence and extremism policy areas. Design scalable enforcement systems, develop and maintain evals to measure performance and surface regressions, and partner with Engineering and Data Science to optimize detection. Review flagged content and refine enforcement guidelines, reviewer documentation, and policy feedback based on real enforcement scenarios. Continuously track emerging threats and misuse patterns to improve workflows and escalations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
2 months ago

Safeguards Enforcement Analyst, Violence & Extremism

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Build and execute operational workflows to assess model behavior for violence and extremism policy areas. Design scalable enforcement systems, develop and maintain evals to measure performance and surface regressions, and partner with Engineering and Data Science to optimize detection. Review flagged content and refine enforcement guidelines, reviewer documentation, and policy feedback based on real enforcement scenarios. Continuously track emerging threats and misuse patterns to improve workflows and escalations.
Location: San Francisco, New York, Washington
Workplace: Hybrid
Employment Type: Full time

Key Responsibilities

  • •Design and architect automated enforcement systems and review workflows that scale with high accuracy.
  • •Develop and maintain evals to measure model performance, surface regressions, and inform policy and model improvements.
  • •Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations.
  • •Review flagged content to drive enforcement decisions, surface policy gaps, and focus on novel, technically sophisticated misuse and emerging extremist movements.
  • •Develop and maintain enforcement guidelines and reviewer documentation, and keep workflows and evals updated with emerging threats, regulatory changes, and best practices.

Pay and Benefits

Salary: USD 285,000 - 330,000 annually
Perks:Parental Leave

Key Requirements

  • •Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field with exposure to harmful content or violent extremism.
  • •Experience standing up and scaling policy enforcement or content review workflows.
  • •Proficiency in SQL and/or other data analysis tools to derive insights from large datasets and monitor workflow health.
  • •Experience identifying emerging risks and threat actors and communicating findings to stakeholders across Product, Policy, Engineering, and Legal.
  • •Experience working with generative AI products, including writing effective prompts for content review and enforcement.
Experience:Threat intelligenceCounterterrorismContent moderationPolicy enforcementGenerative AI
Education:Bachelor's
Skills:Communication
Languages:English
Tech Stack:SQLPythonMITRE ATT&CKOSINT

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn