Safeguards Enforcement Analyst, Integrity & Authenticity

Anthropic
San Francisco, New York, Washington
Workplace: HybridFull timeUSD 285,000 - 330,000 annuallyFunction: CybersecurityEducation: bachelorsSkills: ["Communication","Stakeholder management","Collaboration"]

Build and execute enforcement workflows to detect and mitigate coordinated inauthentic behavior, election manipulation, and targeting/surveillance harms. Design automated enforcement systems and review processes that scale with high accuracy, partnering with Engineering and Data Science to improve detection models. Review flagged content, support safeguards policy design with real enforcement feedback, and stay current on evolving AI policy enforcement practices and threat tactics.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
1 month ago

Safeguards Enforcement Analyst, Integrity & Authenticity

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and execute enforcement workflows to detect and mitigate coordinated inauthentic behavior, election manipulation, and targeting/surveillance harms. Design automated enforcement systems and review processes that scale with high accuracy, partnering with Engineering and Data Science to improve detection models. Review flagged content, support safeguards policy design with real enforcement feedback, and stay current on evolving AI policy enforcement practices and threat tactics.
Location: San Francisco, New York, Washington
Workplace: Hybrid
Employment Type: Full time
Job Function: Cybersecurity
Seniority: Mid level

Key Responsibilities

  • •Design and architect automated enforcement systems and scalable review workflows while maintaining high accuracy.
  • •Partner with Engineering and Data Science to optimize detection models for policy violations and enforcement systems.
  • •Review flagged content to drive enforcement and policy improvements.
  • •Enforce usage policies focused on misuse including AI-enabled influence operations, coordinated inauthentic behavior, election interference, and targeting/surveillance harms.
  • •Provide detailed feedback to safeguards policy design based on real enforcement scenarios and emerging threat tactics and regulatory developments.

Pay and Benefits

Salary: USD 285,000 - 330,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •Experience in trust & safety, policy enforcement, threat intelligence, or a closely related field focused on influence operations, disinformation, election integrity, or privacy/surveillance harms.
  • •Experience standing up and scaling policy enforcement or content review workflows.
  • •Proficiency in SQL and/or other data analysis tools to draw insights from large datasets.
  • •Experience identifying emerging risks and threat actors and communicating findings to stakeholders across Product, Policy, Engineering, and Legal teams.
  • •Experience working with generative AI products, including writing effective prompts for content review and enforcement.
Experience:Trust & safetyPolicy enforcementThreat intelligenceElectionsPrivacySurveillanceDisinformationInfluence operations
Education:Bachelor's in A field relevant to the role as demonstrated through coursework, training, or professional experience
Skills:CommunicationStakeholder managementCollaboration
Languages:English
Tech Stack:SQLPythonLarge language modelsLLMsOSINTNetwork analysis

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn