Safeguards Enforcement Analyst, Cyber Harm

Anthropic
San Francisco, New York, Washington
Workplace: HybridFull timeUSD 285,000 - 330,000 annuallyFunction: CybersecurityEducation: bachelorsSkills: ["Communication","Stakeholder collaboration","Risk identification"]

Review flagged content and accounts to make accurate enforcement decisions focused on detecting and mitigating attempts to misuse AI systems for malicious cyber operations. Triage ambiguous and high-severity cases, escalate when needed, and provide detailed feedback to the Safeguards policy team. Partner with Engineering and Data Science by surfacing detection model errors and quality signals to improve precision and recall, while staying current on evolving cyber threats and AI enforcement best practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
1 month ago

Safeguards Enforcement Analyst, Cyber Harm

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Review flagged content and accounts to make accurate enforcement decisions focused on detecting and mitigating attempts to misuse AI systems for malicious cyber operations. Triage ambiguous and high-severity cases, escalate when needed, and provide detailed feedback to the Safeguards policy team. Partner with Engineering and Data Science by surfacing detection model errors and quality signals to improve precision and recall, while staying current on evolving cyber threats and AI enforcement best practices.
Location: San Francisco, New York, Washington
Workplace: Hybrid
Employment Type: Full time
Job Function: Cybersecurity

Key Responsibilities

  • •Review flagged content and accounts to make accurate, well-documented enforcement decisions aligned with usage policies.
  • •Detect and mitigate potential misuse of AI systems that facilitates cyberattacks, malware creation, exploitation tooling, and other harmful cyber operations.
  • •Triage and escalate novel, ambiguous, or high-severity cases to appropriate stakeholders.
  • •Provide detailed feedback to the Safeguards policy design team on policy gaps surfaced through real enforcement scenarios.
  • •Partner with Engineering and Data Science to surface detection model errors and quality signals from review to improve precision and recall.

Pay and Benefits

Salary: USD 285,000 - 330,000 annually
Equity and Bonus:Equity
Perks:EquityPaid LeaveParental Leave

Key Requirements

  • •Experience in cybersecurity, including knowledge of offensive techniques, exploit development, malware analysis, or vulnerability research.
  • •Experience performing content review, abuse investigations, or policy enforcement at volume.
  • •Proficiency in SQL and/or Python for data analysis and threat detection.
  • •Experience identifying emerging risks and communicating findings to stakeholders such as Product, Policy, Engineering, and Legal teams.
  • •Experience working with generative AI products, including writing effective prompts for content review and enforcement.
Education:Bachelor's
Skills:CommunicationStakeholder collaborationRisk identification
Languages:English
Tech Stack:SQLPythonGenerative AILarge language models

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn