Safeguards Enforcement Lead, Cyber Harms

Anthropic
Washington, San Francisco, New York
Workplace: OnsiteFull timeUSD 285,000 - 330,000 annuallyFunction: CybersecurityEducation: bachelorsSkills: ["Leadership","Stakeholder management","Communication","Collaboration","Risk identification"]

Lead and execute safeguards enforcement across products and services, focusing on detecting and mitigating attempts to misuse AI for malicious cyber operations. Develop enforcement frameworks for cyberattacks, malware creation, and offensive exploitation, and manage a team of Cyber Enforcement Analysts and contractors. Partner with Engineering, Data Science, and Policy teams to handle high-severity and ambiguous cases, improve tooling and measurement, and stay current on evolving AI policy enforcement and cyber threat tactics.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
2 days ago

Safeguards Enforcement Lead, Cyber Harms

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Lead and execute safeguards enforcement across products and services, focusing on detecting and mitigating attempts to misuse AI for malicious cyber operations. Develop enforcement frameworks for cyberattacks, malware creation, and offensive exploitation, and manage a team of Cyber Enforcement Analysts and contractors. Partner with Engineering, Data Science, and Policy teams to handle high-severity and ambiguous cases, improve tooling and measurement, and stay current on evolving AI policy enforcement and cyber threat tactics.
Location: Washington, San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Cybersecurity
Seniority: Manager level

Key Responsibilities

  • •Manage a team of Cyber Enforcement Analysts and contractors, overseeing the vision of Cyber Enforcement strategy.
  • •Create strategies to detect and mitigate AI misuse for cyberattacks, malware creation, exploitation tooling, and related harmful cyber operations.
  • •Collaborate with stakeholders on novel, ambiguous, or high-severity cases.
  • •Partner with Safeguards Policy Design, Engineering, and Data Science to address policy gaps and provide tooling and measurement support.
  • •Stay current on AI policy enforcement best practices, threat actor tactics, and the evolving cyber threat landscape.

Pay and Benefits

Salary: USD 285,000 - 330,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •Experience as a people manager.
  • •Experience in cybersecurity, including knowledge of offensive techniques, exploit development, malware analysis, or vulnerability research.
  • •Experience performing content review, abuse investigations, or policy enforcement at volume.
  • •Proficiency in SQL and/or Python for data analysis and threat detection.
  • •Experience identifying emerging risks and communicating findings to stakeholders across Product, Policy, Engineering, and Legal.
Experience:CybersecurityTrust & safetyAbuse investigationsThreat intelligenceGenerative AI
Education:Bachelor's
Skills:LeadershipStakeholder managementCommunicationCollaborationRisk identification
Languages:English
Tech Stack:SQLPythonGenerative AILarge language modelsLLMsContent moderation

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn