Red Team Engineer, Safeguards

Anthropic
San Francisco
Workplace: RemoteFull timeUSD 320,000 - 405,000 annuallyFunction: Administration & Executive AssistanceEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving","Written communication","Verbal communication"]

Take an adversarial approach to test Anthropic’s deployed AI systems and products, uncovering vulnerabilities across the product ecosystem before malicious actors can exploit them. You’ll research novel abuse patterns unique to advanced AI, design full kill-chain attack scenarios, and build automated testing frameworks for continuous, at-scale evaluation. Work closely with Product, Engineering, and Policy teams to translate findings into concrete safety improvements and measure detection effectiveness.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
1 month ago

Red Team Engineer, Safeguards

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Take an adversarial approach to test Anthropic’s deployed AI systems and products, uncovering vulnerabilities across the product ecosystem before malicious actors can exploit them. You’ll research novel abuse patterns unique to advanced AI, design full kill-chain attack scenarios, and build automated testing frameworks for continuous, at-scale evaluation. Work closely with Product, Engineering, and Policy teams to translate findings into concrete safety improvements and measure detection effectiveness.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: Administration & Executive Assistance

Key Responsibilities

  • •Conduct comprehensive adversarial testing across Anthropic product surfaces, creating multi-technique attack scenarios.
  • •Research and implement novel testing approaches for emerging AI capabilities such as agent systems and tool use.
  • •Design and execute full kill-chain attacks that emulate real-world threat actors and malicious objectives.
  • •Build and maintain systematic testing methodologies and automate testing frameworks for continuous assessment at scale.
  • •Collaborate with Product, Engineering, and Policy teams to turn findings into concrete improvements and help define detection metrics.
Travel: Medium travel

Pay and Benefits

Salary: USD 320,000 - 405,000 annually
Perks:Parental LeavePaid LeaveEquity

Key Requirements

  • •Experience in penetration testing, red teaming, or application security.
  • •Experience in model jailbreaking and testing large-scale agentic workflows for prompt injection vulnerabilities.
  • •Hands-on web application security skills, including tools such as Burp Suite, Metasploit, and custom testing frameworks.
  • •Experience building custom automation and LLM-specific testing frameworks.
  • •A strong track record (e.g., CVEs, blog posts, or bug bounty reports) and the ability to communicate technical concepts clearly.
Experience:AI safetyAdversarial MLTrust & safetyAbuse prevention
Education:Bachelor's
Skills:CommunicationCollaborationProblem-solvingWritten communicationVerbal communication
Languages:English
Tech Stack:Burp SuiteMetasploitLLMAgent systemsTool usePrompt injectionAPI securityRate-limiting

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn