Red Team Engineer, Safeguards
Anthropic
San Francisco
Workplace: RemoteFull timeUSD 320,000 - 405,000 annuallyFunction: Administration & Executive AssistanceEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving","Written communication","Verbal communication"]Take an adversarial approach to test Anthropic’s deployed AI systems and products, uncovering vulnerabilities across the product ecosystem before malicious actors can exploit them. You’ll research novel abuse patterns unique to advanced AI, design full kill-chain attack scenarios, and build automated testing frameworks for continuous, at-scale evaluation. Work closely with Product, Engineering, and Policy teams to translate findings into concrete safety improvements and measure detection effectiveness.

