AI Content Red Team Analyst - Trust and Safety
San Jose
Workplace: OnsiteFull timeFunction: Content & Editorial (Writing/Editing)Experience: 3+ yearsSkills: ["Judgment","Creativity","Analytical rigor","Evidence-based decision making","Cross-functional collaboration"]Conduct structured adversarial testing of generative AI models and features to uncover emerging trust and safety risks. Probe behavior across contexts and user journeys, investigate jailbreaks, evasions, and prompt-based attacks, and document evidence with reproduction steps and severity assessments. Partner with policy, product, business, and engineering stakeholders to validate mitigations and drive root-cause closure, while maintaining testing playbooks and staying current on evolving abuse trends.

