AI Content Red Team Analyst - Trust and Safety

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Content & Editorial (Writing/Editing)Experience: 3+ yearsSkills: ["Analytical rigor","Strong judgment","Creativity","Communication","Cross-functional collaboration"]

Conduct unstructured adversarial testing of generative AI models and product experiences to uncover emerging content-safety risks. Investigate jailbreaks, evasions, prompt-based attacks, and failure modes across contexts and user journeys. Document vulnerabilities with reproduction steps, severity, and mitigation recommendations, and partner with policy, product, engineering, data science, operations, and business teams to validate mitigations and drive root-cause closure. Stay current on evolving adversarial trends and contribute to testing playbooks.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

AI Content Red Team Analyst - Trust and Safety

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Conduct unstructured adversarial testing of generative AI models and product experiences to uncover emerging content-safety risks. Investigate jailbreaks, evasions, prompt-based attacks, and failure modes across contexts and user journeys. Document vulnerabilities with reproduction steps, severity, and mitigation recommendations, and partner with policy, product, engineering, data science, operations, and business teams to validate mitigations and drive root-cause closure. Stay current on evolving adversarial trends and contribute to testing playbooks.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Content & Editorial (Writing/Editing)
Seniority: Mid level

Key Responsibilities

  • •Conduct structured adversarial testing on AI models, features, and policies to identify vulnerabilities and emerging risks.
  • •Explore product behavior across contexts and user journeys to uncover model failure modes not covered by standard evaluations.
  • •Investigate jailbreaks, evasions, prompt-based attacks, and other content-safety adversarial techniques.
  • •Document findings clearly with risk descriptions, reproduction steps, severity assessments, and mitigation recommendations.
  • •Partner with cross-functional stakeholders to validate mitigations and support root cause closure, including developing testing playbooks and taxonomies.

Key Requirements

  • •3+ years in Trust & Safety, cybersecurity, risk/adversarial testing, or related fields.
  • •Experience with prompt testing, jailbreak analysis, LLM evaluation, or adversarial QA.
  • •Familiarity with AI safety risks including jailbreaks, hallucinations, bias, and misuse patterns.
  • •Strong interest in GenAI safety and how AI systems can be compromised under adversarial conditions.
  • •Ability to independently investigate ambiguous problems, identify non-obvious failure modes/abuse patterns, and produce evidence-based conclusions.
Experience:3+ yearsGenAILLMTrust and safetyAdversarial testingCybersecurity
Skills:Analytical rigorStrong judgmentCreativityCommunicationCross-functional collaboration

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn