Research Engineer, Safety
San Francisco, New York
Workplace: OnsiteFull timeUSD 200,000 - 400,000 annuallyFunction: Research & Scientific (R&D)Experience: 2+ yearsSkills: ["Experimental judgment","Risk tradeoffs","Ownership","Handling ambiguity"]Build and evaluate safety mechanisms for conversational AI agents end-to-end, from spotting real-world failure modes to shipping safeguards in production. You’ll design adversarial evaluations and red-team datasets, develop classifiers/judges and post-training or runtime safeguards, and analyze production traces to prevent prompt injection, unsafe tool use, sensitive-data disclosure, policy violations, and hallucinated commitments. Work cross-functionally with Security, Product, Infrastructure, Legal, and enterprise stakeholders.
Loading
Loading job details...
Preparing the role view and application actions.

