Research, Safety
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication"]Work on AI safety research to make models safe and trustworthy, exploring how training and data shape refusal and engagement with harmful or dual-use requests. Contribute across the stack—from pre-training data curation and safety-focused fine-tuning to evaluations and red-teaming—designing experiments, building safety evaluations (including long-horizon/agentic tasks), and developing mitigations based on failure modes you discover.
Loading
Loading job details...
Preparing the role view and application actions.

