Research Scientist, Interpretability
San Francisco
Workplace: OnsiteFull timeUSD 350,000 - 850,000 annuallyFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving","Writing"]Join Anthropic's Interpretability team to reverse engineer how trained models work and develop mechanistic explanations for neural networks. You will design robust experiments, build infrastructure for running experiments and visualizing results, create interpretability features, and communicate findings with colleagues and the broader community. Familiarity with Python and a strong research background in ML/NLP are preferred.
Loading
Loading job details...
Preparing the role view and application actions.

