Research Scientist, Interpretability

Anthropic
San Francisco
Workplace: OnsiteFull timeUSD 350,000 - 850,000 annuallyFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving","Writing"]

Join Anthropic's Interpretability team to reverse engineer how trained models work and develop mechanistic explanations for neural networks. You will design robust experiments, build infrastructure for running experiments and visualizing results, create interpretability features, and communicate findings with colleagues and the broader community. Familiarity with Python and a strong research background in ML/NLP are preferred.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
10 months ago

Research Scientist, Interpretability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Join Anthropic's Interpretability team to reverse engineer how trained models work and develop mechanistic explanations for neural networks. You will design robust experiments, build infrastructure for running experiments and visualizing results, create interpretability features, and communicate findings with colleagues and the broader community. Familiarity with Python and a strong research background in ML/NLP are preferred.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Develop methods for understanding LLMs by reverse engineering algorithms learned in their weights.
  • •Design and run robust experiments, both quickly in toy scenarios and at scale in large models.
  • •Create and analyze new interpretability features and circuits to better understand how models work.
  • •Build infrastructure for running experiments and visualizing results.
  • •Work with colleagues to communicate results internally and publicly.

Pay and Benefits

Salary: USD 350,000 - 850,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •Familiarity with Python is required for this role.
  • •Strong track record of scientific research (in any field), and some work on Interpretability.
  • •Ability to design and run robust experiments, both in toy scenarios and at scale.
  • •Ability to build infrastructure for running experiments and visualizing results.
  • •Excellent communication skills to discuss motivations and results, including writing up null results.
Experience:AI researchML interpretabilityNLP
Education:Bachelor's
Skills:CommunicationCollaborationProblem-solvingWriting
Languages:English
Tech Stack:Python

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn