Research Engineer, Interpretability

Anthropic
San Francisco
Workplace: OnsiteFull timeUSD 315,000 - 560,000 annuallyFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Collaboration","Communication","Prioritization","Problem-solving"]

Join the Interpretability team to design and analyze experiments on mechanistic interpretability of large language models, build scalable research workflows and tooling, and help other teams apply interpretability techniques to improve model safety. You’ll work with researchers and engineers on large-scale AI systems, contributing to cutting-edge projects and reporoducing results across production models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
9 months ago

Research Engineer, Interpretability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Join the Interpretability team to design and analyze experiments on mechanistic interpretability of large language models, build scalable research workflows and tooling, and help other teams apply interpretability techniques to improve model safety. You’ll work with researchers and engineers on large-scale AI systems, contributing to cutting-edge projects and reporoducing results across production models.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Implement and analyze research experiments, both quickly in toy scenarios and at scale in large models.
  • •Set up and optimize research workflows to run efficiently and reliably at large scale.
  • •Build tools and abstractions to support rapid pace of research experimentation.
  • •Develop and improve tools and infrastructure to support other teams in using Interpretability’s work to improve model safety

Pay and Benefits

Salary: USD 315,000 - 560,000 annually

Key Requirements

  • •Have 5-10+ years of experience building software.
  • •Are highly proficient in at least one programming language (e.g., Python, Rust, Go, Java) and productive with python
  • •Have some experience contributing to empirical AI research projects
  • •Have a strong ability to prioritize and direct effort toward the most impactful work and are comfortable operating with ambiguity and questioning assumptions
  • •Prefer fast-moving collaborative projects to extensive solo efforts
Experience:Ai researchMachine learningTransformers
Education:Bachelor's
Skills:CollaborationCommunicationPrioritizationProblem-solving
Languages:English
Tech Stack:PythonRustGoJavaPyTorchGPUsDistributed systems

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn