Researcher, Alignment Interpretability

OpenAI
San Francisco
Workplace: OnsiteFull timeUSD 295,000 - 500,000 annuallyFunction: Research & Scientific (R&D)Experience: 2+ yearsEducation: phdSkills: ["Collaboration","Quantitative reasoning","Curiosity","Long-term thinking","Research process focus"]

Join the Interpretability team to develop and publish research on techniques for understanding deep neural network representations. You’ll also engineer infrastructure to study model internals at scale, collaborating across teams to pursue projects uniquely suited to OpenAI. The work directly supports OpenAI’s mission by helping ensure future models remain safe as they grow in capability, leveraging mechanistic interpretability and quantitative research rigor.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
3 days ago

Researcher, Alignment Interpretability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join the Interpretability team to develop and publish research on techniques for understanding deep neural network representations. You’ll also engineer infrastructure to study model internals at scale, collaborating across teams to pursue projects uniquely suited to OpenAI. The work directly supports OpenAI’s mission by helping ensure future models remain safe as they grow in capability, leveraging mechanistic interpretability and quantitative research rigor.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Develop and publish research techniques for understanding representations of deep networks.
  • •Engineer infrastructure for studying model internals at scale.
  • •Collaborate across teams on interpretability and alignment projects.
  • •Guide research directions toward demonstrable usefulness and long-term scalability.

Pay and Benefits

Salary: USD 295,000 - 500,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Experience conducting research on deep networks or related technical research areas.
  • •A strong background in mechanistic interpretability and/or AI safety & alignment.
  • •Proficiency in Python (or similar languages) and experience with research engineering.
  • •A Ph.D. or research experience in computer science, machine learning, or a related field.
  • •Experience thriving in large-scale AI systems environments and a curiosity-driven approach.
Experience:2+ yearsAI safety & alignmentMechanistic interpretabilityDeep learningResearch engineering
Education:PhD / Doctorate
Skills:CollaborationQuantitative reasoningCuriosityLong-term thinkingResearch process focus
Tech Stack:Python

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor