Researcher, Alignment CoT Monitorability

OpenAI
San Francisco
Workplace: HybridFull timeUSD 250,000 - 445,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Communication","Problem-solving","Teamwork","Independence","Curiosity"]

Join OpenAI’s CoT Monitorability team to design and run empirical studies that measure and improve how chain-of-thought monitorability is affected by training interventions. You will build evaluations, analyze model behavior, and translate findings into practical oversight strategies for frontier LLMs, collaborating with researchers and engineers across training and alignment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
2 months ago

Researcher, Alignment CoT Monitorability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Join OpenAI’s CoT Monitorability team to design and run empirical studies that measure and improve how chain-of-thought monitorability is affected by training interventions. You will build evaluations, analyze model behavior, and translate findings into practical oversight strategies for frontier LLMs, collaborating with researchers and engineers across training and alignment.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Design and run empirical studies of chain-of-thought monitorability across frontier reasoning models and training settings.
  • •Build evaluations that measure whether monitors can reliably predict properties of interest, including high-stakes forms of misbehavior.
  • •Investigate how pre-training, synthetic data, mid-training, post-training, reinforcement learning, and other interventions improve or degrade monitorability.
  • •Analyze model behavior and turn observations from monitoring into hypotheses, experiments, and recommendations.
  • •Translate research findings into practical monitoring and oversight approaches that can inform real training runs.

Pay and Benefits

Salary: USD 250,000 - 445,000 annually
Equity and Bonus:Equity
Perks:EquityRelocation

Key Requirements

  • •Strong, proven experience in empirical machine learning with a track record of hands-on work training, evaluating, or debugging large ML models (LLMs preferred).
  • •Deep interest in model behavior, alignment, or interpretability with ability to design rigorous experiments.
  • •Ability to move from ambiguous research questions to concrete experimental setups, including hypothesizing, building evaluations, and analyzing results.
  • •Solid collaboration skills, able to work with researchers and engineers across model training, alignment evaluations, monitoring, and frontier-risk work.
  • •Excellent Analytical thinking and autonomy to produce externally publishable research when results advance the science of alignment.
Experience:AI researchMachine learningAlignmentInterpretability
Skills:CommunicationProblem-solvingTeamworkIndependenceCuriosity
Languages:English
Tech Stack:MLPythonResearchLLMsEmpricial MLEvaluation

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor