Machine Learning Research Scientist, Evaluations

Scale AI
San Francisco, Seattle, New York
Workplace: OnsiteFull timeUSD 180,600 - 225,750 annuallyFunction: Data Science & Machine LearningEducation: phdSkills: ["Written communication","Verbal communication","Research publication","Problem diagnosis","Collaboration"]

Build and lead evaluation-driven research for frontier LLMs and agents, focusing on root-cause diagnosis of failure modes across text and multimodal settings. Develop and run rigorous benchmarks and evaluation methods, applying post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to data and training interventions. Publish findings at top AI conferences and partner with foundation model labs to inform next-generation model development.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
2 days ago

Machine Learning Research Scientist, Evaluations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Build and lead evaluation-driven research for frontier LLMs and agents, focusing on root-cause diagnosis of failure modes across text and multimodal settings. Develop and run rigorous benchmarks and evaluation methods, applying post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to data and training interventions. Publish findings at top AI conferences and partner with foundation model labs to inform next-generation model development.
Location: San Francisco, Seattle, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Develop rigorous evaluations and diagnostic methods to reveal where frontier models fail and why.
  • •Analyze model behavior to identify, characterize, and diagnose failure modes using RCA for LLMs and agents.
  • •Design and build benchmarks and evaluation methods for text and multimodal modalities.
  • •Apply post-training expertise (SFT, RLHF, reward modeling) to connect failures to data and training interventions.
  • •Publish research findings in top-tier AI conferences and partner with foundation model labs to inform next-generation model development.

Pay and Benefits

Salary: USD 180,600 - 225,750 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRetirement BenefitsLearning BudgetPaid LeaveCommuter Benefits

Key Requirements

  • •Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.
  • •Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
  • •Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.
  • •Excellent written and verbal communication skills.
  • •Published research in machine learning at major conferences and/or journals.
Experience:GenAILLMReinforcement learningMachine learning research
Education:PhD / Doctorate in Computer Science, Machine Learning, AI
Skills:Written communicationVerbal communicationResearch publicationProblem diagnosisCollaboration
Languages:English
Tech Stack:LLMLLMsPost-trainingSFTRLHFReinforcement learningReward modelingPreference modelingInstruction tuningBenchmark developmentMultimodalAgentsDeep learning

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn