Senior Research Scientist, Model Evaluation

Cohere
Toronto, New York, Seattle, San Francisco, London, United States, Canada
Workplace: RemoteFull timeFunction: Data Science & Machine LearningSkills: ["Communication","Problem-solving","Collaboration","Rigor","Software engineering"]

Senior Research Scientist, Model Evaluation leads the creation of advanced evaluation benchmarks and scalable tools to measure LLM progress. You’ll work across cross-functional teams to translate model feedback into repeatable evaluations, conduct cutting-edge research in LLM evaluation methods, and build infrastructure to assessAI capabilities at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
10 months ago

Senior Research Scientist, Model Evaluation

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Senior Research Scientist, Model Evaluation leads the creation of advanced evaluation benchmarks and scalable tools to measure LLM progress. You’ll work across cross-functional teams to translate model feedback into repeatable evaluations, conduct cutting-edge research in LLM evaluation methods, and build infrastructure to assessAI capabilities at scale.
Location: Toronto, New York, Seattle, San Francisco, London, United States, Canada
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish.
  • •Work on cross-functional teams to translate model feedback into trustworthy, repeatable evaluations.
  • •Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges and refining data synthesis pipelines.
  • •Improve evaluation efficiency and build scalable, reusable tools for analyzing model performance.
  • •Prototype and validate resources to measure AI capabilities and ensure alignment with desired outcomes.

Pay and Benefits

Perks:Health InsuranceDentalParentral LeaveRemote WorkMeal AllowancePaid Leave

Key Requirements

  • •Strong software engineering skills with a track record of building prototypes and production-grade tooling
  • •Extensive experience analyzing complex data and LLM outputs to ensure high data quality
  • •Experience designing and evaluating AI systems with rigorous measurement of capabilities
  • •Ability to work on highly cross-functional teams to translate model feedback into reliable evaluations
  • •Familiarity with evaluation benchmarks, data synthesis pipelines, and scalable evaluation infrastructure
Experience:AIMLNLPResearch
Skills:CommunicationProblem-solvingCollaborationRigorSoftware engineering
Languages:English
Tech Stack:PythonLLMData synthesisBenchmarksEvaluation tooling

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor