Researcher, Evals

Cartesia
California, San Francisco
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Skills: ["Creativity","Communication","Problem-solving","Teamwork","Analytical thinking"]

Evaluations Lead will design evaluation frameworks and pipelines for next-generation interactive AI models, measuring not just knowledge but how models reason, remember, and interact over time. You’ll bridge research, product, and infrastructure to create robust metrics, studies, and systems that define intelligence in real-world use, influencing model development and evaluation at Cartesia.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cartesia
Cartesia
10 months ago

Researcher, Evals

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Evaluations Lead will design evaluation frameworks and pipelines for next-generation interactive AI models, measuring not just knowledge but how models reason, remember, and interact over time. You’ll bridge research, product, and infrastructure to create robust metrics, studies, and systems that define intelligence in real-world use, influencing model development and evaluation at Cartesia.
Location: California, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Identify and define key model capabilities and behaviors that matter for next-generation model evals
  • •Develop and implement new evaluation pipelines with robust statistical analysis and clear reporting
  • •Partner closely with model training and research teams to embed evaluation systems directly into model development loops
  • •Prototype new user studies and behavioral experiments to ground evaluations in real-world use
  • •Define and refine quantitative metrics that capture subjective or behavioral qualities of models

Pay and Benefits

Perks:Health InsuranceDentalVision401kRelocationImmigration Support

Key Requirements

  • •Experience designing or implementing evaluation frameworks for generative models (audio, text, or multimodal)
  • •Strong technical and analytical skills, ability to take open-ended research ideas and translate them into production-ready systems
  • •Creativity in defining novel quantitative metrics for subjective or behavioral qualities
  • •Excitement for building evaluation systems that bridge research and real-world use
  • •Curiosity and rigor in equal measure; motivation driven by discovering how to measure meaningful progress in intelligent behavior
Experience:AIResearchMachine learningMultimodal
Skills:CreativityCommunicationProblem-solvingTeamworkAnalytical thinking
Languages:English

Company Brief

Cartesia
Builds real-time multimodal and voice AI (Sonic) that generates expressive, low-latency speech for conversational agents and on-device experiences, serving developers and enterprises with APIs and SDKs.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2023
WebsiteLinkedInGlassdoor