Member of Technical Staff - Evaluations

Reflection AI
San Francisco, New York, London
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Collaboration","Attention to detail","Analytical thinking","Communication","Teamwork"]

Join a fast-paced AI research startup to design and drive evaluation frameworks for open foundational models. You will conduct rigorous statistical analyses, build evaluation systems linking data to model behavior, and collaborate with pre-training, post-training, and applied teams to translate insights into model improvements, advancing capabilities through measurable benchmarks and human feedback.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
8 months ago

Member of Technical Staff - Evaluations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Join a fast-paced AI research startup to design and drive evaluation frameworks for open foundational models. You will conduct rigorous statistical analyses, build evaluation systems linking data to model behavior, and collaborate with pre-training, post-training, and applied teams to translate insights into model improvements, advancing capabilities through measurable benchmarks and human feedback.
Location: San Francisco, New York, London
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Conduct critical comparative analysis to advance our understanding of model capabilities.
  • •Build and refine evaluation systems and processes that create tight feedback loops between data, evals, and model behavior.
  • •Develop generalizable evaluation frameworks that capture what matters for reasoning, alignment, and usefulness.
  • •Collaborate closely with pre-training, post-training, and applied teams to translate insights into model improvements.
  • •Push the boundaries of what’s measurable, from synthetic evals to human feedback and real-world interaction data.

Pay and Benefits

Perks:Health InsuranceDentalVisionLife InsuranceDisability InsuranceRelocation SupportParental LeaveEquity

Key Requirements

  • •Strong statistical analysis and experimental design skills to rigorously measure model improvements.
  • •Familiarity with LLM evaluation methodologies: static benchmarks, human preference evals, and/or agentic tasks.
  • •High agency and thrive in a fast-paced startup environment; bias for impact over process.
  • •Excited to work in a new frontier lab, defining how we measure and accelerate progress toward more capable models.
  • •Collaborative, detail-oriented, and motivated by building the feedback loops that make models truly improve.
Experience:AIMLOpen models
Skills:CollaborationAttention to detailAnalytical thinkingCommunicationTeamwork

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor