Evaluation Research Manager

Aaru
New York
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Skills: ["Leadership","Communication","Independent judgment","Prioritization","Truth-seeking"]

Lead a focused team of evaluation researchers and research engineers to turn the evaluation charter into a portfolio of rigorous, decision-relevant studies. You’ll design and run analyses, build reusable evaluation infrastructure, set standards for baselines and uncertainty, and diagnose failures that aggregate metrics miss. Work across simulation, population, prediction, engineering, deployment, and leadership to ensure improvements translate into better real-world decision outcomes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Aaru
Aaru
3 days ago

Evaluation Research Manager

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Lead a focused team of evaluation researchers and research engineers to turn the evaluation charter into a portfolio of rigorous, decision-relevant studies. You’ll design and run analyses, build reusable evaluation infrastructure, set standards for baselines and uncertainty, and diagnose failures that aggregate metrics miss. Work across simulation, population, prediction, engineering, deployment, and leadership to ensure improvements translate into better real-world decision outcomes.
Location: New York
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Manager level

Key Responsibilities

  • •Build, lead, and develop a high-performing team of evaluation researchers and research engineers, setting priorities and recruiting exceptional talent.
  • •Translate evaluation questions into measurable constructs, decisive experiments, and explicit decision criteria that support real claims and useful diagnostics.
  • •Set standards for evaluation infrastructure and rigor, including baselines, temporal holdouts, prospective testing, contamination control, uncertainty, subgroup analysis, and reproducibility.
  • •Design and run evaluations that test population quality, forecast calibration, ranking quality, selective prediction, temporal validity, and real cost of errors, then find failures hidden by aggregates.
  • •Partner with engineering teams to make evaluations repeatable and integrated into development and release workflows, and communicate negative or inconclusive results with precision and urgency.

Key Requirements

  • •Lead evaluation, measurement, or empirical research in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or similarly rigorous environments.
  • •Define difficult constructs precisely enough to measure them without losing the underlying question, and choose effective evidence to support claims.
  • •Comfortable with experimental design and observational work including sampling, statistical power, uncertainty, causal threats, leakage, and condition shift.
  • •Ability to write code, analyze large datasets, inspect individual failures, and review technical work of researchers and engineers.
  • •Experience managing or technically leading strong researchers, providing clear feedback and independent research judgment.
Experience:Machine learning evaluationBehavioral scienceComputational social scienceStatisticsEconometrics
Skills:LeadershipCommunicationIndependent judgmentPrioritizationTruth-seeking
Tech Stack:Machine learningEmpirical researchStatistical powerUncertainty quantificationCausal inferenceExperimental designCalibrationProper scoring rulesForecastingBacktestingLLM agentsMulti-agent systemsSynthetic populationsProbabilistic modelsEconometricsPsychometrics

Company Brief

Aaru
Aaru builds multi-agent simulation software that models entire populations to predict behavior and future events, replacing traditional research with decision-ready forecasts for enterprises, governments, and agencies across industries.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Revenue: USD 1M to 5M
Growth: Early Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn