Evaluation Researcher

Aaru
New York
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Skills: ["Intellectual honesty","High ownership","Clear communication","Independent judgment","Analytical rigor"]

Own difficult evaluation and measurement problems at the intersection of machine learning, statistics, behavioral science, and product decision-making. Define constructs and decision criteria, assemble evaluation data, design historical and prospective studies, build evaluation code and diagnostics, and quantify uncertainty. Collaborate closely with research and engineering teams while protecting independence and holdouts to ensure evidence is valid, reproducible, and actionable for real decisions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Aaru
Aaru
3 days ago

Evaluation Researcher

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Own difficult evaluation and measurement problems at the intersection of machine learning, statistics, behavioral science, and product decision-making. Define constructs and decision criteria, assemble evaluation data, design historical and prospective studies, build evaluation code and diagnostics, and quantify uncertainty. Collaborate closely with research and engineering teams while protecting independence and holdouts to ensure evidence is valid, reproducible, and actionable for real decisions.
Location: New York
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Own an evaluation or measurement area across population construction, predictive systems, individual agent behavior, group dynamics, or end-to-end simulations.
  • •Turn realism, accuracy, calibration, usefulness, and decision quality questions into measurable constructs and explicit decision criteria.
  • •Design studies and evaluations using methods like historical backtests, temporal holdouts, prospective outcomes, controlled experiments, and mixed methods as needed.
  • •Build diagnostic tests for individual and population quality, quantify uncertainty, and inspect failures hidden by aggregate metrics.
  • •Write clear technical reports and develop reusable evaluation infrastructure (datasets, harnesses, libraries, graders, schemas, leaderboards, and reporting tools) while protecting holdouts from leakage.

Key Requirements

  • •A record of rigorous work in ML evaluation, statistics, behavioral science, computational social science, economics, psychometrics, or experimental design (or equivalent research depth).
  • •Ability to define a difficult construct precisely enough to measure it without losing the underlying question.
  • •Comfort with observational data, sampling, statistical power, uncertainty, causal threats, selection, leakage, and condition shift.
  • •Capability to write code, analyze large datasets, and inspect individual examples beyond aggregate dashboards.
  • •Clear technical communication of evidence, uncertainty, limitations, and alternative interpretations, including negative or inconclusive findings.
Experience:Machine learning evaluationStatisticsBehavioral scienceComputational social scienceEconometrics
Skills:Intellectual honestyHigh ownershipClear communicationIndependent judgmentAnalytical rigor

Company Brief

Aaru
Aaru builds multi-agent simulation software that models entire populations to predict behavior and future events, replacing traditional research with decision-ready forecasts for enterprises, governments, and agencies across industries.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Revenue: USD 1M to 5M
Growth: Early Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn