Research Engineer, Benchmarks

hud
San Francisco, Singapore
Workplace: OnsiteFull timeUSD 140,000 - 250,000 annuallyFunction: Research & Scientific (R&D)Experience: 3+ yearsSkills: ["Detail-oriented","Communication","Independent work","Curiosity","Reasoning"]

Build and own rigorous benchmark suites to evaluate frontier AI agents on realistic, domain-specific tasks. Work with subject-matter experts to design tasks and scoring, implement benchmark infrastructure to run models reliably, and develop metrics that characterize difficulty, reliability, and failure modes. Validate benchmark relevance against real-world evals and customer expectations, and communicate results through clear technical documentation and benchmark reports.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
hud
hud
1 month ago

Research Engineer, Benchmarks

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Build and own rigorous benchmark suites to evaluate frontier AI agents on realistic, domain-specific tasks. Work with subject-matter experts to design tasks and scoring, implement benchmark infrastructure to run models reliably, and develop metrics that characterize difficulty, reliability, and failure modes. Validate benchmark relevance against real-world evals and customer expectations, and communicate results through clear technical documentation and benchmark reports.
Location: San Francisco, Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design, implement, and ensure the quality of HUD’s internal agent benchmarks.
  • •Partner with subject-matter experts to define tasks and create domain-specific benchmarks for realistic workflows.
  • •Build infrastructure to reliably run models and agents against benchmark tasks.
  • •Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes.
  • •Validate whether benchmark performance correlates with real-world evals and document results for technical audiences.

Pay and Benefits

Salary: USD 140,000 - 250,000 annually
Perks:Health InsuranceDentalVision401kMeal AllowancePaid LeaveCommuter BenefitsEquity

Key Requirements

  • •Proficiency in Python, Docker, and Linux environments.
  • •Experience working on environments and evals for models/agents.
  • •Strong understanding of what makes a benchmark realistic, reliable, and useful.
  • •Published papers or written technical blogs related to benchmarks and model failure modes (link in application).
  • •Curiosity and ability to understand how workflows in various domains work.
Experience:3+ years
Skills:Detail-orientedCommunicationIndependent workCuriosityReasoning
Tech Stack:PythonDockerLinuxChatGPTClaude CodeCursor

Eligibility

Work Authorization:Sponsorship available.

Company Brief

hud
Hud provides a lightweight developer-focused tool that captures and shares UI components and interactive design specs from the browser to streamline design-to-development handoff and collaboration between designers and engineers.
Industry: Developer Tools
Website