Research Engineer (General)

hud
San Francisco, Singapore
Workplace: HybridFull timeUSD 140,000 - 250,000 annuallyFunction: Research & Scientific (R&D)Experience: 3+ yearsSkills: ["Attention to detail","Independent work","Communication"]

Build the technical foundation for training and evaluating frontier AI agents by creating and improving agent training environments and the full lifecycle of agent training data. Design experiments to diagnose model behavior, agent failure modes, and data quality issues, develop tools for higher-quality tasks and trajectories, and create metrics to validate whether HUD’s environments and evals improve RL training. Partner with external vendors to improve the data engine’s quality and throughput.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
hud
hud
4 months ago

Research Engineer (General)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Build the technical foundation for training and evaluating frontier AI agents by creating and improving agent training environments and the full lifecycle of agent training data. Design experiments to diagnose model behavior, agent failure modes, and data quality issues, develop tools for higher-quality tasks and trajectories, and create metrics to validate whether HUD’s environments and evals improve RL training. Partner with external vendors to improve the data engine’s quality and throughput.
Location: San Francisco, Singapore
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Build systems for creating, running, evaluating, and improving agent training environments.
  • •Design experiments to understand model behavior, agent failure modes, and data quality issues.
  • •Develop tools to help researchers, engineers, and data vendors create higher-quality tasks, trajectories, and feedback loops.
  • •Own the full lifecycle of agent training data, including task design, environment setup, trajectory collection, evaluation, and validation.
  • •Build metrics and analyses to determine whether tasks, environments, and evals are useful for training frontier agents.

Pay and Benefits

Salary: USD 140,000 - 250,000 annually
Perks:Health InsuranceDentalVision401kMeal AllowanceGym MembershipCommuter Benefits

Key Requirements

  • •Proficiency in Python, Docker, and Linux environments.
  • •Experience reasoning about benchmarks and evals, including what makes tasks realistic, rubrics reliable, environments usable, and trajectories useful for RL training.
  • •Strong attention to detail to spot subtle inconsistencies in data, model behavior, or task design.
  • •Experience building tools, pipelines, experiments, or infrastructure without a fully prescribed roadmap.
  • •Early-stage startup experience and the ability to work independently in fast-paced environments.
Experience:3+ yearsRL trainingAI researchBenchmarksEvaluationRL environments
Skills:Attention to detailIndependent workCommunication
Tech Stack:PythonDockerLinuxChatGPTClaude CodeCursor

Company Brief

hud
Hud provides a lightweight developer-focused tool that captures and shares UI components and interactive design specs from the browser to streamline design-to-development handoff and collaboration between designers and engineers.
Industry: Developer Tools
Website