Member of Technical Staff - Data Quality Engineer (Post-training)

Reflection AI
San Francisco, New York, London
Workplace: OnsiteFull timeFunction: QA, Test & Release EngineeringSkills: ["Communication","Detail-oriented","Problem-solving","Analytical","Teamwork"]

Join the Data Team to own upstream data quality for LLM post-training and evaluation, building automated QA methods, QA pipelines, and measurable quality signals. You’ll collaborate with researchers and engineers to translate requirements into concrete data quality standards, shaping model training and evaluation. This IC role emphasizes data quality impact on agentic use cases, long-horizon reasoning, and safety alignment in open foundational models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
8 months ago

Member of Technical Staff - Data Quality Engineer (Post-training)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Join the Data Team to own upstream data quality for LLM post-training and evaluation, building automated QA methods, QA pipelines, and measurable quality signals. You’ll collaborate with researchers and engineers to translate requirements into concrete data quality standards, shaping model training and evaluation. This IC role emphasizes data quality impact on agentic use cases, long-horizon reasoning, and safety alignment in open foundational models.
Location: San Francisco, New York, London
Workplace: Onsite
Employment Type: Full time
Job Function: QA, Test & Release Engineering

Key Responsibilities

  • •Own upstream data quality for LLM post-training and evaluation by analyzing expert-developed datasets and operationalizing quality standards for reasoning, alignment, and agentic use cases
  • •Partner closely with research and post-training teams to translate requirements into measurable quality signals, and provide actionable feedback to external data vendors
  • •Design, validate, and scale automated QA methods, including LLM-as-a-Judge frameworks, to reliably measure data quality across large campaigns
  • •Build reusable QA pipelines that reliably deliver high-quality data to post-training teams for model training and evaluation
  • •Monitor and report on data quality over time, driving continuous iteration on quality standards, processes, and acceptance criteria

Pay and Benefits

Perks:Health InsuranceDentalVisionParental LeaveRelocationMeal Allowance

Key Requirements

  • •Strong engineering fundamentals with experience building data pipelines, QA systems, or evaluation workflows for post-training data and agentic environments
  • •Detail-oriented with an analytical mindset, able to identify failure modes, inconsistencies, and subtle issues that affect data quality
  • •Solid understanding of how data quality impacts training (SFT and RL) and evaluation, with the ability to translate quality concerns into concrete signals, decisions, and feedback
  • •Experience designing and validating automated quality checks, including rule-based systems, statistical methods, or model-assisted approaches such as LLM-as-a-Judge
  • •Comfortable working autonomously, owning problems end-to-end, and collaborating effectively with researchers, engineers, and operations partners
Experience:AiMlData qualityPost-training
Skills:CommunicationDetail-orientedProblem-solvingAnalyticalTeamwork
Tech Stack:PythonMLLLMData pipelinesQAStatistical methods

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor