Member of Technical Staff - Data Quality Engineer (Pre-training)

Reflection AI
San Francisco, New York, London
Workplace: OnsiteFull timeFunction: QA, Test & Release EngineeringSkills: ["Communication","Detail-oriented","Analytical thinking","Collaboration"]

Join the Data Team to own and improve data quality for pre-training of large language models. You’ll build QA pipelines, define measurable quality signals, collaborate with research and vendor teams, and scale automated checks to ensure high-quality training data across languages and modalities, directly influencing model performance and reliability in a fast-moving AI research environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
8 months ago

Member of Technical Staff - Data Quality Engineer (Pre-training)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Join the Data Team to own and improve data quality for pre-training of large language models. You’ll build QA pipelines, define measurable quality signals, collaborate with research and vendor teams, and scale automated checks to ensure high-quality training data across languages and modalities, directly influencing model performance and reliability in a fast-moving AI research environment.
Location: San Francisco, New York, London
Workplace: Onsite
Employment Type: Full time
Job Function: QA, Test & Release Engineering

Key Responsibilities

  • •Own upstream data quality for LLM pre-training; as a specialist or generalist across languages and modalities
  • •Partner closely with research and pre-training teams to translate requirements into measurable quality signals, and provide actionable feedback to external data vendors
  • •Design, validate, and scale automated QA methods to reliably measure data quality across large campaigns
  • •Build reusable QA pipelines that reliably deliver high-quality data to pre-training teams for model training
  • •Monitor and report on data quality over time, driving continuous iteration on quality standards, processes, and acceptance criteria

Pay and Benefits

Perks:Health InsuranceDentalVisionLife InsuranceDisability InsuranceParential LeaveRelocation

Key Requirements

  • •Proficiency in Python and building ML / LLM workflows. Must be comfortable debugging and writing scalable code
  • •Strong engineering fundamentals with experience building data pipelines, QA systems, or evaluation workflows for pre-training data
  • •Experience working with large datasets and automated evaluation or quality-checking systems
  • •Familiarity with how LLMs work and can describe how models are trained and evaluated
  • •Excellent communication skills with the ability to clearly articulate complex technical concepts across teams
Experience:AIMLPre-trainingLLMData quality
Skills:CommunicationDetail-orientedAnalytical thinkingCollaboration
Tech Stack:PythonMLLLMData pipelinesQA systems

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor