Member of Technical Staff - Multilingual Data

Reflection AI
San Francisco, London, New York
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Ownership","Measurement mindset","Collaboration","Problem-solving","Curiosity"]

Design and operate large-scale multilingual data pipelines, including sourcing, cleaning, deduplication, language identification, and script normalization across high- and low-resource languages. Define quality bars for multilingual corpora (translation quality, cultural fidelity, toxicity, contamination) and run experiments to improve multilingual data efficiency. Build evaluation sets and diagnostics to pinpoint degradation by language and domain, and collaborate across pre-, mid-, and post-training teams to deliver measurable improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
16 hours ago

Member of Technical Staff - Multilingual Data

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Design and operate large-scale multilingual data pipelines, including sourcing, cleaning, deduplication, language identification, and script normalization across high- and low-resource languages. Define quality bars for multilingual corpora (translation quality, cultural fidelity, toxicity, contamination) and run experiments to improve multilingual data efficiency. Build evaluation sets and diagnostics to pinpoint degradation by language and domain, and collaborate across pre-, mid-, and post-training teams to deliver measurable improvements.
Location: San Francisco, London, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and operate large-scale multilingual data pipelines across high- and low-resource languages.
  • •Define and enforce quality bars for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks.
  • •Design and run scientific experiments to improve multilingual data efficiency for large language models.
  • •Lead small research projects independently while collaborating on larger initiatives.
  • •Build evaluation sets and diagnostics to identify multilingual degradation and close gaps with targeted data.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionLife InsuranceEquityPaid LeaveParental LeaveVisa Sponsorship

Key Requirements

  • •Strong software engineering fundamentals and experience processing web-scale datasets in distributed environments.
  • •Experience building large-scale data pipelines for language models, machine translation, speech, or search, ideally across multiple languages.
  • •Fluency or working proficiency in at least one language other than English and curiosity about linguistic differences.
  • •A rigorous, measurement-first mindset to validate that data changes improve outcomes.
  • •Comfort balancing research objectives with practical engineering trade-offs in an ambiguous, high-ownership environment.
Skills:OwnershipMeasurement mindsetCollaborationProblem-solvingCuriosity

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor