Member of Technical Staff - Data Ingestion Engineer

Reflection AI
San Francisco, New York, London
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Communication","Collaboration","Problem-solving","Experimentation","Analytical thinking"]

The role focuses on designing, building, and operating large-scale data ingestion systems to turn open web and other sources into structured corpora for pre-training frontier models. You will own ingestion machinery, run experiments on crawling and extraction strategies, and collaborate with researchers, infra, and operations to improve data quality and model performance across distributed systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
7 months ago

Member of Technical Staff - Data Ingestion Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

The role focuses on designing, building, and operating large-scale data ingestion systems to turn open web and other sources into structured corpora for pre-training frontier models. You will own ingestion machinery, run experiments on crawling and extraction strategies, and collaborate with researchers, infra, and operations to improve data quality and model performance across distributed systems.
Location: San Francisco, New York, London
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build and operate large-scale data ingestion systems for pre-training, including web crawling, extraction, and dataset delivery.
  • •Run experiments to evaluate crawling strategies, extraction methods, and ingestion tradeoffs
  • •Analyze ingested data to identify gaps, redundancy, and areas to improve
  • •Build ingestion pipelines that scale reliably across large data campaigns
  • •Develop specialized crawlers for high-priority data sources

Pay and Benefits

Perks:Health InsuranceDentalVisionLife InsuranceParantal LeaveRelocation

Key Requirements

  • •Experience building web crawling, data ingestion, or large-scale data acquisition systems using Ray, Beam, Spark, or similar technologies.
  • •Familiarity with how LLMs are trained and evaluated, and an intuition for what makes data useful for training
  • •Comfortable working with very large datasets (multi-TB to PB scale) and building systems that are observable, testable, and maintainable
  • •Comfortable designing experiments and using data to guide system improvements
  • •Excellent communication skills. You can explain system behavior and communicate tradeoffs clearly
Experience:AIData engineeringMachine learningOpen dataResearch
Skills:CommunicationCollaborationProblem-solvingExperimentationAnalytical thinking
Tech Stack:RayBeamSparkWeb crawlingData ingestionDistributed systems

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor