Member of Technical Staff - Engineering Lead, Data Ingestion

Reflection AI
San Francisco, London, New York
Workplace: OnsiteFull timeFunction: Administration & Executive AssistanceSkills: ["Technical leadership","Mentorship","Prioritization","Communication","Execution"]

Lead Reflection’s Data Ingestion team responsible for turning the web and other large-scale sources into structured, versioned, auditable training corpora. You’ll build, mentor, and grow data ingestion engineers; guide technical and architectural decisions across web crawl, extraction/normalization pipelines, and data lakes; and stay hands-on with the stack. Partner with research, data quality, and partnerships teams to connect ingestion choices to downstream model impact and run evaluation experiments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
1 month ago

Member of Technical Staff - Engineering Lead, Data Ingestion

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Lead Reflection’s Data Ingestion team responsible for turning the web and other large-scale sources into structured, versioned, auditable training corpora. You’ll build, mentor, and grow data ingestion engineers; guide technical and architectural decisions across web crawl, extraction/normalization pipelines, and data lakes; and stay hands-on with the stack. Partner with research, data quality, and partnerships teams to connect ingestion choices to downstream model impact and run evaluation experiments.
Location: San Francisco, London, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Manager level

Key Responsibilities

  • •Build, mentor, and grow a high-performing team of data ingestion engineers.
  • •Provide front-line leadership across the ingestion stack: web crawling/acquisition, extraction/normalization pipelines, and data lakes for versioned delivery.
  • •Stay hands-on and contribute as an individual contributor by maintaining deep stack familiarity.
  • •Manage day-to-day execution by prioritizing work and running data acquisition/ingestion campaigns in a fast-paced environment.
  • •Guide technical and architectural decisions for scalability, reliability, auditability, cost, and dataset versioning/delivery, partnering with research and data quality teams.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionLifeEquityPaid LeaveParental LeaveMeal AllowanceWellness Stipend

Key Requirements

  • •Experience building, mentoring, and growing data or infrastructure engineering teams while staying technically hands-on.
  • •Deep experience building web-scale data acquisition or ingestion systems with ownership of production pipelines at multi-TB to PB scale.
  • •Strong coding ability with credibility to earn technical trust.
  • •Expertise in at least one area: web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes/corpus storage & delivery, with ability to learn the rest.
  • •Fluency with large-scale data tooling including distributed compute (Ray/Beam/Spark), orchestration (Airflow/Prefect), and data formats (Parquet/JSONL/WARC).
Experience:Data infrastructureLarge-scale dataLLM dataWeb-scale dataData ingestion
Skills:Technical leadershipMentorshipPrioritizationCommunicationExecution
Tech Stack:RayBeamSparkAirflowPrefectParquetJSONLWARCObject storesLakehouseRobots.txtLLMs

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor