Member of Technical Staff, Data Infrastructure

Inception Labs
San Mateo
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsSkills: ["Python","SQL","NoSQL","Data processing","Privacy","ETL","Data ingestion","Data catalogs","Web scraping","Crawling"]

Experienced data infrastructure engineer to architect and scale core infrastructure for distributed training pipelines and petabyte-scale data catalogs. Collaborate with AI researchers to accelerate experiments, build high-throughput data ingestion and processing systems, and ensure privacy-compliant data collection across distributed storage and compute resources.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Inception Labs
Inception Labs
6 months ago

Member of Technical Staff, Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Experienced data infrastructure engineer to architect and scale core infrastructure for distributed training pipelines and petabyte-scale data catalogs. Collaborate with AI researchers to accelerate experiments, build high-throughput data ingestion and processing systems, and ensure privacy-compliant data collection across distributed storage and compute resources.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design, build, and operate scalable, fault-tolerant infrastructure for LLM research: distributed compute, data orchestration, and storage across modalities.
  • •Develop high-throughput systems for data ingestion, processing, and transformation — including training data catalogs, deduplication, quality checks, and search.
  • •Build systems for web crawling, data ingestion, and real-time data processing to support model training operations.
  • •Develop tools and frameworks for efficient data storage, retrieval, and versioning across distributed systems.
  • •Ensure data collection adheres to privacy regulations.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionEquityCommuter BenefitsPaid Leave

Key Requirements

  • •BS/MS/PhD in Computer Science, Machine Learning, or a related field (or equivalent experience).
  • •3+ years of experience building data processing pipelines at scale, particularly with AI/ML applications.
  • •Strong proficiency in Python and experience with data processing frameworks (Apache Spark, Beam, Airflow).
  • •Familiarity with synthetic data generation techniques and data augmentation strategies.
  • •Familiarity with web scraping, crawling technologies, and Common Crawl datasets.
Experience:3+ yearsAI/MLData processingDistributed systems
Skills:PythonSQLNoSQLData processingPrivacyETLData ingestionData catalogsWeb scrapingCrawling
Tech Stack:PythonApache SparkBeamAirflowPyTorchTensorFlowSQLNoSQLHDFSS3BigQueryCommon Crawl

Company Brief

Inception Labs
Develops artificial intelligence solutions and research-driven products, focusing on machine learning models, AI tools, and enterprise AI integrations to help organizations automate workflows, extract insights from data, and build intelligent applications.
Industry: AI & Machine Learning
Website