Web Crawling - Research Engineer

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Research & Scientific (R&D)Experience: 8+ yearsSkills: ["Ownership","Reliability focus","Technical direction","Cross-team collaboration","Efficiency improvement"]

Own Thinking Machines’ web-crawling and ingestion stack, from distributed collection to filtering, deduplication, and decisioning on stored data. Design and scale production crawlers and infrastructure at internet scale and petabyte scale, building pipelines that convert raw crawls into high-quality pretraining data. Partner with pretraining and data teams to understand how crawled-data changes impact model quality, while improving reliability, efficiency, and technical direction as the area grows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 day ago

Web Crawling - Research Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own Thinking Machines’ web-crawling and ingestion stack, from distributed collection to filtering, deduplication, and decisioning on stored data. Design and scale production crawlers and infrastructure at internet scale and petabyte scale, building pipelines that convert raw crawls into high-quality pretraining data. Partner with pretraining and data teams to understand how crawled-data changes impact model quality, while improving reliability, efficiency, and technical direction as the area grows.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design and scale the web crawler and ingestion infrastructure that sources pretraining data.
  • •Build pipelines for large-scale extraction, deduplication, and data quality filtering.
  • •Build specialized crawlers for high-value or hard-to-reach data sources.
  • •Work with pretraining teams to understand how changes in crawled data affect model performance.
  • •Improve reliability and efficiency of crawling and ingestion infrastructure at petabyte scale.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeavePaid ParentalRelocation

Key Requirements

  • •8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems.
  • •Demonstrated ownership of crawler or data-acquisition infrastructure at internet scale.
  • •Strong software engineering skills in Python, Go, or Rust, with real experience in distributed systems.
  • •Working knowledge of practical and legal considerations for web data collection (robots.txt, rate limiting, licensing).
  • •Track record of owning production crawling and ingestion systems and delivering usable pretraining data pipelines.
Experience:8+ yearsWeb crawlingData acquisitionDistributed systemsInternet scalePetabyte-scale data
Skills:OwnershipReliability focusTechnical directionCross-team collaborationEfficiency improvement
Tech Stack:PythonGoRustDistributed systemsRobots.txtRate limitingWeb crawlingScrapingDeduplicationData pipelinesPetabyte-scale storageData quality filteringMachine learning

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website