Member of Technical Staff - Web Crawl Engineer

Reflection AI
San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Ray","Spark","Beam","Flink","HTML parsing","Browser automation","Rendering","Distributed systems","URLs","Crawl scheduling"]

Build and operate large-scale web crawling infrastructure to continuously discover, acquire, and process content from billions of URLs. Collaborate with researchers and infrastructure teams to optimize URL discovery, scheduling, and crawl orchestration, while ensuring quality, coverage, and efficiency across diverse web formats.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
3 months ago

Member of Technical Staff - Web Crawl Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build and operate large-scale web crawling infrastructure to continuously discover, acquire, and process content from billions of URLs. Collaborate with researchers and infrastructure teams to optimize URL discovery, scheduling, and crawl orchestration, while ensuring quality, coverage, and efficiency across diverse web formats.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build and operate web-scale crawling infrastructure capable of continuously collecting data across billions of URLs.
  • •Design and optimize URL discovery, prioritization, scheduling, and crawl orchestration systems.
  • •Develop distributed crawlers that efficiently acquire content while respecting site constraints and operational requirements.
  • •Build systems for content extraction, rendering, parsing, and normalization across diverse web formats.
  • •Improve crawl coverage, freshness, efficiency, and quality through measurement and experimentation.

Key Requirements

  • •Experience building large-scale web crawling, search indexing, content acquisition, or internet-scale data collection systems.
  • •Strong understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination.
  • •Experience with large-scale distributed systems using technologies such as Ray, Spark, Beam, Flink, or similar frameworks.
  • •Familiarity with content extraction, HTML parsing, browser automation, rendering systems, and modern web technologies.
  • •Experience operating systems that process petabyte-scale datasets.
Experience:Web crawlingDistributed systemsData collection
Skills:RaySparkBeamFlinkHTML parsingBrowser automationRenderingDistributed systemsURLsCrawl scheduling
Languages:English
Tech Stack:RaySparkBeamFlinkHTMLBrowser automationRenderingDistributed systems

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor