Research Engineer - Web Crawlers

ElevenLabs
United Kingdom, London, New York, San Francisco, Warsaw, United States, Poland, Bulgaria
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Skills: ["Autonomy","Problem-solving","Quality evaluation"]

Own large-scale web crawling systems that source high-quality open-web data for frontier AI models. Build and operate distributed crawlers to reliably discover, fetch, and extract data at web scale, handling messy HTML, deduplication, freshness/recrawl strategies, and politeness/rate limiting. Design targeted crawling pipelines for audio, video, and multilingual content, and create tooling so researchers can request, monitor, and explore newly crawled datasets quickly and safely.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ElevenLabs
ElevenLabs
4 days ago

Research Engineer - Web Crawlers

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Own large-scale web crawling systems that source high-quality open-web data for frontier AI models. Build and operate distributed crawlers to reliably discover, fetch, and extract data at web scale, handling messy HTML, deduplication, freshness/recrawl strategies, and politeness/rate limiting. Design targeted crawling pipelines for audio, video, and multilingual content, and create tooling so researchers can request, monitor, and explore newly crawled datasets quickly and safely.
Location: United Kingdom, London, New York, San Francisco, Warsaw, United States, Poland, Bulgaria
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Build and operate large-scale, distributed web crawlers to discover, fetch, and extract data reliably and efficiently.
  • •Solve web crawling challenges including content extraction from messy HTML, deduplication at web scale, freshness/recrawl strategies, and politeness/rate-limit handling.
  • •Design targeted crawling pipelines to find high-value sources (audio, video, multilingual content) and produce clean training-ready datasets.
  • •Create tooling and infrastructure enabling researchers to request, monitor, and explore newly crawled web data quickly and reliably.

Pay and Benefits

Perks:Learning BudgetTravel AllowanceCo-working Stipend

Key Requirements

  • •Hands-on experience building and scaling web crawlers or scraping systems for machine learning training data.
  • •Strong engineering skills in distributed systems at scale, such as Kubernetes, queue-based architectures, or custom pipelines processing billions of documents.
  • •Ability to autonomously evaluate the quality, coverage, and compliance of crawled data and build tooling to measure it.
  • •Showcase past projects, designs, or GitHub contributions demonstrating solving hard technical problems.
  • •No formal certifications or degrees required; enthusiasm and engineering craft are key.
Experience:Machine learningAIData engineeringWeb crawlingDistributed systems
Skills:AutonomyProblem-solvingQuality evaluation
Tech Stack:KubernetesQueue-based architecturesGitHub

Company Brief

ElevenLabs
Develops advanced AI audio models and tools for realistic text-to-speech, voice cloning, dubbing, music generation, and conversational voice agents for creators and enterprises.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: London, United Kingdom
Founded: 2022
Glassdoor
Glassdoor: 4.2
WebsiteLinkedInGlassdoor