Software Engineer, Pretraining

Cursor
San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Ownership","Debugging","Experimentation","Partnership","High throughput systems"]

Build the data systems that power frontier model pretraining. You’ll work across Data Quality, Data Platform, and Crawling to process large-scale web crawls, construct training-ready datasets, and run experiments that validate data quality at extreme throughput. Own high-throughput, observable pipelines with end-to-end traceability, improve crawl success and parsing quality, and partner with researchers to close the loop between data changes and model loss and evals.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cursor
Cursor
2 days ago

Software Engineer, Pretraining

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Build the data systems that power frontier model pretraining. You’ll work across Data Quality, Data Platform, and Crawling to process large-scale web crawls, construct training-ready datasets, and run experiments that validate data quality at extreme throughput. Own high-throughput, observable pipelines with end-to-end traceability, improve crawl success and parsing quality, and partner with researchers to close the loop between data changes and model loss and evals.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build and own high-throughput, fully telemetered data pipelines with end-to-end traceability for frontier-scale data processing.
  • •Train and ship models to classify, rank, filter, clean, and identify data at extreme throughput in the critical path.
  • •Design and run scaling-ladder experiments for data mixture, repeatability, and quality depth, validating datasets with evidence.
  • •Build and operate platform pipelines that transform raw web/code/multimodal/acquired data into training-ready datasets, with orchestration, tooling, and reproducibility.
  • •Develop and scale web crawling systems that discover, schedule, fetch, and parse high-quality documents; improve URL seeding, scoring, host scheduling, crawl success, parsing quality, and automate delivery into the data pipeline.

Key Requirements

  • •Strong infrastructure or data platform background, with outlier depth in areas like crawling or search infrastructure (a plus).
  • •Ability to architect and ship end-to-end with high ownership, and debug complex systems independently.
  • •Experience designing and running experiments that turn qualitative dataset judgments into hard evidence.
  • •High-slope track record—moving unusually fast and taking staff-level ownership within a few years (example provided).
  • •Strong intuitions about large-scale distributed systems and interest in how pretraining data shapes model quality.
Experience:Machine learningDistributed systemsData engineeringWeb crawling
Skills:OwnershipDebuggingExperimentationPartnershipHigh throughput systems

Company Brief

Cursor
Builds Cursor, an AI-powered development platform and code editor that helps developers write, review, and maintain code using large language models and developer-focused automation tools for enterprise engineering teams.
Industry: Developer Tools
Company Size: Large (251 to 1,000 employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2022
Glassdoor
Glassdoor: 3.0
WebsiteLinkedInGlassdoor