Data Infrastructure Engineer, Pre-training

Anthropic
San Francisco
Workplace: OnsiteFull timeUSD 500,000 - 850,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Attention to detail","Communication","Collaboration","Ownership"]

Build highly performant, reproducible data processing infrastructure for large language model pre-training. Develop and scale core processing primitives (tokenization, deduplication, chunking), implement data quality assurance, and design end-to-end pipelines that transform web-scale corpora into training-ready datasets. Collaborate with research teams to translate novel research requirements into robust distributed systems, leveraging frameworks like Apache Spark to deliver fault-tolerant, traceable infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
10 months ago

Data Infrastructure Engineer, Pre-training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Build highly performant, reproducible data processing infrastructure for large language model pre-training. Develop and scale core processing primitives (tokenization, deduplication, chunking), implement data quality assurance, and design end-to-end pipelines that transform web-scale corpora into training-ready datasets. Collaborate with research teams to translate novel research requirements into robust distributed systems, leveraging frameworks like Apache Spark to deliver fault-tolerant, traceable infrastructure.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design and implement data processing infrastructure for large language model training with performance, reproducibility, and traceability.
  • •Develop and maintain core processing primitives such as tokenization, deduplication, and chunking at scale.
  • •Build robust systems for data quality assurance and validation at scale.
  • •Collaborate with research teams to implement novel data processing architectures.
  • •Build and operate end-to-end data pipelines converting raw web-scale corpora into training-ready datasets.

Pay and Benefits

Salary: USD 500,000 - 850,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •5+ years of experience outside of internships.
  • •Strong software engineering skills building high-throughput, fault-tolerant distributed systems.
  • •Hands-on experience with distributed computing frameworks, particularly Apache Spark.
  • •Advanced degree in Computer Science or a related field.
  • •Experience with language model training infrastructure and/or data infrastructure, MLOps, or ML infrastructure.
Experience:5+ yearsLarge language modelsData infrastructureMLOpsML infrastructure
Education:Bachelor's
Skills:Problem-solvingAttention to detailCommunicationCollaborationOwnership
Languages:English
Tech Stack:Apache SparkPythonRust

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn