Research Engineer, Data Infrastructure (Language Modeling)

Cartesia
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 350,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Engineering execution","Experimentation","Data quality standards","Collaboration"]

Build and operate scalable data infrastructure for pretraining, acquiring, ingesting, processing, and curating massive text datasets. Design high-throughput, reproducible pipelines for ingestion through filtering, deduplication, and augmentation, and run ablation experiments to quantify how data sources and mixture weights impact model quality. Partner with research to co-design data loading, versioning, and experimentation workflows, while enforcing rigorous data quality standards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cartesia
Cartesia
1 day ago

Research Engineer, Data Infrastructure (Language Modeling)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 40 minutes agoStatus: Live

Job Summary

Build and operate scalable data infrastructure for pretraining, acquiring, ingesting, processing, and curating massive text datasets. Design high-throughput, reproducible pipelines for ingestion through filtering, deduplication, and augmentation, and run ablation experiments to quantify how data sources and mixture weights impact model quality. Partner with research to co-design data loading, versioning, and experimentation workflows, while enforcing rigorous data quality standards.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Build and operate performant, scalable data processing infrastructure to acquire, ingest, and combine massive text datasets.
  • •Design and operate high-throughput, reproducible data pipelines for ingestion, preprocessing, filtering, deduplication, and augmentation.
  • •Design and run ablation experiments to assess how data sources, processing choices, and mixture weights affect model quality.
  • •Partner with research and infrastructure teams on data loading, versioning, and experimentation pipelines.
  • •Establish rigorous data quality standards with a feedback loop between dataset characteristics and model behavior.

Pay and Benefits

Salary: USD 200,000 - 350,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kParental LeavePaid LeaveCommuter Benefits

Key Requirements

  • •Hands-on experience building ML data infrastructure, including training data pipelines, dataset versioning, and large-scale data loading.
  • •Strong engineering execution with clean, well-tested code and fluency with current tooling.
  • •Familiarity building and evaluating datasets for generative models, including how they are trained and used for inference.
  • •Experience with training-data workflows involving ingestion, preprocessing, filtering, deduplication, and augmentation.
  • •Ability to select appropriate tools and systems based on the problem rather than relying on familiar patterns.
Skills:Engineering executionExperimentationData quality standardsCollaboration
Tech Stack:RaySparkKubernetes

Company Brief

Cartesia
Builds real-time multimodal and voice AI (Sonic) that generates expressive, low-latency speech for conversational agents and on-device experiences, serving developers and enterprises with APIs and SDKs.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2023
WebsiteLinkedInGlassdoor