Member of Technical Staff, Pre-Training Data

Cohere
Toronto, San Francisco, New York, London, Paris, Montreal
Workplace: RemoteFull timeFunction: Education & TrainingSkills: ["Collaboration","Teamwork","Problem-solving","Communication"]

Drive data-centric ML initiatives by building and refining data pipelines for pre-training language models. You will conduct data ablations, develop robust data models, curate data, and collaborate with researchers and engineers to optimize training throughput and accelerator utilization in a remote-friendly, multi-office environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
1 year ago

Member of Technical Staff, Pre-Training Data

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Drive data-centric ML initiatives by building and refining data pipelines for pre-training language models. You will conduct data ablations, develop robust data models, curate data, and collaborate with researchers and engineers to optimize training throughput and accelerator utilization in a remote-friendly, multi-office environment.
Location: Toronto, San Francisco, New York, London, Paris, Montreal
Workplace: Remote
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Conduct data ablations to assess data quality and experiment with data mixtures to enhance model performance.
  • •Develop robust data modeling techniques to ensure datasets are structured and formatted for optimal training efficiency.
  • •Research and implement innovative data curation methods, leveraging Cohere’s infrastructure to drive advancements in natural language processing.
  • •Collaborate with cross-functional teams, including researchers and engineers, to ensure data pipelines meet the demands of cutting-edge language models.
  • •Knowledge of data quality assessment techniques and experimentation with data mixtures to improve training outcomes.

Pay and Benefits

Perks:Health InsuranceDentalMeal AllowanceRemote WorkCo-working StipendParental Leave

Key Requirements

  • •Strong software engineering skills, with proficiency in Python and experience building data pipelines.
  • •Familiarity with curriculum learning, data mixing and data attribution.
  • •Familiarity with data processing frameworks such as Apache Spark, Apache Beam, Pandas, or similar tools.
  • •Experience working with large-scale datasets, including web data, code data, and multilingual corpora.
  • •Knowledge of data quality assessment techniques and experimentation with data mixtures.
Experience:AIMLNLPLanguage models
Skills:CollaborationTeamworkProblem-solvingCommunication
Tech Stack:PythonApache SparkApache BeamPandas

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor