Member of Technical Staff, Training Infra Engineer

Cohere
Paris, London, Toronto, New York, Montreal, San Francisco
Workplace: RemoteFull timeFunction: Education & TrainingSkills: ["Python","ML frameworks","JAX","PyTorch","XLA/MLIR","Kubernetes","Slurm","Ray","Distributed training","Large-scale training"]

Design and maintain scalable training software and infrastructure for Cohere’s model training pipelines, bridging research and production. Contribute to high-performance training systems, optimize infrastructure, and collaborate with top researchers to accelerate large-scale AI model training across a distributed compute environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
1 year ago

Member of Technical Staff, Training Infra Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Design and maintain scalable training software and infrastructure for Cohere’s model training pipelines, bridging research and production. Contribute to high-performance training systems, optimize infrastructure, and collaborate with top researchers to accelerate large-scale AI model training across a distributed compute environment.
Location: Paris, London, Toronto, New York, Montreal, San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Design and write high-performant and scalable software for training.
  • •Improve our training setup from an infrastructure and codebase performance standpoint.
  • •Craft and implement tools to speed up our training cycles and improve the overall efficacy of our training infrastructure
  • •Research, implement, and experiment with ideas on our supercompute and data infrastructure.
  • •Learn from and work with the best researchers in the field.

Pay and Benefits

Perks:Health InsuranceDentalParental LeaveCo-working Stipend

Key Requirements

  • •Extremely strong software engineering skills.
  • •Proficiency in Python and related ML frameworks such as JAX, Pytorch and XLA/MLIR.
  • •Experience with distributed training infrastructures (Kubernetes, Slurm) and associated frameworks (Ray)
  • •Experience using large-scale distributed training strategies.
  • •Hands on experience on training large model at scale and having contributed to the tooling and/or setup of the training infrastructure
Experience:AIMLDistributed trainingPythonResearch to production
Skills:PythonML frameworksJAXPyTorchXLA/MLIRKubernetesSlurmRayDistributed trainingLarge-scale training
Languages:English
Tech Stack:PythonJAXPyTorchXLA/MLIRKubernetesSlurmRay

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor