Senior ML Systems Engineer, Frameworks & Tooling

Cohere
London, Paris, New York, Toronto, Montreal, San Francisco
Workplace: HybridFull timeFunction: IT Operations (Systems/Network Admin)Skills: ["Collaboration","Problem-solving","Communication","Teamwork"]

Senior ML Systems Engineer will build and maintain the training framework for large-scale language models, designing distributed training abstractions and boosting throughput on multi-node clusters. You’ll develop tooling for monitoring and debugging, collaborate with infra to optimize Slurm, containers, and hardware, and help ensure reproducible, scalable runs. You’ll work across ML systems with autonomy and high impact.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
9 months ago

Senior ML Systems Engineer, Frameworks & Tooling

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Senior ML Systems Engineer will build and maintain the training framework for large-scale language models, designing distributed training abstractions and boosting throughput on multi-node clusters. You’ll develop tooling for monitoring and debugging, collaborate with infra to optimize Slurm, containers, and hardware, and help ensure reproducible, scalable runs. You’ll work across ML systems with autonomy and high impact.
Location: London, Paris, New York, Toronto, Montreal, San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Build and own the training framework responsible for large-scale LLM training.
  • •Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing).
  • •Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100).
  • •Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
  • •Collaborate closely with infra teams to ensure Slurm setups, container environments, and hardware configurations support high-performance training.

Pay and Benefits

Perks:Health InsuranceDentalParental LeaveMeal AllowanceWellness StipendRemote WorkPaid LeaveCo-working Stipend

Key Requirements

  • •Strong engineering experience in large-scale distributed training or HPC systems.
  • •Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).
  • •Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.
  • •Experience working with containerized environments (Docker, Singularity/Apptainer).
  • •A track record of building tools that increase developer velocity for ML teams.
Experience:AIMLDistributed systems
Skills:CollaborationProblem-solvingCommunicationTeamwork
Languages:English
Tech Stack:JAXPyTorchDeepSpeedMegatronXFormersCUDANCCLSlurmRayKubernetesDockerSingularityApptainer

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor