Member of Technical Staff, Training Performance Engineer

Cohere
London, New York, Toronto, Montreal, Paris
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: ["Python","CUDA","Triton","JAX","PyTorch","XLA","MLIR","GPU kernels","Distributed training","Transformers"]

Performance Engineer in Cohere's Pre-Training team focusing on optimizing training performance and throughput for large language models. You will design high-performance software, write CUDA/Triton kernels, explore supercompute infrastructure, and collaborate with top researchers to push the efficiency and quality of model training.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
1 year ago

Member of Technical Staff, Training Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Performance Engineer in Cohere's Pre-Training team focusing on optimizing training performance and throughput for large language models. You will design high-performance software, write CUDA/Triton kernels, explore supercompute infrastructure, and collaborate with top researchers to push the efficiency and quality of model training.
Location: London, New York, Toronto, Montreal, Paris
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Design and write high-performant and scalable software for training.
  • •Understand architectural modifications and design choices and their effects on training throughput and quality.
  • •Write low-level CUDA, triton kernels to squeeze every last bit of performance from our accelerators.
  • •Research, implement, and experiment with ideas on our supercompute and data infrastructure.
  • •Learn from and work with the best researchers in the field.

Pay and Benefits

Perks:Health InsuranceDentalParental LeaveRemote WorkMeal AllowanceCo-working StipendPaid Leave

Key Requirements

  • •Extremely strong software engineering skills.
  • •Proficiency in Python and related ML frameworks such as JAX, Pytorch and XLA/MLIR.
  • •Experience writing kernels for GPUs using CUDA, triton, etc
  • •Experience using large-scale distributed training strategies.
  • •Familiarity with autoregressive sequence models, such as Transformers.
Skills:PythonCUDATritonJAXPyTorchXLAMLIRGPU kernelsDistributed trainingTransformers
Languages:English
Tech Stack:PythonML frameworksJAXPyTorchXLA/MLIRCUDATritonGPUDistributed trainingTransformers

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor