Member of Technical Staff, Model Efficiency

Cohere
New York, Toronto, San Francisco, Montreal
Workplace: OnsiteFull timeFunction: QA, Test & Release EngineeringExperience: 5+ yearsSkills: ["C++","Python","Rust","Go","CUDA","GPU"]

Join Cohere’s Model Efficiency team to optimize ML inference performance across the execution stack. You’ll drive low-latency, high-throughput improvements, collaborating with modeling and systems teams, and applying GPU/CUDA and kernel-level optimizations to advanced transformer architectures like MoE and large-scale models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
9 months ago

Member of Technical Staff, Model Efficiency

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Join Cohere’s Model Efficiency team to optimize ML inference performance across the execution stack. You’ll drive low-latency, high-throughput improvements, collaborating with modeling and systems teams, and applying GPU/CUDA and kernel-level optimizations to advanced transformer architectures like MoE and large-scale models.
Location: New York, Toronto, San Francisco, Montreal
Workplace: Onsite
Employment Type: Full time
Job Function: QA, Test & Release Engineering

Key Responsibilities

  • •Improve core performance metrics across the model execution stack by diagnosing bottlenecks and implementing optimizations.
  • •Collaborate with modeling and systems teams to experiment, measure, and ship improvements for faster inference.
  • •Develop and apply techniques for GPU/CUDA optimizations and kernel-level improvements.
  • •Explore model execution strategies for MoE and large-scale architectures.
  • •Ship features with a bias for action, measuring impact and iterating quickly.

Pay and Benefits

Perks:Health InsuranceDentalParental LeaveMeal AllowanceCo-working StipendRemote WorkPaid LeaveWellness Stipend

Key Requirements

  • •5+ years of experience writing high-performance, production-quality code
  • •Strong programming skills in C++ or Python (Rust/Go also welcome)
  • •Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.)
  • •Ability to diagnose and resolve performance bottlenecks across the model execution stack
  • •A strong bias for action — you ship fast, measure impact, and iterate
Experience:5+ yearsAIMLLLMInference
Skills:C++PythonRustGoCUDAGPU
Languages:English
Tech Stack:C++PythonRustGoCUDAGPU

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor