Member of Technical Staff, Model Efficiency
New York, Toronto, San Francisco, Montreal
Workplace: OnsiteFull timeFunction: QA, Test & Release EngineeringExperience: 5+ yearsSkills: ["C++","Python","Rust","Go","CUDA","GPU"]Join Cohere’s Model Efficiency team to optimize ML inference performance across the execution stack. You’ll drive low-latency, high-throughput improvements, collaborating with modeling and systems teams, and applying GPU/CUDA and kernel-level optimizations to advanced transformer architectures like MoE and large-scale models.
Loading
Loading job details...
Preparing the role view and application actions.

