Staff Research Engineer, Model Efficiency

Cohere
New York, San Francisco, Toronto, Montreal
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Communication","Problem-solving","Team collaboration"]

Lead research and engineering efforts to accelerate inference efficiency for Cohere’s foundation models. Develop, prototype, and deploy techniques to speed up model runtime, optimize MoE routing and decoding, and collaborate across a distributed team to push the boundaries of AI performance in production.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
9 months ago

Staff Research Engineer, Model Efficiency

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Lead research and engineering efforts to accelerate inference efficiency for Cohere’s foundation models. Develop, prototype, and deploy techniques to speed up model runtime, optimize MoE routing and decoding, and collaborate across a distributed team to push the boundaries of AI performance in production.
Location: New York, San Francisco, Toronto, Montreal
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Develop, prototype, and deploy techniques that materially improve how fast and efficiently models run in production.
  • •Collaborate with the Model Efficiency team to optimize model architecture, MoE routing, and decoding algorithms while ensuring model quality.
  • •Contribute to software/hardware co-design for GPU acceleration and performance optimization across the stack.
  • •Mentor teammates and share knowledge to uplift the broader team’s capabilities.
  • •Contribute to publications and presentations at top-tier conferences when applicable.

Pay and Benefits

Perks:Health InsuranceDentalRemote WorkParental LeaveCo-working StipendPaid Leave

Key Requirements

  • •PhD in Machine Learning or a related field
  • •Understand LLM architecture and how to optimize LLM inference given resource constraints
  • •Significant experience with techniques that enhance model efficiency
  • •Strong software engineering skills
  • •Publications at top-tier conferences (ICLR, ACL, NeurIPS)
Experience:Artificial IntelligenceMachine LearningLLMs
Education:PhD / Doctorate
Skills:CommunicationProblem-solvingTeam collaboration
Tech Stack:GPUMoELLMInference

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor