Staff Software Engineer, Inference Infrastructure

Cohere
San Francisco, New York, Toronto, Montreal
Workplace: RemoteFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Collaboration","Troubleshooting","Problem-solving","Communication","Teamwork"]

Build, deploy, and operate a high-performance ML infrastructure platform delivering Cohere’s large language models via API endpoints. Collaborate across teams to deploy optimized NLP models with low latency, high throughput, and high availability, while interfacing with customers to tailor deployments and ensure smooth, scalable operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
7 months ago

Staff Software Engineer, Inference Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Build, deploy, and operate a high-performance ML infrastructure platform delivering Cohere’s large language models via API endpoints. Collaborate across teams to deploy optimized NLP models with low latency, high throughput, and high availability, while interfacing with customers to tailor deployments and ensure smooth, scalable operations.
Location: San Francisco, New York, Toronto, Montreal
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Develop, deploy, and operate Cohere's AI platform delivering Cohere's large language models through API endpoints.
  • •Work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments.
  • •Interface with customers and create customized deployments to meet their specific needs.
  • •Collaborate across teams to ensure smooth operations and efficient teamwork for mission-critical systems.
  • •Investigate and troubleshoot complex infrastructure issues to improve reliability and performance.

Pay and Benefits

Perks:Health InsuranceDentalRemote WorkMeal AllowanceParental LeaveWellness Stipend

Key Requirements

  • •5+ years of engineering experience running production infrastructure at a large scale
  • •Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters
  • •Experience with Kubernetes dev and production coding and support
  • •Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving
  • •Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments
Experience:5+ yearsAIMLNLPSaaS
Skills:CollaborationTroubleshootingProblem-solvingCommunicationTeamwork
Tech Stack:KubernetesGPUGCPAzureAWSOCIMulti-cloudLinuxGolangC++

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor