Site Reliability Engineer, Inference Infrastructure

Cohere
Toronto, San Francisco, New York, Montreal
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Collaboration","Troubleshooting","Communication","Problem-solving","Teamwork"]

Join Cohere as a Site Reliability Engineer focused on inference infrastructure. You’ll build scalable, high-availability systems for ML platforms, develop Kubernetes-based operators for language model deployments, automate observability, participate in on-call SLOs, and collaborate across teams to improve the infrastructure roadmap and knowledge sharing.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
7 months ago

Site Reliability Engineer, Inference Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join Cohere as a Site Reliability Engineer focused on inference infrastructure. You’ll build scalable, high-availability systems for ML platforms, develop Kubernetes-based operators for language model deployments, automate observability, participate in on-call SLOs, and collaborate across teams to improve the infrastructure roadmap and knowledge sharing.
Location: Toronto, San Francisco, New York, Montreal
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Build self-service systems that automate managing, deploying and operating services.
  • •This includes our custom Kubernetes operators that support language model deployments.
  • •Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems.
  • •Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation.
  • •Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback.

Pay and Benefits

Perks:Health InsuranceDentalRemote WorkMeal AllowanceParental Leave

Key Requirements

  • •5+ years of engineering experience running production infrastructure at a large scale
  • •Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters
  • •Experience with Kubernetes dev and production coding and support
  • •Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving
  • •Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments
Experience:5+ yearsAIMachine LearningCloud Platforms
Skills:CollaborationTroubleshootingCommunicationProblem-solvingTeamwork
Languages:English
Tech Stack:KubernetesGPUsTPUsGolangC++LinuxAWSGCPAzureOCIMulti-cloudOn-premKubernetes operators

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor