Engineering Manager, GPU Infrastructure

Cohere
San Francisco, United States, Canada
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Technical mentorship","Communication","Data-driven decision-making","Problem-solving","Cross-team collaboration"]

Lead the GPU Clusters team to build and operate superclusters powering Cohere’s frontier AI models. You’ll mentor and manage engineers, define and execute roadmaps for GPU cluster deployment, scheduling, fault detection, and performance optimization, and ensure reliability, scalability, and security. Collaborate with cloud providers and AI researchers to validate new GPU architectures, implement observability and infrastructure-as-code automation, drive cost optimization, and manage vendor relationships.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
1 month ago

Engineering Manager, GPU Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Lead the GPU Clusters team to build and operate superclusters powering Cohere’s frontier AI models. You’ll mentor and manage engineers, define and execute roadmaps for GPU cluster deployment, scheduling, fault detection, and performance optimization, and ensure reliability, scalability, and security. Collaborate with cloud providers and AI researchers to validate new GPU architectures, implement observability and infrastructure-as-code automation, drive cost optimization, and manage vendor relationships.
Location: San Francisco, United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Lead and mentor a team of engineers specializing in GPU infrastructure, including performance management, career development, and hiring.
  • •Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • •Oversee implementation of topology-aware scheduling, hardware fault detection, and performance optimization systems.
  • •Establish observability and monitoring frameworks for GPU utilization, performance, and reliability, and implement infrastructure-as-code automation for cluster provisioning.
  • •Collaborate with AI researchers and cloud providers to validate and deploy new GPU architectures, while ensuring infrastructure reliability, scalability, and security.

Pay and Benefits

Perks:Meal AllowanceHealth InsuranceDental401kPensionParental LeavePaid LeaveLearning StipendHome OfficeCo-working Benefit

Key Requirements

  • •Experience managing engineering teams with a focus on technical mentorship and growth.
  • •Deep expertise in ML/HPC infrastructure, including GPU/TPU clusters and high-performance computing environments.
  • •Proven experience with Kubernetes at scale for AI workloads in multi-cloud environments.
  • •Knowledge of infrastructure monitoring tools such as Prometheus and Grafana.
  • •Familiarity with infrastructure-as-code tools like Terraform and ArgoCD, plus experience with cost optimization and capacity planning for GPU infrastructure.
Experience:AIMLHPCCloud infrastructureMulti-cloudKubernetesDistributed systems
Skills:Technical mentorshipCommunicationData-driven decision-makingProblem-solvingCross-team collaboration
Tech Stack:GPUTPUHPCJAXPyTorchTensorFlowKubernetesPrometheusGrafanaTerraformArgoCDInfrastructure-as-codeTopology-aware schedulingDistributed training

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor