Engineering Manager, GPU Infrastructure
San Francisco, United States, Canada
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Technical mentorship","Communication","Data-driven decision-making","Problem-solving","Cross-team collaboration"]Lead the GPU Clusters team to build and operate superclusters powering Cohere’s frontier AI models. You’ll mentor and manage engineers, define and execute roadmaps for GPU cluster deployment, scheduling, fault detection, and performance optimization, and ensure reliability, scalability, and security. Collaborate with cloud providers and AI researchers to validate new GPU architectures, implement observability and infrastructure-as-code automation, drive cost optimization, and manage vendor relationships.
Loading
Loading job details...
Preparing the role view and application actions.

