Engineering Manager, GPU Infrastructure
Cohere
San Francisco
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Technical mentorship","Communication","Data-driven decision-making","Problem-solving","Cross-team collaboration"]Lead the GPU Clusters team to build and operate superclusters powering Cohere’s frontier AI models. You’ll mentor and manage engineers, define and execute roadmaps for GPU cluster deployment, scheduling, fault detection, and performance optimization, and ensure reliability, scalability, and security. Collaborate with cloud providers and AI researchers to validate new GPU architectures, implement observability and infrastructure-as-code automation, drive cost optimization, and manage vendor relationships.

