Software Engineer, GPU Infrastructure (HPC)

Cohere
Canada, United States, San Francisco
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["Mentorship","Collaboration","Problem-solving"]

Staff Software Engineer roles involve building and scaling ML-optimized HPC infrastructure, deploying Kubernetes-based GPU/TPU superclusters across clouds, and collaborating with AI researchers to optimize training workloads. You’ll drive performance, reliability, and observability while enabling researchers with self-service tooling and advancing ML infrastructure through IaC practices and open-source collaboration.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
7 months ago

Software Engineer, GPU Infrastructure (HPC)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Staff Software Engineer roles involve building and scaling ML-optimized HPC infrastructure, deploying Kubernetes-based GPU/TPU superclusters across clouds, and collaborating with AI researchers to optimize training workloads. You’ll drive performance, reliability, and observability while enabling researchers with self-service tooling and advancing ML infrastructure through IaC practices and open-source collaboration.
Location: Canada, United States, San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build and scale ML-optimized HPC infrastructure: Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads.
  • •Optimize for AI/ML training: Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance, leveraging technologies like RDMA, NCCL, and high-speed interconnects.
  • •Troubleshoot and resolve complex issues: Proactively identify and resolve infrastructure bottlenecks, performance degradation, and system failures to ensure minimal disruption to AI/ML workflows.
  • •Enable researchers with self-service tools: Design intuitive interfaces and workflows that allow researchers to monitor, debug, and optimize their training jobs independently.
  • •Drive innovation in ML infrastructure: Work closely with AI researchers to understand emerging needs (e.g., JAX, PyTorch, distributed training) and translate them into robust, scalable infrastructure solutions.

Pay and Benefits

Perks:Remote WorkHealth InsuranceParental LeaveMeal AllowanceCo-working StipendPaid Leave

Key Requirements

  • •Deep expertise in ML/HPC infrastructure: experience with GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and HPC environments.
  • •Kubernetes at scale: ability to deploy, manage, and troubleshoot cloud-native Kubernetes clusters for AI workloads.
  • •Strong programming skills: proficiency in Python for ML tooling and Go for systems engineering, with preference for open-source contributions.
  • •Low-level systems knowledge: familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
  • •Research collaboration experience: track record of working closely with AI researchers or ML engineers to solve infrastructure challenges.
Experience:MLHPCInfrastructureKubernetesGPUTPUDistributed training
Skills:MentorshipCollaborationProblem-solving
Tech Stack:KubernetesRDMANCCLJAXPyTorchTensorFlowLinuxPythonGo

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor