HPC Infrastructure Engineer - GPU Clusters
London, New York, San Francisco, Warsaw
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Automation","Troubleshooting","Ownership"]Own end-to-end operations for ElevenLabs’ NVIDIA GPU clusters, making compute fast, reliable, and “boring.” You’ll provision, schedule, monitor, and upgrade fleets; build automation for node health and remediation; and manage the software stack beneath training (CUDA, NCCL, drivers, container runtimes, high-speed networking). You’ll tune Slurm scheduling for throughput and fairness, troubleshoot performance issues, and validate rented GPU providers while keeping clusters secure by default.
Loading
Loading job details...
Preparing the role view and application actions.

