HPC Infrastructure Engineer - GPU Clusters

ElevenLabs
London, New York, San Francisco, Warsaw
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Automation","Troubleshooting","Ownership"]

Own end-to-end operations for ElevenLabs’ NVIDIA GPU clusters, making compute fast, reliable, and “boring.” You’ll provision, schedule, monitor, and upgrade fleets; build automation for node health and remediation; and manage the software stack beneath training (CUDA, NCCL, drivers, container runtimes, high-speed networking). You’ll tune Slurm scheduling for throughput and fairness, troubleshoot performance issues, and validate rented GPU providers while keeping clusters secure by default.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ElevenLabs
ElevenLabs
1 day ago

HPC Infrastructure Engineer - GPU Clusters

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Own end-to-end operations for ElevenLabs’ NVIDIA GPU clusters, making compute fast, reliable, and “boring.” You’ll provision, schedule, monitor, and upgrade fleets; build automation for node health and remediation; and manage the software stack beneath training (CUDA, NCCL, drivers, container runtimes, high-speed networking). You’ll tune Slurm scheduling for throughput and fairness, troubleshoot performance issues, and validate rented GPU providers while keeping clusters secure by default.
Location: London, New York, San Francisco, Warsaw
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Operate and improve the GPU fleet end to end: provisioning, scheduling, monitoring, upgrades, and capacity planning.
  • •Build automation to keep the fleet healthy without human intervention (node health checks, automated draining/remediation, burn-in pipelines).
  • •Own the stack beneath training code: OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and high-speed networking (InfiniBand/RoCE).
  • •Run and tune job scheduling (Slurm or similar) to deliver fair and fast compute for researchers.
  • •Troubleshoot performance issues and validate rented GPU capacity (benchmarking, SLA validation), including hands-on hardware work and security-by-default controls.

Pay and Benefits

Perks:Learning BudgetTravel AllowanceCo-working StipendAnnual Offsite

Key Requirements

  • •Run large-scale Linux server or GPU environments in production and enjoy both building and operating.
  • •Know the NVIDIA stack well (drivers, CUDA, NCCL, DCGM) or have deep systems experience and can learn hardware stacks fast.
  • •Be comfortable with bare-metal environments, server hardware, and high-speed networking.
  • •Write solid automation in Python and/or Bash, with IaC tools like Ansible or Terraform.
  • •Troubleshoot using metrics/logs (including PromQL) and take end-to-end ownership, including datacenter trips when needed.
Experience:GPU infrastructureLinux serversML training workloads
Skills:AutomationTroubleshootingOwnership
Tech Stack:LinuxNVIDIAGPUCUDANCCLDCGMInfiniBandRoCESlurmPythonBashAnsibleTerraformIaCOS imagesNVIDIA driversContainer runtimesHigh-speed networkingPromQLPXE provisioning

Company Brief

ElevenLabs
Develops advanced AI audio models and tools for realistic text-to-speech, voice cloning, dubbing, music generation, and conversational voice agents for creators and enterprises.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: London, United Kingdom
Founded: 2022
Glassdoor
Glassdoor: 4.2
WebsiteLinkedInGlassdoor