Technical Lead - GPU Infrastructure

Tether.io
Barcelona
Workplace: RemoteFull timeFunction: Administration & Executive AssistanceExperience: 8+ yearsEducation: bachelorsSkills: ["Leadership","Technical decision-making","Communication","Incident response"]

Own the end-to-end architecture and delivery of a self-hosted GPU compute and managed inference platform. Lead a distributed engineering team to run bare-metal Slurm scheduling and a Kubernetes-based control plane for NVIDIA GPU enablement, day-2 operations, and multi-tenant inference at scale. Drive platform observability, incident response, SLOs, and partner/vendor escalations, while translating internal research and model-training workloads into reliable infrastructure requirements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
2 days ago

Technical Lead - GPU Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Own the end-to-end architecture and delivery of a self-hosted GPU compute and managed inference platform. Lead a distributed engineering team to run bare-metal Slurm scheduling and a Kubernetes-based control plane for NVIDIA GPU enablement, day-2 operations, and multi-tenant inference at scale. Drive platform observability, incident response, SLOs, and partner/vendor escalations, while translating internal research and model-training workloads into reliable infrastructure requirements.
Location: Barcelona
Workplace: Remote
Employment Type: Full time · Permanent
Job Function: Administration & Executive Assistance
Seniority: Manager level

Key Responsibilities

  • •Own the platform architecture end to end, creating and maintaining architecture proposals and high-/low-level designs as the baseline.
  • •Lead and line-manage a distributed engineering team across backend (Node.js), frontend (React), DevOps, QA, and documentation, including standards, review, releases, and one-to-ones.
  • •Design, build, and operate a managed Slurm service for research users, including controller/accounting, partitions/login nodes, node onboarding, driver/CUDA baselines, health detection, drain/autohealing, storage visibility, and identity/isolation.
  • •Own Kubernetes control plane and GPU enablement on partner-provided bare metal, including cluster bootstrap/lifecycle, NVIDIA GPU Operator and Network Operator, KubeVirt/VFIO GPU isolation, and day-2 operations (upgrades, backup/recovery, node replacement).
  • •Design and operate managed inference at scale with serving architecture, multi-GPU/multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
Travel: Low travel

Key Requirements

  • •8+ years of hands-on engineering, including at least three years leading teams that build and operate infrastructure platforms; bachelor’s or master’s in computer science/engineering or equivalent experience.
  • •Hands-on operation of Slurm at scale (slurmctld/slurmdbd), including partitions/QoS/accounting, prolog/epilog, node health scripting, and upgrades with running jobs.
  • •GPU bare-metal operations: NVIDIA driver and CUDA lifecycle; Fabric Manager/NVSwitch behavior on SXM; DCGM health/utilization; MIG; node burn-in/acceptance.
  • •High-performance interconnects expertise (InfiniBand fabric/subnet configuration, RDMA, SR-IOV) and diagnosing multi-node NCCL performance problems.
  • •Production Kubernetes operations beyond deployment, including control plane lifecycle, upgrades, CNI/CSI, operators/custom controllers, and multi-tenancy design.
Experience:8+ yearsHPCGPU infrastructureKubernetesInfrastructure platformsResearch computingNVIDIA
Education:Bachelor's
Skills:LeadershipTechnical decision-makingCommunicationIncident response
Languages:English
Tech Stack:JavaScriptNode.jsReactKubernetesSlurmNVIDIACUDAFabric ManagerNVSwitchDCGMMIGInfiniBandRDMASR-IOVLinuxLinux kernel modulesVfio-pciCgroupsNamespacesPrometheus

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website