Technical Lead - GPU Infrastructure

Tether.io
Bucharest
Workplace: RemoteFull timeFunction: Administration & Executive AssistanceExperience: 8+ yearsEducation: bachelorsSkills: ["Leadership","Written communication","Cross-team collaboration","Architecture decision-making","Incident response"]

Lead the architecture and delivery of a self-hosted GPU compute and managed inference platform for internal research and external inference tenancy. Own end-to-end platform designs, line-manage a distributed engineering team, and build/operate a managed Slurm scheduling layer on bare-metal GPUs. Run and evolve a JavaScript-based Kubernetes control plane with GPU enablement, multi-tenancy, autoscaling, observability, and incident response, while partnering closely with infrastructure vendors.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
2 days ago

Technical Lead - GPU Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Lead the architecture and delivery of a self-hosted GPU compute and managed inference platform for internal research and external inference tenancy. Own end-to-end platform designs, line-manage a distributed engineering team, and build/operate a managed Slurm scheduling layer on bare-metal GPUs. Run and evolve a JavaScript-based Kubernetes control plane with GPU enablement, multi-tenancy, autoscaling, observability, and incident response, while partnering closely with infrastructure vendors.
Location: Bucharest
Workplace: Remote
Employment Type: Full time · Permanent
Job Function: Administration & Executive Assistance
Seniority: Manager level

Key Responsibilities

  • •Own the platform architecture end to end, including architecture proposals and high/low-level designs kept current as the baseline.
  • •Line-manage and lead a distributed engineering team (backend Node.js, frontend React, DevOps, QA and documentation) with engineering standards, review processes, release gates, and one-to-ones.
  • •Design, build, and operate a managed Slurm scheduling layer on bare-metal GPUs: controller/accounting, partitions/login nodes, node onboarding, driver/CUDA baseline/upgrade, health detection, drain/autohealing, and storage/identity/isolation.
  • •Own Kubernetes control plane and GPU enablement: cluster bootstrap/lifecycle on partner bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt/VFIO), and day-2 operations (upgrades, backup/recovery, node replacement).
  • •Deliver managed inference at scale with serving architecture, multi-GPU/multi-node parallelism, autoscaling, request routing and endpoint reliability, plus observability, incident response, and post-incident reviews.
Travel: Low travel

Key Requirements

  • •8+ years of hands-on engineering, including at least three years leading teams that build and operate infrastructure platforms others depend on; Bachelor’s/Master’s in CS/engineering (or equivalent).
  • •Hands-on operation of Slurm at scale (slurmctld/slurmdbd, partitions/QoS/accounting, prolog/epilog, node health scripting, upgrades) and ideally experience running HPC/GPU research clusters.
  • •Operate bare-metal NVIDIA GPU fleets end to end (driver/CUDA lifecycle, Fabric Manager/NVSwitch behavior, DCGM health/utilisation, MIG, node burn-in/acceptance).
  • •Hands-on experience with high-performance interconnects and diagnosing GPU training issues (InfiniBand subnet/config, RDMA, SR-IOV, troubleshooting multi-node NCCL performance).
  • •Deep Linux and production orchestration skills: kernel modules/drivers, PCIe passthrough/vfio-pci, cgroups/namespaces; production Kubernetes operations including control plane lifecycle, operators/custom controllers, and multi-tenancy design.
Experience:8+ yearsHPCGPU infrastructureInferenceKubernetesDistributed systems
Education:Bachelor's
Skills:LeadershipWritten communicationCross-team collaborationArchitecture decision-makingIncident response
Languages:English
Tech Stack:KubernetesJavaScriptNode.jsReactDevOpsSlurmSlurmctldSlurmdbdNVIDIACUDAFabric ManagerNVSwitchDCGMMIGInfiniBandNCCLRDMASR-IOVLinuxKernel modules

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website