Technical Lead - GPU Infrastructure

Tether.io
Yerevan
Workplace: RemoteFull timeFunction: Administration & Executive AssistanceEducation: bachelorsSkills: ["Leadership","Written communication","Incident response","Architecture decision-making","Cross-functional collaboration"]

Own end-to-end architecture and delivery of a GPU compute and managed inference platform, spanning bare-metal Slurm scheduling and a Kubernetes control plane with NVIDIA GPU enablement. Lead and line-manage a distributed engineering team while partnering with infrastructure vendors and internal research/model-training users. Build for multi-GPU/multi-node scalability, autoscaling, request routing, and reliability—backed by strong observability, SLOs, and hands-on incident response within a fixed delivery window.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
2 days ago

Technical Lead - GPU Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Own end-to-end architecture and delivery of a GPU compute and managed inference platform, spanning bare-metal Slurm scheduling and a Kubernetes control plane with NVIDIA GPU enablement. Lead and line-manage a distributed engineering team while partnering with infrastructure vendors and internal research/model-training users. Build for multi-GPU/multi-node scalability, autoscaling, request routing, and reliability—backed by strong observability, SLOs, and hands-on incident response within a fixed delivery window.
Location: Yerevan
Workplace: Remote
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Manager level

Key Responsibilities

  • •Own platform architecture end to end, including proposals and high/low-level designs, and keep designs current as the baseline.
  • •Line-manage and lead a distributed team across backend (Node.js), frontend (React), DevOps, QA, and documentation: standards, reviews, releases, one-to-ones, and performance input.
  • •Design, build, and operate a managed Slurm scheduling layer on bare metal for research users, including accounting, partitions/login nodes, onboarding/acceptance, driver/CUDA baseline and upgrades, health detection, draining/autohealing, and storage visibility and isolation.
  • •Own Kubernetes control plane and GPU enablement: cluster bootstrap/lifecycle on partner bare metal, NVIDIA GPU operator and Network operator, GPU isolation (KubeVirt/VFIO), and day-2 operations like upgrades, backup/recovery, and node replacement.
  • •Deliver managed inference at scale with serving architecture, multi-GPU/multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity; run observability and incident response with SLOs and post-incident reviews.
Travel: Low travel

Key Requirements

  • •8+ years of hands-on engineering, including 3+ years leading teams building and operating infrastructure platforms; Bachelor’s or Master’s in computer science/engineering (or equivalent).
  • •Run Slurm at scale hands-on (slurmctld/slurmdbd), including partitions/QoS/accounting, prolog/epilog, node health scripting, and upgrades; ideally operated HPC/GPU clusters for research users.
  • •Operate bare-metal NVIDIA GPU infrastructure (driver/CUDA lifecycle, Fabric Manager/NVSwitch behaviour, DCGM health/utilisation, MIG, node burn-in/acceptance).
  • •Diagnose high-performance interconnect issues (InfiniBand fabric, subnet/RDMA, SR-IOV) and multi-node NCCL performance problems.
  • •Strong Linux and Kubernetes depth: kernel modules/drivers, PCIe passthrough/vfio-pci, cgroups/namespaces; production Kubernetes operations with control plane lifecycle, CNI/CSI, operators/custom controllers, and multi-tenancy design.
Experience:FintechBlockchainAIHPCGPU infrastructure
Education:Bachelor's
Skills:LeadershipWritten communicationIncident responseArchitecture decision-makingCross-functional collaboration
Languages:English
Tech Stack:JavaScriptNode.jsReactKubernetesSlurmNVIDIA GPU OperatorNetwork OperatorKubeVirtVFIOCUDANVSwitchFabric ManagerDCGMInfiniBandNCCLRDMASR-IOVLinuxPCIeVfio-pci

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website