Technical Lead - GPU Infrastructure

Tether.io
Bengaluru
Workplace: RemoteFull timeFunction: Administration & Executive AssistanceExperience: 8+ yearsEducation: bachelorsSkills: ["Leadership","Written and spoken communication","Architecture decision-making","Incident response","Cross-time-zone collaboration"]

Own the end-to-end architecture and delivery of a self-hosted GPU compute and managed inference platform (Cosmic AC). Lead a distributed engineering team (~12) building Kubernetes-based control plane and a managed Slurm scheduling layer for research workloads on bare-metal NVIDIA infrastructure. Drive GPU enablement, cluster bootstrap/lifecycle, autoscaling and reliability for multi-tenant inference, and observability/incident response. Serve as primary technical interface to infrastructure partners and vendors.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
1 day ago

Technical Lead - GPU Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Own the end-to-end architecture and delivery of a self-hosted GPU compute and managed inference platform (Cosmic AC). Lead a distributed engineering team (~12) building Kubernetes-based control plane and a managed Slurm scheduling layer for research workloads on bare-metal NVIDIA infrastructure. Drive GPU enablement, cluster bootstrap/lifecycle, autoscaling and reliability for multi-tenant inference, and observability/incident response. Serve as primary technical interface to infrastructure partners and vendors.
Location: Bengaluru
Workplace: Remote
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Manager level

Key Responsibilities

  • •Own platform architecture end to end, including architecture proposals and high/low-level designs, and keep the baseline current.
  • •Line-manage and lead a distributed team (~12) across backend, frontend, DevOps, QA, and documentation; drive engineering standards, code/design review, release gates, and one-to-ones.
  • •Design, build, and operate a managed Slurm scheduling service for research users (controller/accounting, partitions/login nodes, onboarding/acceptance, upgrades, health detection, drain/autohealing, storage visibility, and identity/isolation).
  • •Own Kubernetes control-plane and GPU enablement on partner-provided bare metal, including cluster bootstrap/lifecycle, NVIDIA GPU/network operator enablement, upgrades, backup/recovery, and node replacement.
  • •Deliver managed inference at scale and lead observability/operations (metrics/logging/alerting/SLOs, incident response and post-incident review, and keeping on-call sustainable).
Travel: Low travel

Key Requirements

  • •8+ years hands-on engineering experience, including at least three years leading teams building and operating infrastructure platforms; Bachelor’s or Master’s in computer science/engineering (or equivalent).
  • •Hands-on operation of Slurm at scale (slurmctld/slurmdbd) for real users: partitions, QoS/priority, accounting, prolog/epilog, node-health scripting, and upgrades; ideally operated an HPC/GPU research training cluster.
  • •Bare-metal NVIDIA GPU fleet operation: driver and CUDA lifecycle, Fabric Manager/NVSwitch behavior, DCGM health/utilization, MIG, and node burn-in/acceptance.
  • •High-performance interconnect expertise: InfiniBand fabric/subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues.
  • •Deep Linux systems knowledge (kernel modules/drivers, PCIe passthrough/vfio-pci, cgroups/namespaces) and production Kubernetes operation (control plane, upgrades, CNI/CSI, operators/custom controllers) with multi-tenancy design.
Experience:8+ yearsFintechGPU infrastructureKubernetesHPCAIDistributed systems
Education:Bachelor's
Skills:LeadershipWritten and spoken communicationArchitecture decision-makingIncident responseCross-time-zone collaboration
Languages:English
Tech Stack:JavaScriptNode.jsReactKubernetesSlurmNVIDIA GPU OperatorNetwork OperatorKubeVirtVFIOLinuxCUDAFabric ManagerNVSwitchDCGMMIGInfiniBandRDMASR-IOVNCCLPCIe

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website