Director of Infrastructure Engineering

RunPod
United States
Workplace: RemoteFull timeUSD 225,000 - 325,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Leadership","Stakeholder management","Communication","Ownership","Incident leadership"]

Lead and scale the core cloud and bare-metal infrastructure powering Runpod’s AI Developer Cloud. Own SRE, global networking, ultra-low-latency HPC cluster networks, and distributed storage engines to keep the platform highly available and low-latency at massive GPU scale. Hire and grow engineering managers and senior ICs, establish SRE practices, and partner across product and program to forecast capacity and deliver measurable reliability and performance improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
RunPod
RunPod
1 month ago

Director of Infrastructure Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Lead and scale the core cloud and bare-metal infrastructure powering Runpod’s AI Developer Cloud. Own SRE, global networking, ultra-low-latency HPC cluster networks, and distributed storage engines to keep the platform highly available and low-latency at massive GPU scale. Hire and grow engineering managers and senior ICs, establish SRE practices, and partner across product and program to forecast capacity and deliver measurable reliability and performance improvements.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Director level

Key Responsibilities

  • •Own core infrastructure and SRE by leading multiple teams responsible for reliability practices, SLA/SLOs, incident response, observability, and automated remediation.
  • •Architect HPC and global networking, scaling the global network backbone and ultra-low-latency HPC cluster networks, optimizing InfiniBand and RDMA over RoCE.
  • •Drive storage engine innovation by directing architecture and performance tuning for scalable distributed storage systems to maximize IOPS and throughput.
  • •Build a high-output organization by hiring, mentoring, and growing engineering managers and senior ICs, and establishing a remote-first culture of ownership and operational excellence.
  • •Translate scale into strategy by partnering with program management and product to forecast capacity, shape technical roadmaps, and define measurable outcomes.

Pay and Benefits

Salary: USD 225,000 - 325,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPaid LeaveEquityHome OfficeEquipment Stipend

Key Requirements

  • •7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, scaling high-availability cloud environments.
  • •8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms.
  • •Deep familiarity with ultra-low-latency networking, including InfiniBand and/or RoCE, spine-leaf architectures, and BGP.
  • •Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) for heavy AI/ML I/O.
  • •Strong SRE/DevOps foundation, including reliability engineering, IaC (Terraform, Ansible), Kubernetes, and modern observability stacks.
Experience:7+ years
Skills:LeadershipStakeholder managementCommunicationOwnershipIncident leadership
Tech Stack:Site Reliability Engineering (SRE)InfiniBandRoCESpine-leaf architecturesBGPCephLustreWekaNVMe-oFTerraformAnsibleKubernetesSlackInfrastructure as code (IaC)

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

RunPod
Provides on-demand GPU cloud and marketplace services for machine learning workloads, offering rentable GPU instances, scalable compute for training and inference, and tools to run ML jobs cost-effectively without long-term commitments.
Industry: Cloud Computing
Website