Datacenter Infrastructure Specialist

RunPod
United States
Workplace: RemoteFull timeUSD 120,000 - 160,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 3-5 yearsSkills: ["Communication","Problem-solving","Proactive","Detail-oriented","Ownership"]

Own the technical lifecycle and operational health of Runpod’s rapidly expanding, high-density GPU fleet. Serve as the technical authority between hardware partners and engineering teams, ensuring uptime and SLA enforcement through deep HPC systems engineering, advanced network troubleshooting (including RDMA fabrics), and process automation. Use Linux, Docker, observability stacks, and AI-driven runbooks to validate new hardware, diagnose incidents, and scale reliable infrastructure for demanding AI workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
RunPod
RunPod
2 days ago

Datacenter Infrastructure Specialist

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live
Reposted: similar role first listed 3 weeks ago

Job Summary

Own the technical lifecycle and operational health of Runpod’s rapidly expanding, high-density GPU fleet. Serve as the technical authority between hardware partners and engineering teams, ensuring uptime and SLA enforcement through deep HPC systems engineering, advanced network troubleshooting (including RDMA fabrics), and process automation. Use Linux, Docker, observability stacks, and AI-driven runbooks to validate new hardware, diagnose incidents, and scale reliable infrastructure for demanding AI workloads.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Validate new hardware and benchmark deployments to ensure they meet Runpod specifications for distributed AI/ML workloads.
  • •Monitor fleet health to detect performance degradation, audit downtime, and provide technical data to protect customer SLAs.
  • •Automate network triage and generate dynamic runbooks using an AI-first approach with LLMs and AI agents.
  • •Coordinate technical incident communications and translate outages into actionable resolutions.
  • •Provide technical support to infrastructure partners to support their growth.

Pay and Benefits

Salary: USD 120,000 - 160,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPaid LeaveHome Office

Key Requirements

  • •3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
  • •Strong proficiency in datacenter networking and performance troubleshooting; exposure to RDMA, InfiniBand, or RoCE is preferred.
  • •Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and multi-node performance tuning.
  • •Strong Linux system administration skills and experience with containerization (Docker), including kernel/hardware interface troubleshooting and performance tuning.
  • •Clear written and verbal communication skills to explain hardware/networking issues to technical partners and internal leadership.
Experience:3-5 years
Skills:CommunicationProblem-solvingProactiveDetail-orientedOwnership
Tech Stack:LinuxDockerNVIDIA Software StackRDMAInfiniBandRoCEGrafanaPrometheusDatadogPythonGo (Golang)BashSlack

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

RunPod
Provides on-demand GPU cloud and marketplace services for machine learning workloads, offering rentable GPU instances, scalable compute for training and inference, and tools to run ML jobs cost-effectively without long-term commitments.
Industry: Cloud Computing
Website