Support Engineer, GPU Infrastructure

Hydra Host
Miami
Workplace: RemoteFull timeUSD 95,000 - 130,000 annuallyFunction: Product ManagementExperience: 3+ yearsSkills: ["Written communication","Sound judgment","Diagnostic reasoning"]

Own tier 2/3 production incidents for bare-metal GPU compute, diagnosing issues end-to-end from Linux through networking, storage, and hardware. Troubleshoot NVIDIA GPU servers and out-of-band management failures, coordinate partner remote hands, and provide engineering escalations with clear evidence packs. Document findings, improve runbooks and operational procedures, and automate repetitive support tasks using Python/Bash while collaborating to strengthen monitoring and platform reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Hydra Host
Hydra Host
1 day ago

Support Engineer, GPU Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Own tier 2/3 production incidents for bare-metal GPU compute, diagnosing issues end-to-end from Linux through networking, storage, and hardware. Troubleshoot NVIDIA GPU servers and out-of-band management failures, coordinate partner remote hands, and provide engineering escalations with clear evidence packs. Document findings, improve runbooks and operational procedures, and automate repetitive support tasks using Python/Bash while collaborating to strengthen monitoring and platform reliability.
Location: Miami
Workplace: Remote
Employment Type: Full time
Job Function: Product Management
Seniority: Mid level

Key Responsibilities

  • •Diagnose production Linux and infrastructure issues end-to-end, including boot/network boot failures, kernel/driver problems, filesystem/storage pressure, and instability.
  • •Troubleshoot server-side networking and isolate host vs network problems using routing/NIC/VLAN/MTU/DNS/DHCP/bonding knowledge and packet captures.
  • •Investigate hardware and GPU issues using out-of-band management and diagnostics, including NVIDIA GPU availability/thermal throttling, driver/VBIOS mismatches, PCIe issues, and XID errors.
  • •Own incidents through resolution or clean handoff, set severity by blast radius, coordinate remote hands with data center partners, and escalate to engineering with an evidence pack.
  • •Document solutions by writing runbooks, improve operational procedures through automation and infrastructure-as-code support, and help strengthen monitoring/alerting.

Pay and Benefits

Salary: USD 95,000 - 130,000 annually

Key Requirements

  • •3+ years supporting production servers, data center infrastructure, or bare metal and cloud environments.
  • •Strong hands-on Linux troubleshooting using logs/console tools, including dmesg, systemd, storage, and networking utilities.
  • •Real experience with server hardware (CPU/memory, storage/filesystems, RAID, PCIe, NICs, BIOS/UEFI, firmware, and drivers).
  • •Out-of-band management experience with IPMI, Redfish, iDRAC, iLO, or similar systems.
  • •Clear written English and experience working in ticketing, monitoring, incident management, or infrastructure management systems.
Experience:3+ years
Skills:Written communicationSound judgmentDiagnostic reasoning
Languages:English
Tech Stack:LinuxDmesgSystemdIPMIRedfishIDRACILOTCP/IPDNSDHCPVLANMTUNVIDIA GPUCUDANCCLNVLinkDCGMInfiniBandHigh performance EthernetPrometheus

Company Brief

Hydra Host
Hydra Host provides bare metal GPU infrastructure and an operating system for AI factories. It helps customers and data center operators deploy, manage, and monetize GPU compute across distributed data centers, with a focus on AI workloads and sovereign infrastructure.
Industry: Cloud Computing
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: Miami, United States
WebsiteLinkedIn