Reliability Engineer, Supercomputing
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Data Analytics & Business IntelligenceEducation: bachelorsSkills: ["Collaboration","Initiative","Analytical thinking","Debugging","Writing"]Own reliability for a GPU supercomputing fleet by diagnosing and remediating long-tail hardware issues that can derail frontier AI experiments. Work across hardware, firmware, and operating system boundaries—tracking root causes down to NIC/HBM/kernel-driver edge cases, building monitoring and analytics, and driving firmware lifecycle through qualification and rollout. Engage GPU/server/NIC/storage vendors, manage RMA flows, and publish postmortems that move fixes forward.
Loading
Loading job details...
Preparing the role view and application actions.

