Reliability Engineer, Supercomputing

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Data Analytics & Business IntelligenceEducation: bachelorsSkills: ["Collaboration","Initiative","Analytical thinking","Debugging","Writing"]

Own reliability for a GPU supercomputing fleet by diagnosing and remediating long-tail hardware issues that can derail frontier AI experiments. Work across hardware, firmware, and operating system boundaries—tracking root causes down to NIC/HBM/kernel-driver edge cases, building monitoring and analytics, and driving firmware lifecycle through qualification and rollout. Engage GPU/server/NIC/storage vendors, manage RMA flows, and publish postmortems that move fixes forward.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 month ago

Reliability Engineer, Supercomputing

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Own reliability for a GPU supercomputing fleet by diagnosing and remediating long-tail hardware issues that can derail frontier AI experiments. Work across hardware, firmware, and operating system boundaries—tracking root causes down to NIC/HBM/kernel-driver edge cases, building monitoring and analytics, and driving firmware lifecycle through qualification and rollout. Engage GPU/server/NIC/storage vendors, manage RMA flows, and publish postmortems that move fixes forward.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence

Key Responsibilities

  • •Investigate, reproduce, and remediate issues across large GPU clusters.
  • •Own drivers, kernel surface, and diagnostics spanning hardware, firmware, and OS.
  • •Automate fleet reliability monitoring and analyze error rates to measure failure reductions.
  • •Drive firmware lifecycle (tracking, qualification, staged rollout, regression analysis).
  • •Engage vendors directly and manage RMA flows when hardware needs replacement; write postmortems and vendor cases.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeaveParental LeaveRelocation

Key Requirements

  • •Bachelor’s degree (or equivalent) in computer science, engineering, or similar.
  • •Proficiency in at least one backend language (Python or Rust).
  • •Experience operating large-scale clusters and container orchestration systems (Kubernetes or Slurm).
  • •Comfort operating across the stack and owning end-to-end projects.
  • •Fluency with Linux systems and debugging tools, plus ability to analyze reliability and debug issues to root cause.
Education:Bachelor's
Skills:CollaborationInitiativeAnalytical thinkingDebuggingWriting
Tech Stack:PythonRustKubernetesSlurmLinuxLinux kernelBMCIDRACIPMIRedfishDCGMNVLinkNVSwitchFabric managerKernel driversFirmwareGPUNICHBM

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website