Customer Reliability Engineer

FluidStack
San Francisco
Workplace: OnsiteFull timeUSD 204,000 - 284,000 annuallyFunction: Data Analytics & Business IntelligenceSkills: ["Extreme ownership","Autonomy","Methodical debugging","Incident communication","Cross-team follow-through"]

Own reliability for named customer workloads, including clusters, SLAs, and escalations. Debug training issues across the full stack—from hardware to fabric to scheduler—and handle customer-facing incident communication with technical depth. Convert recurring customer pain into durable engineering fixes by partnering with production teams, helping operate data-center infrastructure at extreme speed and scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
1 month ago

Customer Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Own reliability for named customer workloads, including clusters, SLAs, and escalations. Debug training issues across the full stack—from hardware to fabric to scheduler—and handle customer-facing incident communication with technical depth. Convert recurring customer pain into durable engineering fixes by partnering with production teams, helping operate data-center infrastructure at extreme speed and scale.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence

Key Responsibilities

  • •Own reliability for named customer workloads, including clusters, SLAs, and escalations.
  • •Debug issues across the full stack (hardware, fabric, scheduler) when training runs degrade.
  • •Run customer-facing incident communications with technical depth and no spin.
  • •Turn recurring customer pain into engineering fixes in partnership with production teams.

Pay and Benefits

Salary: USD 204,000 - 284,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • •Debug distributed systems methodically across layers you don't own.
  • •Write incident updates customers trusted more after reading.
  • •Push internal teams to fix causes, not symptoms, and follow up until they do.
  • •Bonus experience with GPU training workloads, including InfiniBand or RoCE, plus Slurm or Kubernetes and NCCL debugging.
Experience:HPCCloudAI labs
Skills:Extreme ownershipAutonomyMethodical debuggingIncident communicationCross-team follow-through
Tech Stack:GPU training workloadsInfiniBandRoCESlurmKubernetesNCCLClustersFabricScheduler

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor