Site Reliability Engineer, Compute
FluidStack
San Francisco, New York, Austin, Seattle
Workplace: OnsiteFull timeUSD 175,000 - 300,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Extreme ownership","First-principles thinking","Incident management","Continuous learning"]Own compute fleet health end to end for an AI infrastructure platform, building metrics pipelines, alerting, and unified visibility across Kubernetes-orchestrated workloads and bare metal. Turn deployments and repairs into automated pipelines covering detection, triage, parts management, and return to service. Expand GPU qualification and burn-in workflows, and operate low-level reliability tooling built on Redfish and BMC telemetry—at hyperscale speeds.

