Site Reliability Engineer, Compute

FluidStack
San Francisco, New York, Austin, Seattle
Workplace: OnsiteFull timeUSD 175,000 - 300,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Extreme ownership","First-principles thinking","Incident management","Continuous learning"]

Own compute fleet health end to end for an AI infrastructure platform, building metrics pipelines, alerting, and unified visibility across Kubernetes-orchestrated workloads and bare metal. Turn deployments and repairs into automated pipelines covering detection, triage, parts management, and return to service. Expand GPU qualification and burn-in workflows, and operate low-level reliability tooling built on Redfish and BMC telemetry—at hyperscale speeds.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
2 months ago

Site Reliability Engineer, Compute

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own compute fleet health end to end for an AI infrastructure platform, building metrics pipelines, alerting, and unified visibility across Kubernetes-orchestrated workloads and bare metal. Turn deployments and repairs into automated pipelines covering detection, triage, parts management, and return to service. Expand GPU qualification and burn-in workflows, and operate low-level reliability tooling built on Redfish and BMC telemetry—at hyperscale speeds.
Location: San Francisco, New York, Austin, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own compute fleet health end to end by building metrics pipelines, alerting, and a unified health view for every GPU in production.
  • •Automate compute deployments and repairs into a pipeline covering detection, triage, parts management, and return to service.
  • •Design and expand the GPU qualification platform, including burn-in, performance baselining, and NPI execution.
  • •Own Redfish and BMC tooling, including firmware-level telemetry and fleet-scale log collection.
  • •Ensure end-to-end reliability, scalability, and operation of the compute fleet at large scale, with disciplined incident operations.

Pay and Benefits

Salary: USD 175,000 - 300,000 annually
Equity and Bonus:Equity
Perks:PensionHealth InsuranceDentalVisionPaid Leave

Key Requirements

  • •Treat toil as a bug and eliminate manual steps in repair workflows.
  • •Be comfortable reasoning about hardware failure modes at the firmware and silicon level.
  • •Navigate ambiguity: build understanding and explain it to others.
  • •Carry a pager, run incidents, write postmortems, and fix systemic causes.
  • •Be fluent with AI tooling (LLM APIs, MCP servers, agentic frameworks) and use Claude Code, Cursor, or similar daily.
Skills:Extreme ownershipFirst-principles thinkingIncident managementContinuous learning
Tech Stack:KubernetesBare metalRedfishBMCIPMIFirmware telemetryLLM APIsMCP serversAgentic frameworksClaude CodeCursorTemporalCadencePrometheusGrafanaGoPython

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor