Production Engineer, Compute

FluidStack
San Francisco, New York, Seattle, Austin
Workplace: HybridFull timeUSD 175,000 - 300,000 annuallyFunction: Manufacturing & Production OperationsSkills: ["Kubernetes","Go","Python","Prometheus","Grafana","Redfish","IPMI","GPU","TPU","Temporal","Cadence"]

Own and scale the compute fleet for Fluidstack’s AI-grade infrastructure. Build end-to-end health metrics, automate failure repair pipelines, and design the XPU qualification platform for new GPU/TPU generations. Drive fleet reliability and performance across Kubernetes and bare metal at hyperscale, with a focus on low-latency incident response and tooling.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
2 months ago

Production Engineer, Compute

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own and scale the compute fleet for Fluidstack’s AI-grade infrastructure. Build end-to-end health metrics, automate failure repair pipelines, and design the XPU qualification platform for new GPU/TPU generations. Drive fleet reliability and performance across Kubernetes and bare metal at hyperscale, with a focus on low-latency incident response and tooling.
Location: San Francisco, New York, Seattle, Austin
Workplace: Hybrid
Employment Type: Full time
Job Function: Manufacturing & Production Operations

Key Responsibilities

  • •Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU and TPU in production — across Kubernetes-orchestrated workloads and bare metal, at scale.
  • •Turn repair into a pipeline, not a procedure. Build and own the automation that takes a compute failure from detection through triage, parts management, and return to service. No one-off scripts, no heroics.
  • •Design and expand the XPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU and TPU generation. You define what "good" looks like before hardware goes into production.
  • •Own Redfish and BMC tooling. Firmware-level telemetry, log collection at fleet scale, and the low-level access layer that repair automation and health tooling depend on.
  • •Own end-to-end reliability, scalability, and operation of the compute fleet at-scale. Fluidstack is building one of the largest XPU fleets in the world and that can only be accomplished with aggressive automation, tooling, and incident discipline.

Pay and Benefits

Salary: USD 175,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPensionPaid Leave

Key Requirements

  • •Shipped production automation or tooling at scale in a hardware/software environment
  • •Experience designing or operating compute infrastructure (GPU/TPU, Kubernetes, bare metal)
  • •Proficiency with telemetry, metrics, and incident response tooling (e.g., Redfish, BMC, Prometheus, Grafana)
  • •Ability to own complex end-to-end systems and drive projects with minimal supervision
  • •Comfort with ambiguity, rapid iteration, and working across hardware and software teams
Experience:HyperscaleAi hardwareCompute infrastructure
Skills:KubernetesGoPythonPrometheusGrafanaRedfishIPMIGPUTPUTemporalCadence
Languages:English
Tech Stack:KubernetesPrometheusGrafanaRedfishIPMIGPUTPUGoPythonTemporalCadence

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor