Software Engineer, Compute (GPU)

FluidStack
San Francisco, New York, Austin, Seattle
Workplace: OnsiteFull timeUSD 175,000 - 300,000 annuallyFunction: Software EngineeringSkills: ["Extreme ownership","Autonomy","Velocity","First-principles thinking","Incident management"]

Build and automate GPU compute fleet operations end to end for hyperscale AI infrastructure. Own compute fleet health with metrics pipelines, alerting, and unified observability across Kubernetes-orchestrated workloads and bare metal. Turn deployment and repairs into reliable pipelines, design the GPU qualification/burn-in platform, and build Redfish/BMC tooling to support firmware-level telemetry. Drive incident discipline and scalable reliability as the fleet grows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
2 months ago

Software Engineer, Compute (GPU)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and automate GPU compute fleet operations end to end for hyperscale AI infrastructure. Own compute fleet health with metrics pipelines, alerting, and unified observability across Kubernetes-orchestrated workloads and bare metal. Turn deployment and repairs into reliable pipelines, design the GPU qualification/burn-in platform, and build Redfish/BMC tooling to support firmware-level telemetry. Drive incident discipline and scalable reliability as the fleet grows.
Location: San Francisco, New York, Austin, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own compute fleet health end to end, including metrics pipelines, alerting, and unified health views.
  • •Turn deployment and repair into pipelines (not one-off scripts), from detection through triage and return to service.
  • •Design and expand the GPU qualification platform, including burn-in, performance baselining, and NPI execution.
  • •Own Redfish and BMC tooling for firmware-level telemetry and fleet-scale log collection.
  • •Drive end-to-end reliability, scalability, and operational discipline for the at-scale compute fleet.

Pay and Benefits

Salary: USD 175,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRetirementPaid Leave

Key Requirements

  • •Treat toil as a bug and replace manual steps in repair workflows with automation.
  • •Reason about hardware failure modes at the firmware and silicon level.
  • •Move toward ambiguity by building a map and explaining it clearly to others.
  • •Learn quickly in unfamiliar domains and reach real competence fast.
  • •Run incidents reliably: carry a pager, lead response, write postmortems, and fix systemic causes.
Experience:AIInfrastructureProduction automationHardwareKubernetes
Skills:Extreme ownershipAutonomyVelocityFirst-principles thinkingIncident management
Tech Stack:KubernetesRedfishBMCIPMILLM APIsMCP serversAgentic frameworksClaude CodeCursorTemporalCadencePrometheusGrafanaGoPythonRMA

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor