Software Engineer, GPU Infrastructure

FluidStack
San Francisco, New York, Austin, Seattle
Workplace: OnsiteFull timeUSD 175,000 - 300,000 annuallyFunction: Software EngineeringSkills: ["Extreme ownership","First-principles thinking","Velocity","Incident ownership","Learning mindset"]

Build and operate a GPU infrastructure platform that keeps hyperscale AI compute reliable and scalable. Own the compute fleet’s health end to end with metrics pipelines, alerting, and unified observability across Kubernetes-orchestrated workloads and bare metal. Turn deployment and repair into automated pipelines, expand GPU qualification and burn-in workflows, and build low-level Redfish/BMC tooling to support fleet-scale incidents and reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
2 months ago

Software Engineer, GPU Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and operate a GPU infrastructure platform that keeps hyperscale AI compute reliable and scalable. Own the compute fleet’s health end to end with metrics pipelines, alerting, and unified observability across Kubernetes-orchestrated workloads and bare metal. Turn deployment and repair into automated pipelines, expand GPU qualification and burn-in workflows, and build low-level Redfish/BMC tooling to support fleet-scale incidents and reliability.
Location: San Francisco, New York, Austin, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own compute fleet health end to end by building metrics pipelines, alerting, and a unified health view for every GPU in production.
  • •Automate compute deployment and repair as pipelines, handling detection through triage and return to service without one-off scripts.
  • •Design and expand the GPU qualification platform including burn-in, performance baselining, and NPI execution.
  • •Build and maintain Redfish and BMC tooling for firmware-level telemetry and fleet-scale log collection.
  • •Drive end-to-end reliability, scalability, and operational excellence for a large at-scale GPU compute fleet.

Pay and Benefits

Salary: USD 175,000 - 300,000 annually
Equity and Bonus:Equity
Perks:RetirementHealth InsuranceDentalVisionPaid Leave

Key Requirements

  • •Treat toil as a bug and eliminate manual steps in repair workflows.
  • •Reason about hardware failure modes at the firmware and silicon level.
  • •Work effectively in ambiguity by building clarity and explaining it to others.
  • •Learn rapidly in unfamiliar domains and reach competence quickly.
  • •Carry a pager, run incidents, write postmortems, and fix systemic root causes.
Experience:AIInfrastructureHardwareAutomation
Skills:Extreme ownershipFirst-principles thinkingVelocityIncident ownershipLearning mindset
Tech Stack:KubernetesBare metalRedfishBMCIPMILLM APIsMCP serversAgentic frameworksClaude CodeCursorTemporalCadencePrometheusGrafanaGoPython

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor