Production Engineer, Compute Team Lead

FluidStack
San Francisco
Workplace: OnsiteFull timeUSD 269,000 - 335,000 annuallyFunction: Manufacturing & Production OperationsSkills: ["Extreme ownership","Autonomy","Velocity","First principles thinking","Leadership"]

Lead compute production engineering to keep tens of thousands of GPUs serving customers. Own compute fleet availability by defining SLOs, building tooling, and driving performance. Build automation for the node lifecycle—provisioning, health checks, remediation, and return to service—aiming for zero-touch operations. Set the on-call and escalation model to maintain sharp response while preventing team burnout, at AI infrastructure scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
1 month ago

Production Engineer, Compute Team Lead

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Lead compute production engineering to keep tens of thousands of GPUs serving customers. Own compute fleet availability by defining SLOs, building tooling, and driving performance. Build automation for the node lifecycle—provisioning, health checks, remediation, and return to service—aiming for zero-touch operations. Set the on-call and escalation model to maintain sharp response while preventing team burnout, at AI infrastructure scale.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Manager level

Key Responsibilities

  • •Lead the compute production engineering team operating tens of thousands of GPUs serving customers.
  • •Own compute fleet availability by defining SLOs, building tooling, and driving the availability number.
  • •Build automation for the node lifecycle: provisioning, health checks, remediation, and return to service without human touch.
  • •Set the on-call and escalation model to keep response sharp without burning out the team.

Pay and Benefits

Salary: USD 269,000 - 335,000 annually

Key Requirements

  • •Led SRE or production engineering teams running large fleets.
  • •Moved an availability metric and can explain how it was improved.
  • •Shipped automated remediation that retired a runbook.
  • •Hired and grown strong engineers who can speak to the team’s impact.
  • •Experience with GPU or HPC fleets and customer-facing reliability (bonus).
Experience:GPUHPCSREProduction engineeringCustomer-facing reliability
Skills:Extreme ownershipAutonomyVelocityFirst principles thinkingLeadership
Tech Stack:SREGPUHPCKubernetesSlurm

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor