Production Engineer, IaaS

FluidStack
San Francisco, New York, Austin, Seattle
Workplace: OnsiteFull timeUSD 175,000 - 300,000 annuallyFunction: Software EngineeringSkills: ["Extreme ownership","Velocity","First principles","Incident response","Ability to learn quickly"]

Build and operate the observability and control-plane infrastructure that makes large-scale AI compute fleets legible and manageable. Own production data pipelines, correlation/decoration engines, health checks, and a versioned API surface used across the company. Design unified machine management, distributed command execution, and fleet state “source of truth,” including onboarding new hardware through ZTP, DHCP, DNS, and IaaS.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
2 months ago

Production Engineer, IaaS

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live
Reposted: similar role first listed 3 months ago

Job Summary

Build and operate the observability and control-plane infrastructure that makes large-scale AI compute fleets legible and manageable. Own production data pipelines, correlation/decoration engines, health checks, and a versioned API surface used across the company. Design unified machine management, distributed command execution, and fleet state “source of truth,” including onboarding new hardware through ZTP, DHCP, DNS, and IaaS.
Location: San Francisco, New York, Austin, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own the observability platform, building and operating data pipelines, decoration/correlation, and healthcheck frameworks from site to device and link.
  • •Define and build the API surface for infrastructure by designing contracts between production infrastructure and the tools that operate the hyperscale fleet.
  • •Build the production control plane, including unified machine management, state inspection, and distributed command execution with Kubernetes-based infrastructure.
  • •Own fleet state as the source of truth, covering SLOs, site lifecycle state, and integrations so system reality matches what the platform reports.
  • •Land new hardware into the platform via ZTP, DHCP, DNS, and artifacts, ensuring each new XPU generation and site integration goes through IaaS before production.

Pay and Benefits

Salary: USD 175,000 - 300,000 annually
Perks:RetirementHealth InsuranceDentalVisionPaid Leave

Key Requirements

  • •Treat toil as a bug by automating repeat manual work and building solutions that remove human effort.
  • •Design APIs that age well and avoid leaky abstractions at scale by learning from production pain points.
  • •Comfortable moving toward ambiguity—build the map in unclear domains and explain it to others.
  • •Learn quickly in unfamiliar domains and reach real competence fast.
  • •Experience shipping production services other teams depend on at scale, and familiarity with AI tooling (LLM APIs, MCP servers, agentic frameworks).
Experience:AI infrastructureDistributed systemsData pipeline engineering
Skills:Extreme ownershipVelocityFirst principlesIncident responseAbility to learn quickly
Tech Stack:KubernetesPrometheusThanosVictoriaMetricsTemporalCadenceBMCRedfishGoPythonPostgresLLM APIsMCP serversAgentic frameworksClaude CodeCursorZTPDHCPDNSIaaS

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor