Distributed Systems Engineer

FluidStack
San Francisco, New York, Austin, Seattle
Workplace: OnsiteFull timeUSD 175,000 - 300,000 annuallyFunction: IT Operations (Systems/Network Admin)Skills: ["Extreme ownership","Autonomy","First-principles thinking","Velocity","Incident response"]

Own and build the observability platform, including data pipelines, correlation, and health checks that make the compute fleet legible from site down to device and link. Define and deliver stable, versioned infrastructure APIs used across the company, and help build the production control plane for unified machine management and distributed command execution. Own fleet state as the source of truth and ensure new hardware lands cleanly via automated provisioning.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
2 months ago

Distributed Systems Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Own and build the observability platform, including data pipelines, correlation, and health checks that make the compute fleet legible from site down to device and link. Define and deliver stable, versioned infrastructure APIs used across the company, and help build the production control plane for unified machine management and distributed command execution. Own fleet state as the source of truth and ensure new hardware lands cleanly via automated provisioning.
Location: San Francisco, New York, Austin, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Own the observability platform, building and operating data pipelines, correlation, and healthcheck frameworks for site-to-device visibility.
  • •Define and build the API surface for infrastructure, designing contracts between production infrastructure and tools across the company.
  • •Build the production control plane, including unified machine management, state inspection, and distributed command execution supported by Kubernetes-based infrastructure.
  • •Own fleet state as the source of truth, including SLOs, site lifecycle state, and integration with internal and customer-facing operations platforms.
  • •Land new hardware cleanly into the platform via ZTP, DHCP, DNS, and artifact-based workflows through IaaS before production.

Pay and Benefits

Salary: USD 175,000 - 300,000 annually
Perks:RetirementHealth InsuranceDentalVisionPaid Leave

Key Requirements

  • •Treat toil as a bug and automate work that would otherwise be repeated by humans.
  • •Design APIs that age well and avoid leaky abstractions at scale.
  • •Move toward ambiguity by building the map and communicating it clearly to others.
  • •Learn quickly in unfamiliar domains and reach real competence fast.
  • •Be able to carry a pager, run incidents, write postmortems, and fix systemic causes.
Experience:Distributed systemsData pipelinesTime-series observabilityProduction services
Skills:Extreme ownershipAutonomyFirst-principles thinkingVelocityIncident response
Tech Stack:KubernetesLLM APIsMCP serversClaude CodeCursorPrometheusThanosVictoriaMetricsTemporalCadenceBMC/RedfishGoPythonPostgresZTPDHCPDNSArtifactsIaaSDistributed command execution

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor