Reliability Engineer, R&D

FluidStack
Austin, New York, San Francisco, Seattle
Workplace: OnsiteFull timeUSD 203,000 - 232,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Extreme ownership","First principles thinking","Velocity","Cross-discipline collaboration","Data-driven decision making"]

Own reliability engineering for the reference design, modeling availability, identifying weak points, and engineering fixes before deployment. Build RAM models quantifying failure rates, redundancy, and maintainability by configuration. Run cross-discipline FMEAs to catalog failure modes and design them out with partner engineering teams. Close the loop with fleet field failures by feeding real-world data back into models and driving design changes for civilization-scale AI infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FluidStack
FluidStack
1 month ago

Reliability Engineer, R&D

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Own reliability engineering for the reference design, modeling availability, identifying weak points, and engineering fixes before deployment. Build RAM models quantifying failure rates, redundancy, and maintainability by configuration. Run cross-discipline FMEAs to catalog failure modes and design them out with partner engineering teams. Close the loop with fleet field failures by feeding real-world data back into models and driving design changes for civilization-scale AI infrastructure.
Location: Austin, New York, San Francisco, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Own reliability engineering for the reference design, including availability modeling, weak-point identification, and pre-deployment fixes.
  • •Build RAM models to quantify failure rates, redundancy, and maintainability per configuration.
  • •Run FMEAs across disciplines to catalog failure modes and design them out with engineering teams.
  • •Close the loop with the fleet by feeding field failures back into models and design changes.

Pay and Benefits

Salary: USD 203,000 - 232,000 annually

Key Requirements

  • •Have done reliability engineering for infrastructure, energy, or complex hardware.
  • •Have built availability models used by decision-makers.
  • •Have led cross-discipline FMEAs that changed designs.
  • •Mine field data to validate failure rates.
  • •Argue redundancy trade-offs in dollars and nines.
Experience:InfrastructureEnergyComplex hardwareData centersAI compute infrastructure
Skills:Extreme ownershipFirst principles thinkingVelocityCross-discipline collaborationData-driven decision making

Company Brief

FluidStack
Builds and deploys large-scale GPU cloud infrastructure for AI labs, enterprises and governments, providing high-performance AI training and inference capacity and rapid data-center deployment services.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York, United States
Founded: 2017
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor