Site Reliability Engineer (High Performance Computing)

SpaceX
United States
Workplace: OnsiteFull timeUSD 125,000 - 195,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Ownership","Self-critical","Fairness","Clear communication","Incident readiness"]

Own the end-to-end operating model for a shared HPC compute platform, building reliability through Linux administration, infrastructure-as-code, storage/resource management, and production automation. Drive observability across clusters, nodes, and storage while reducing toil via software tooling. Partner with HPC systems engineers to improve cluster reliability, incidents, and capacity planning so engineers can run mission-critical work with minimal friction.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
20 hours ago

Site Reliability Engineer (High Performance Computing)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Own the end-to-end operating model for a shared HPC compute platform, building reliability through Linux administration, infrastructure-as-code, storage/resource management, and production automation. Drive observability across clusters, nodes, and storage while reducing toil via software tooling. Partner with HPC systems engineers to improve cluster reliability, incidents, and capacity planning so engineers can run mission-critical work with minimal friction.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Participate in an on-call rotation with sustainable incident response and blameless postmortems.
  • •Manage node lifecycle using infrastructure as code, including OS images, firmware, configuration management, and kernel/driver stacks.
  • •Build observability for administrators and end users across cluster, node, and storage health plus job/workflow-level signals.
  • •Reduce toil through automation while balancing operations with building software that makes production work smaller.
  • •Lead capacity planning with users and collaborate with HPC systems engineers to ensure operable, maintainable infrastructure.

Pay and Benefits

Salary: USD 125,000 - 195,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kPaid ParentalDisability InsuranceLife InsurancePaid Leave

Key Requirements

  • •2+ years of professional experience operating production infrastructure (servers, services, or networks) including monitoring, debugging, and repair.
  • •2+ years of experience with Linux operating systems in production.
  • •2+ years of experience with monitoring and debugging what you own, with strong production instincts.
  • •Experience deploying and maintaining infrastructure as code or configuration management (e.g., Ansible, Puppet, Terraform).
  • •Bachelor's degree in computer science, engineering, math, or a scientific discipline, or 2+ years of professional experience operating production infrastructure in lieu of a degree.
Experience:2+ yearsProduction infrastructureHPC
Education:Bachelor's
Skills:OwnershipSelf-criticalFairnessClear communicationIncident readiness
Languages:English
Tech Stack:LinuxInfrastructure as CodePrometheusGrafanaNagiosAnsiblePuppetTerraformPythonDockerPodmanSingularityApptainerKubernetesVASTSlurmPBSLSFPyTorchTensorFlow

Eligibility

Nationality:US National
Security Clearance:TS/SCI with polygraph

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn