Site Reliability Engineer — HPC & Automation (Silicon Engineering)

SpaceX
Redmond, Seattle
Workplace: OnsiteFull timeUSD 125,000 - 175,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Proactive","Intellectually curious","Strong communication","Problem-solving","Ability to learn"]

Design, operate, and scale SpaceX’s high-performance computing infrastructure used to develop Starlink silicon. As a Site Reliability Engineer on the Silicon Engineering team, you’ll deploy and maintain HPC clusters, automate full simulation workflows, and manage infrastructure-as-code with modern observability. You’ll also run CI/CD pipelines and build/release systems, eliminate performance bottlenecks through measurement, and accelerate regression and simulation turnaround times for chip teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
2 months ago

Site Reliability Engineer — HPC & Automation (Silicon Engineering)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Design, operate, and scale SpaceX’s high-performance computing infrastructure used to develop Starlink silicon. As a Site Reliability Engineer on the Silicon Engineering team, you’ll deploy and maintain HPC clusters, automate full simulation workflows, and manage infrastructure-as-code with modern observability. You’ll also run CI/CD pipelines and build/release systems, eliminate performance bottlenecks through measurement, and accelerate regression and simulation turnaround times for chip teams.
Location: Redmond, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Deploy, upgrade, operate, maintain, and scale a suite of clusters and services.
  • •Collaborate to build automated, turnkey silicon simulation workflows that speed up project timelines.
  • •Manage infrastructure as code and use observability tools to provide full visibility into cluster and infrastructure health.
  • •Operate continuous integration pipelines, build and release systems, and version control across the environment.
  • •Identify and eliminate performance bottlenecks using measurement and engineering techniques.

Pay and Benefits

Salary: USD 125,000 - 175,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kLife InsurancePaid ParentalPaid LeaveSick TimeEquity

Key Requirements

  • •Bachelor’s degree in computer science, information systems, or an engineering discipline, or 2+ years of professional experience in system administration, high performance computing, or site reliability engineering.
  • •1+ years of development experience with Bash, Python, and/or other programming languages.
  • •1+ years of experience with Linux operating systems.
  • •Ability to deploy, upgrade, operate, maintain, and scale clusters and services.
  • •Ability to work extended hours and weekends as needed to meet critical milestones.
Experience:2+ yearsHigh performance computingSystem administrationSite reliability engineering
Education:Bachelor's
Skills:ProactiveIntellectually curiousStrong communicationProblem-solvingAbility to learn
Languages:English
Tech Stack:BashPythonLinuxDockerKubernetesMySQLPostgreSQLSQLiteTCP/IPSlurmLSFTerraformAnsiblePuppetGrafanaPrometheusJenkinsBambooREST APINetApp ONTAP

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn