Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
Washington, California, Redmond, Palo Alto
Workplace: OnsiteFull timeUSD 165,000 - 265,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communications","Mentorship","Technical leadership","Problem-solving"]

Design, operate, and scale AI infrastructure supporting Starshield’s national security missions. Manage GPU/CPU deployments to secure data centers, provide GPU-as-a-service for external customers, and productize AI cluster solutions at 100k+ GPU scale. Build automation for on-prem Kubernetes/AI clusters, core infrastructure (databases, monitoring, distributed storage), and high-availability systems—while collaborating with AI engineers and mentoring others.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
20 hours ago

Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Design, operate, and scale AI infrastructure supporting Starshield’s national security missions. Manage GPU/CPU deployments to secure data centers, provide GPU-as-a-service for external customers, and productize AI cluster solutions at 100k+ GPU scale. Build automation for on-prem Kubernetes/AI clusters, core infrastructure (databases, monitoring, distributed storage), and high-availability systems—while collaborating with AI engineers and mentoring others.
Location: Washington, California, Redmond, Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • •Provide GPU-as-a-service support for external customers on bare metal and virtualized platforms.
  • •Design, validate, and productize solutions for AI clusters at 100k+ GPU scale.
  • •Build automation to deploy and manage on-prem Kubernetes/AI clusters, operating systems, and core infrastructure (databases, monitoring, distributed storage).
  • •Monitor/alert for high availability, improve system performance, mentor junior engineers, and lead technical excellence.
Travel: High travel

Pay and Benefits

Salary: USD 165,000 - 265,000 annually
Equity and Bonus:Equity
Perks:MedicalVisionDental401kLife InsuranceLong-term DisabilityPaid ParentalPaid Leave

Key Requirements

  • •Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 5+ years of professional experience with Linux, or 7+ years in software/DevOps/SRE in lieu of a degree.
  • •5+ years of experience with Kubernetes.
  • •5+ years of experience managing Linux operating systems.
  • •Experience with Terraform, Ansible, or other infrastructure tools.
  • •Development experience in Python, C++, or Go, plus scripting in Bash or Python.
Education:Bachelor's
Skills:CommunicationsMentorshipTechnical leadershipProblem-solving
Tech Stack:LinuxKubernetesTerraformAnsibleOCI containersBashPythonC++GoKubernetes/AI clustersDatabasesMonitoringDistributed storageBazelMakefilesTCP/IPNVIDIA GPUBlackwellRubin

Eligibility

Nationality:US National
Work Authorization:Authorization required. Sponsorship not provided.
Security Clearance:Top SecretTop Secret SCIDOE Level Q

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn