Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
Washington, California, Redmond
Full timeUSD 165,000 - 270,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Communications","Collaboration","Mentoring","Training","Technical leadership"]

Design, operate, and scale AI and GPU infrastructure supporting critical Starshield national security missions. Own on-prem deployments for GPU/CPU clusters, Kubernetes/AI environments, and core infrastructure like databases, monitoring, and distributed storage. Build automation with tools such as Terraform/Ansible, ensure high availability through monitoring and alerting, and collaborate with AI engineers to deliver reliable, maintainable systems at 100k+ GPU scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
2 days ago

Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Design, operate, and scale AI and GPU infrastructure supporting critical Starshield national security missions. Own on-prem deployments for GPU/CPU clusters, Kubernetes/AI environments, and core infrastructure like databases, monitoring, and distributed storage. Build automation with tools such as Terraform/Ansible, ensure high availability through monitoring and alerting, and collaborate with AI engineers to deliver reliable, maintainable systems at 100k+ GPU scale.
Location: Washington, California, Redmond
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • •Manage and support GPU-as-a-service for external customers on bare metal hardware and virtualized platforms.
  • •Design, validate, and productize solutions for AI clusters at 100k+ GPU scale.
  • •Develop automation to deploy and manage on-premise Kubernetes/AI clusters and operating systems.
  • •Closely collaborate with AI engineers across the full service lifecycle, ensuring monitoring/alerting and high availability while mentoring and training junior engineers.
Travel: Extensive travel

Pay and Benefits

Salary: USD 165,000 - 270,000 annually
Equity and Bonus:Equity
Perks:MedicalVisionDental401kDisability InsuranceLife InsurancePaid ParentalPaid LeavePaid HolidaysPaid Sick

Key Requirements

  • •Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 5+ years with Linux operating systems, or 7+ years in software, DevOps, or site reliability engineering in lieu of a degree.
  • •5+ years of experience with Kubernetes.
  • •5+ years managing Linux operating systems.
  • •Experience with Terraform, Ansible, or other infrastructure tools, plus containerization technologies such as OCI containers and Kubernetes.
  • •Development experience in Python, C++, or Go, including scripting in Bash or Python.
  • •Active Top Secret, Top Secret SCI, or DOE Level Q clearance.
  • •Knowledge of monitoring/alerting and high-availability systems, plus strong networking knowledge of TCP/IP.
Experience:5+ years
Education:Bachelor's
Skills:CommunicationsCollaborationMentoringTrainingTechnical leadership
Tech Stack:LinuxKubernetesTerraformAnsibleOCI containersBashPythonC++GoBazelMakefilesTCP/IPDistributed databasesDistributed storageNVIDIA GPUBlackwellRubinOn-premise computeKubernetes AI clusters

Eligibility

Nationality:US National
Work Authorization:Authorization required. Sponsorship not provided.
Security Clearance:Top SecretTop Secret SCIDOE Level Q

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn