Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
California, Redmond, Washington
Workplace: OnsiteFull timeUSD 125,000 - 195,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 1+ yearsEducation: bachelorsSkills: ["Communication"]

Design, operate, and scale AI infrastructure that powers Starshield’s national security missions. Manage GPU/CPU deployments to classified data centers, provide GPU-as-a-service on bare metal and virtualized platforms, and productize solutions for large AI clusters. Build automation for on-prem Kubernetes/AI clusters and OS deployments, own core infrastructure (databases, monitoring, distributed storage), and partner with AI engineers to improve service lifecycle reliability and high availability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
2 days ago

Site Reliability Engineer, AI Infrastructure (Starshield)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Design, operate, and scale AI infrastructure that powers Starshield’s national security missions. Manage GPU/CPU deployments to classified data centers, provide GPU-as-a-service on bare metal and virtualized platforms, and productize solutions for large AI clusters. Build automation for on-prem Kubernetes/AI clusters and OS deployments, own core infrastructure (databases, monitoring, distributed storage), and partner with AI engineers to improve service lifecycle reliability and high availability.
Location: California, Redmond, Washington
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Entry level

Key Responsibilities

  • •Manage GPU/CPU infrastructure deployments to Top Secret data centers and support GPU as a service for external customers across bare metal and virtualized platforms.
  • •Design, validate, and productize solutions for AI clusters at 100k+ GPU scale.
  • •Develop automation to deploy and manage on-premise Kubernetes/AI clusters and operating systems.
  • •Deploy and operate core infrastructure including databases, monitoring, and distributed storage for high availability.
  • •Collaborate with AI engineers to improve the full service lifecycle from inception and design through deployment, operation, and refinement; implement monitoring and alerting.
Travel: Extensive travel

Pay and Benefits

Salary: USD 125,000 - 195,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kParental LeaveDisabilityLife InsurancePaid Leave

Key Requirements

  • •Bachelor’s degree in computer science, information systems/IT, or an engineering discipline plus 1+ years of site reliability engineering or DevOps experience, or 3+ years in SRE/DevOps in lieu of a degree.
  • •1+ years of professional experience with Linux operating systems.
  • •Experience with Terraform, Ansible, or other infrastructure tools.
  • •Experience with containerization technologies including OCI containers and Kubernetes.
  • •Experience scripting in Bash, Python, or similar languages and development experience in Python, C++, or Go.
Experience:1+ yearsAI infrastructureNational securitySatellite
Education:Bachelor's
Skills:Communication
Tech Stack:LinuxTerraformAnsibleKubernetesOCIBashPythonC++GoContinuous integrationMonitoringDatabasesDistributed storageBazelMakefilesTCP/IPNVIDIA GPUBlackwellRubinKubernetes clusters

Eligibility

Nationality:US National
Security Clearance:Top SecretTop Secret SCIDOE Level Q

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn