Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
Palo Alto, Redmond, Washington
Workplace: OnsiteFull timeUSD 125,000 - 195,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 1+ yearsSkills: ["Monitoring","Alerting","High availability","Automation","Communications"]

Design, operate, and scale AI infrastructure supporting Starshield’s national security missions. Manage GPU/CPU deployments to classified data centers, provide GPU-as-a-service on bare metal and virtualized platforms, and build automation for on-prem Kubernetes/AI clusters, OSs, databases, monitoring, and distributed storage. Collaborate with AI engineers to productize solutions at 100k+ GPU scale, ensuring high availability through monitoring, alerting, and continuous lifecycle improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
20 hours ago

Site Reliability Engineer, AI Infrastructure (Starshield)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Design, operate, and scale AI infrastructure supporting Starshield’s national security missions. Manage GPU/CPU deployments to classified data centers, provide GPU-as-a-service on bare metal and virtualized platforms, and build automation for on-prem Kubernetes/AI clusters, OSs, databases, monitoring, and distributed storage. Collaborate with AI engineers to productize solutions at 100k+ GPU scale, ensuring high availability through monitoring, alerting, and continuous lifecycle improvements.
Location: Palo Alto, Redmond, Washington
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Entry level

Key Responsibilities

  • •Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • •Manage and support GPU as a service for external customers on bare metal and virtualized platforms.
  • •Design, validate, and productize AI cluster solutions at 100k+ GPU scale.
  • •Develop automation to deploy and manage on-prem Kubernetes/AI clusters and operating systems.
  • •Deploy and manage core infrastructure (databases, monitoring, distributed storage) and improve service lifecycle for high system availability.
Travel: Medium travel

Pay and Benefits

Salary: USD 125,000 - 195,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kDisability InsuranceLife InsuranceParental LeavePaid Leave

Key Requirements

  • •Bachelor’s degree in computer science/information systems/IT or engineering discipline and 1+ years of site reliability engineering or DevOps experience, or 3+ years of experience in lieu of a degree.
  • •1+ years of professional experience with Linux operating systems.
  • •Experience with Terraform, Ansible, or other infrastructure tools.
  • •Experience with containerization technologies such as OCI containers and Kubernetes.
  • •Development experience with Python, C++, or Go.
Experience:1+ years
Skills:MonitoringAlertingHigh availabilityAutomationCommunications
Tech Stack:LinuxTerraformAnsibleOCI containersKubernetesBashPythonC++GoBazelMakefilesTCP/IPNVIDIA GPUBlackwellRubinDistributed databasesDistributed storageMonitoring

Eligibility

Security Clearance:Top SecretTop Secret SCIDOE Level Q

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn