Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
Washington, California, Redmond
Workplace: OnsiteFull timeUSD 125,000 - 195,000Function: DevOps, Cloud & InfrastructureExperience: 1+ yearsEducation: bachelorsSkills: ["Communication"]

Design, operate, and scale on-premise infrastructure that powers Starshield’s AI and GPU workloads for critical national security missions. Build automation for deployments and manage Kubernetes/AI clusters, GPU-as-a-service on bare metal and virtualized platforms, and core services like databases, monitoring, and distributed storage. Partner with AI engineers to productize highly scalable, maintainable systems with strong availability, lifecycle ownership, and continuous monitoring.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
2 days ago

Site Reliability Engineer, AI Infrastructure (Starshield)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Design, operate, and scale on-premise infrastructure that powers Starshield’s AI and GPU workloads for critical national security missions. Build automation for deployments and manage Kubernetes/AI clusters, GPU-as-a-service on bare metal and virtualized platforms, and core services like databases, monitoring, and distributed storage. Partner with AI engineers to productize highly scalable, maintainable systems with strong availability, lifecycle ownership, and continuous monitoring.
Location: Washington, California, Redmond
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Manage GPU/CPU infrastructure deployments to Top Secret datacenters.
  • •Provide GPU-as-a-service support for external customers on bare metal and virtualized platforms.
  • •Design, validate, and productize solutions for AI clusters at 100k+ GPU scale.
  • •Develop automation to deploy and manage on-premise Kubernetes/AI clusters and operating systems.
  • •Deploy and manage core infrastructure including databases, monitoring, and distributed storage while improving service availability across the lifecycle.
Travel: Medium travel

Pay and Benefits

Salary: USD 125,000 - 195,000
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kPaid ParentalLife InsurancePaid LeaveLong-term Disability

Key Requirements

  • •Bachelor’s degree (computer science/information systems/IT/engineering) with 1+ years of experience in site reliability engineering or DevOps, or 3+ years in SRE/DevOps in lieu of a degree.
  • •1+ years of professional experience with Linux operating systems.
  • •Experience with Terraform, Ansible, or other infrastructure tools.
  • •Experience with containerization technologies including OCI containers and Kubernetes.
  • •Experience scripting in Bash or Python, plus development experience in Python, C++, or Go.
Experience:1+ yearsNational securitySatellite constellationAI infrastructureDevOpsSite reliability engineeringOn-premise infrastructure
Education:Bachelor's
Skills:Communication
Tech Stack:LinuxTerraformAnsibleKubernetesOCI containersBashPythonC++GoBazelMakefilesTCP/IPNVIDIA GPU deployment stacksBlackwellRubinDistributed storageMonitoringDatabases

Eligibility

Security Clearance:Top SecretTop Secret SCIDOE Level Q

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn