Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
Washington, California, Redmond
Workplace: OnsiteFull timeUSD 125,000 - 200,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 1+ yearsEducation: bachelorsSkills: ["Communication","Customer-facing collaboration","Problem-solving","Continuous improvement"]

Design, operate, and scale on-premise GPU/AI infrastructure supporting critical national security missions. Own GPU/CPU deployment and customer-facing GPU-as-a-service on bare metal and virtualized platforms, productize AI cluster solutions at 100k+ scale, and automate deployments using Kubernetes, Linux, and infrastructure tooling. Build reliable services with monitoring, alerting, and continuous improvement in collaboration with AI engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
2 days ago

Site Reliability Engineer, AI Infrastructure (Starshield)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Design, operate, and scale on-premise GPU/AI infrastructure supporting critical national security missions. Own GPU/CPU deployment and customer-facing GPU-as-a-service on bare metal and virtualized platforms, productize AI cluster solutions at 100k+ scale, and automate deployments using Kubernetes, Linux, and infrastructure tooling. Build reliable services with monitoring, alerting, and continuous improvement in collaboration with AI engineers.
Location: Washington, California, Redmond
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Manage GPU/CPU infrastructure deployments to Top Secret datacenters.
  • •Manage and support GPU as a service for external customers on bare metal hardware and virtualized platforms.
  • •Design, validate, and productize solutions for AI clusters (100k+ GPU scale).
  • •Develop automation to deploy and manage on-premise Kubernetes/AI clusters and operating systems.
  • •Deploy and manage core infrastructure (databases, monitoring, distributed storage) with high availability monitoring and alerting.
Travel: Medium travel

Pay and Benefits

Salary: USD 125,000 - 200,000 annually
Perks:Health InsuranceVisionDental401kPaid ParentalLife InsurancePaid LeavePaid Sick

Key Requirements

  • •1+ years of professional experience in site reliability engineering or DevOps (or 3+ years in lieu of a degree)
  • •1+ years of professional experience with Linux operating systems
  • •Experience with Terraform, Ansible, or other infrastructure tools
  • •Experience with containerization technologies, including OCI containers and Kubernetes
  • •Development experience in Bash, Python, and/or similar languages, plus experience in Python, C++, or Go
Experience:1+ years
Education:Bachelor's
Skills:CommunicationCustomer-facing collaborationProblem-solvingContinuous improvement
Tech Stack:LinuxTerraformAnsibleOCI containersKubernetesBashPythonC++GoBazelMakefilesNVIDIA GPUBlackwellRubinDatabasesMonitoringDistributed storageTCP/IP

Eligibility

Nationality:US National
Security Clearance:Top SecretTop Secret SCIDOE Level Q

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn