Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
Washington, California, Redmond
Full timeUSD 165,000 - 265,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Mentorship","Technical leadership","Problem-solving"]

Design, operate, and scale Starshield’s AI infrastructure supporting critical national security missions. Manage on-prem GPU/CPU deployments to classified data centers, provide GPU-as-a-service, and build highly scalable AI clusters at 100k+ GPU scale. Develop automation for Kubernetes/AI clusters, operating systems, and core services like databases, monitoring, and distributed storage. Lead technical excellence, mentor engineers, and ensure high availability through monitoring, alerting, and lifecycle improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
2 days ago

Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Design, operate, and scale Starshield’s AI infrastructure supporting critical national security missions. Manage on-prem GPU/CPU deployments to classified data centers, provide GPU-as-a-service, and build highly scalable AI clusters at 100k+ GPU scale. Develop automation for Kubernetes/AI clusters, operating systems, and core services like databases, monitoring, and distributed storage. Lead technical excellence, mentor engineers, and ensure high availability through monitoring, alerting, and lifecycle improvements.
Location: Washington, California, Redmond
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • •Provide GPU as a service for external customers on bare metal and virtualized platforms.
  • •Design, validate, and productize solutions for AI clusters at 100k+ GPU scale.
  • •Develop automation for deploying and managing on-prem Kubernetes/AI clusters, operating systems, and core infrastructure (databases, monitoring, distributed storage).
  • •Mentor junior engineers and lead the team to technical excellence across service lifecycle and high-availability improvements.
Travel: High travel

Pay and Benefits

Salary: USD 165,000 - 265,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kPaid ParentalLife InsuranceLong-term DisabilityPaid Leave

Key Requirements

  • •Bachelor’s degree in CS/IT/engineering plus 5+ years with Linux OS; or 7+ years in software, DevOps, or site reliability engineering in lieu of a degree.
  • •5+ years of experience with Kubernetes and managing Linux operating systems.
  • •Experience with Terraform, Ansible, or other infrastructure tools for automation.
  • •Experience with containerization technologies (OCI containers, Kubernetes) and scripting in Bash, Python, or similar languages.
  • •Development experience in Python, C++, or Go.
Experience:5+ yearsNational securitySatellite infrastructureOn-prem infrastructureKubernetesAI clustersDevOpsSite reliability engineering
Education:Bachelor's
Skills:CommunicationMentorshipTechnical leadershipProblem-solving
Tech Stack:LinuxKubernetesTerraformAnsibleBashPythonC++GoOCI containersContainerizationDatabasesMonitoringDistributed storageBazelMakefilesTCP/IPCloud virtualizationNVIDIA GPU deployment stacksBlackwellRubin

Eligibility

Security Clearance:Top SecretTop Secret SCIDOE Level Q

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn