Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
California, Redmond, Washington
Workplace: OnsiteFull timeUSD 165,000 - 265,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communications","Mentorship"]

Design, operate, and scale on-prem infrastructure that powers AI clusters for Starshield’s national security missions. You’ll manage GPU/CPU deployments and provide GPU-as-a-service for customers, build automation for Kubernetes and operating systems, and own core platform components like databases, monitoring, and distributed storage. Partner closely with AI engineers to deliver highly scalable, operable services with strong availability, and mentor others while maintaining Top Secret clearance requirements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
2 days ago

Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Design, operate, and scale on-prem infrastructure that powers AI clusters for Starshield’s national security missions. You’ll manage GPU/CPU deployments and provide GPU-as-a-service for customers, build automation for Kubernetes and operating systems, and own core platform components like databases, monitoring, and distributed storage. Partner closely with AI engineers to deliver highly scalable, operable services with strong availability, and mentor others while maintaining Top Secret clearance requirements.
Location: California, Redmond, Washington
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • •Provide and support GPU as a service for external customers on bare metal and virtualized platforms.
  • •Design, validate, and productize solutions for AI clusters at 100k+ GPU scale.
  • •Build automation to deploy and manage on-prem Kubernetes/AI clusters and operating systems, including core infrastructure like databases, monitoring, and distributed storage.
  • •Lead technical excellence by mentoring junior engineers and improving the full service lifecycle from design through deployment and refinement, including monitoring/alerting for high availability.
Travel: High travel

Pay and Benefits

Salary: USD 165,000 - 265,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kLife InsurancePaid ParentalPaid Leave

Key Requirements

  • •Bachelor’s degree in computer science/information systems/IT or engineering, plus 5+ years with Linux OR 7+ years in software, DevOps, or site reliability engineering in lieu of a degree.
  • •5+ years of experience with Kubernetes.
  • •5+ years managing Linux operating systems.
  • •Experience with Terraform, Ansible, or similar infrastructure tools.
  • •Development and scripting experience with Bash, Python (and/or similar), plus development experience in Python, C++, or Go.
Education:Bachelor's
Skills:CommunicationsMentorship
Languages:English
Tech Stack:LinuxKubernetesTerraformAnsibleOCI containersBashPythonC++GoBazelMakefilesTCP/IPDistributed databasesNVIDIA GPU deployment stacksBlackwellRubinOn-premise computeDatabasesDistributed storageMonitoring

Eligibility

Nationality:US National
Security Clearance:Top SecretTop Secret SCIDOE Level Q

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn