Site Reliability Engineer, AI Infrastructure (Starshield)
California, Redmond, Washington
Workplace: OnsiteFull timeUSD 125,000 - 195,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 1+ yearsEducation: bachelorsSkills: ["Communication"]Design, operate, and scale AI infrastructure that powers Starshield’s national security missions. Manage GPU/CPU deployments to classified data centers, provide GPU-as-a-service on bare metal and virtualized platforms, and productize solutions for large AI clusters. Build automation for on-prem Kubernetes/AI clusters and OS deployments, own core infrastructure (databases, monitoring, distributed storage), and partner with AI engineers to improve service lifecycle reliability and high availability.
Loading
Loading job details...
Preparing the role view and application actions.

