Sr. Site Reliability Engineer, AI Infrastructure (Starshield)
Washington, California, Redmond
Full timeUSD 165,000 - 265,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Mentorship","Technical leadership","Problem-solving"]Design, operate, and scale Starshield’s AI infrastructure supporting critical national security missions. Manage on-prem GPU/CPU deployments to classified data centers, provide GPU-as-a-service, and build highly scalable AI clusters at 100k+ GPU scale. Develop automation for Kubernetes/AI clusters, operating systems, and core services like databases, monitoring, and distributed storage. Lead technical excellence, mentor engineers, and ensure high availability through monitoring, alerting, and lifecycle improvements.
Loading
Loading job details...
Preparing the role view and application actions.

