Sr. Site Reliability Engineer, AI Infrastructure (Starshield)
Washington, California, Redmond
Full timeUSD 165,000 - 270,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Communications","Collaboration","Mentoring","Training","Technical leadership"]Design, operate, and scale AI and GPU infrastructure supporting critical Starshield national security missions. Own on-prem deployments for GPU/CPU clusters, Kubernetes/AI environments, and core infrastructure like databases, monitoring, and distributed storage. Build automation with tools such as Terraform/Ansible, ensure high availability through monitoring and alerting, and collaborate with AI engineers to deliver reliable, maintainable systems at 100k+ GPU scale.
Loading
Loading job details...
Preparing the role view and application actions.

