Sr. Site Reliability Engineer, AI Infrastructure (Starshield)
Washington, California, Redmond, Palo Alto
Workplace: OnsiteFull timeUSD 165,000 - 265,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communications","Mentorship","Technical leadership","Problem-solving"]Design, operate, and scale AI infrastructure supporting Starshield’s national security missions. Manage GPU/CPU deployments to secure data centers, provide GPU-as-a-service for external customers, and productize AI cluster solutions at 100k+ GPU scale. Build automation for on-prem Kubernetes/AI clusters, core infrastructure (databases, monitoring, distributed storage), and high-availability systems—while collaborating with AI engineers and mentoring others.
Loading
Loading job details...
Preparing the role view and application actions.

