Site Reliability Engineer, AI Infrastructure (Starshield)
Palo Alto, Redmond, Washington
Workplace: OnsiteFull timeUSD 125,000 - 195,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 1+ yearsSkills: ["Monitoring","Alerting","High availability","Automation","Communications"]Design, operate, and scale AI infrastructure supporting Starshield’s national security missions. Manage GPU/CPU deployments to classified data centers, provide GPU-as-a-service on bare metal and virtualized platforms, and build automation for on-prem Kubernetes/AI clusters, OSs, databases, monitoring, and distributed storage. Collaborate with AI engineers to productize solutions at 100k+ GPU scale, ensuring high availability through monitoring, alerting, and continuous lifecycle improvements.
Loading
Loading job details...
Preparing the role view and application actions.

