Site Reliability Engineer, AI Infrastructure (Starshield)
Washington, California, Redmond
Workplace: OnsiteFull timeUSD 125,000 - 200,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 1+ yearsEducation: bachelorsSkills: ["Communication","Customer-facing collaboration","Problem-solving","Continuous improvement"]Design, operate, and scale on-premise GPU/AI infrastructure supporting critical national security missions. Own GPU/CPU deployment and customer-facing GPU-as-a-service on bare metal and virtualized platforms, productize AI cluster solutions at 100k+ scale, and automate deployments using Kubernetes, Linux, and infrastructure tooling. Build reliable services with monitoring, alerting, and continuous improvement in collaboration with AI engineers.
Loading
Loading job details...
Preparing the role view and application actions.

