AI Research Infrastructure Engineer
AMD
Austin
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Leadership","Communication","Collaboration","Problem-solving","Reliability mindset"]Operate and continuously improve AMD’s shared GPU and HPC compute platform that powers AI/ML and HPC research. Own day-to-day health of SLURM-managed GPU and HPC clusters, partner with researchers to run and optimize large multi-GPU/multi-node workloads, and enhance shared compute services such as storage, networking, containers, and monitoring. Support GPU platform bring-up, coordinate university/external collaborator clusters, and lead incident response with runbooks and reliability practices.

