Principal Cluster Reliability Architect
Austin, Texas
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Strategic thinking","Cross-functional technical leadership","Executive-level communication","Influencing technical direction","Mentorship"]Define and drive end-to-end reliability strategy for AMD’s next-generation AI and HPC cluster platforms. Serve as the technical authority across compute, networking, storage, and control plane architectures—establishing reliability frameworks, requirements, and validation plans. Partner with SRE and Platform Operations to standardize observability, health monitoring, failure detection/recovery, and lifecycle operations, then translate operational learnings into repeatable engineering practices and improved reliability outcomes.
Loading
Loading job details...
Preparing the role view and application actions.

