Principal Cluster Reliability Architect

AMD
Austin, Texas
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Strategic thinking","Cross-functional technical leadership","Executive-level communication","Influencing technical direction","Mentorship"]

Define and drive end-to-end reliability strategy for AMD’s next-generation AI and HPC cluster platforms. Serve as the technical authority across compute, networking, storage, and control plane architectures—establishing reliability frameworks, requirements, and validation plans. Partner with SRE and Platform Operations to standardize observability, health monitoring, failure detection/recovery, and lifecycle operations, then translate operational learnings into repeatable engineering practices and improved reliability outcomes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 days ago

Principal Cluster Reliability Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Define and drive end-to-end reliability strategy for AMD’s next-generation AI and HPC cluster platforms. Serve as the technical authority across compute, networking, storage, and control plane architectures—establishing reliability frameworks, requirements, and validation plans. Partner with SRE and Platform Operations to standardize observability, health monitoring, failure detection/recovery, and lifecycle operations, then translate operational learnings into repeatable engineering practices and improved reliability outcomes.
Location: Austin, Texas
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Define and implement comprehensive cluster reliability architecture spanning the full infrastructure lifecycle, from reference architectures to customer deployments.
  • •Drive architecture decisions to improve fault tolerance, redundancy, failure isolation, and system recoverability across hardware, firmware, software, and operations.
  • •Partner with SRE and Platform Operations to establish sustainable operational models, including observability/telemetry, health monitoring, failure detection/recovery, and lifecycle patching/upgrades.
  • •Define reliability KPIs, SLAs, SLOs, and deployment acceptance criteria, and execute reliability validation at production scale (including failure injection, chaos engineering, and recovery validation).
  • •Provide cross-functional technical leadership as the reliability subject-matter expert, mentor senior engineers, and represent reliability strategy in customer engagements and programs.

Key Requirements

  • •Highly experienced systems architect with deep expertise in distributed systems, large-scale infrastructure, and reliability engineering for production-scale platforms.
  • •Strong understanding of AI/HPC environments, including designing and operating reliable infrastructure platforms.
  • •Demonstrated ability to define reliability frameworks and requirements across the full cluster lifecycle, including architecture, deployment, validation, and operations.
  • •Experience partnering with SRE and Platform Operations to establish day-2 operational models, observability/telemetry, and health monitoring standards.
  • •Bachelor’s degree in a related technical discipline (Computer Engineering, Computer Science, Electrical Engineering, or similar).
Experience:AI/HPCDistributed systemsReliability engineeringSRECloud infrastructurePlatform infrastructure
Education:Bachelor's in Computer Engineering, Computer Science, Electrical Engineering, or a related technical discipline
Skills:Strategic thinkingCross-functional technical leadershipExecutive-level communicationInfluencing technical directionMentorship
Languages:En-us
Tech Stack:KubernetesSlurmROCmLinuxInfiniBandRoCERDMA networkingGPU cluster architecturesAI training infrastructureAI inference infrastructureHigh-performance cluster interconnectsDistributed storage architecturesObservability platformsTelemetry systemsInfrastructure-as-codeChaos engineeringFailure injection testingIncident managementRoot cause analysis

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn