AI Instinct System Management Architect

AMD
Santa Clara
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["Technical leadership","Customer engagement","Architecting","Documentation","Cross-team collaboration"]

Define and drive system management and observability architecture across AMD’s AI datacenter platforms, spanning firmware, operating systems, rack controllers, and orchestration layers. Own reference designs and blueprints to deliver secure, manageable infrastructure at rack/pod scale. Lead standards-based interfaces and telemetry frameworks (Redfish, DMTF), enable lifecycle management workflows, and collaborate with customers/partners to validate designs and influence roadmap priorities for AI workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

AI Instinct System Management Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Define and drive system management and observability architecture across AMD’s AI datacenter platforms, spanning firmware, operating systems, rack controllers, and orchestration layers. Own reference designs and blueprints to deliver secure, manageable infrastructure at rack/pod scale. Lead standards-based interfaces and telemetry frameworks (Redfish, DMTF), enable lifecycle management workflows, and collaborate with customers/partners to validate designs and influence roadmap priorities for AI workloads.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Develop a unified architecture for rack-scale and pod-scale system management, integrating from firmware through orchestration layers.
  • •Define and deliver manageability solutions (e.g., BMC/BSP, rack/pod controllers, APIs) with coherent end-to-end architecture.
  • •Create standards-based interfaces and telemetry frameworks for compute, storage, networking, and accelerators at scale.
  • •Enable rack/pod lifecycle management workflows including discovery, provisioning, firmware upgrades, and decommissioning.
  • •Collaborate with customers and partners (e.g., DCIM/ITSM environments), produce architectural collateral, and influence product strategy with customer roadmaps.

Key Requirements

  • •Expert background in systems or platform software architecture focused on system management and server manageability.
  • •Deep expertise in BMC firmware stacks, telemetry, inventory, alerting, and management protocols.
  • •Strong knowledge of DMTF standards (MCTP, PLDM, SPDM, Redfish), platform security, and management networking.
  • •Experience with PCIe, CXL, NVMe interconnects and cluster schedulers (Kubernetes, Slurm).
  • •Ability to combine technical leadership with customer engagement for scalable AI datacenter deployments.
Experience:AI datacenterTelemetryObservabilityOpen standardsOpen source
Skills:Technical leadershipCustomer engagementArchitectingDocumentationCross-team collaboration
Languages:English
Tech Stack:BMCDMTFMCTPPLDMSPDMRedfishPCIeCXLNVMeKubernetesSlurmLinuxC/C++PythonGoPrometheusLokiELKOpenTelemetryDCIM

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn