Principal AI Solutions Architect

AMD
Santa Clara
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Problem-solving","Collaboration"]

Lead design and deployment of production-grade Kubernetes-based AI infrastructure on AMD GPUs, enabling large-scale LLM training and inference. Architect multi-node, multi-datacenter GPU clusters with topology-aware scheduling, SLURM integration, and Kubernetes controllers (Kubeflow, MPI Operator, Volcano, Kueue). Translate advanced inference frameworks (vLLM, SGLang) into customer-ready solutions, accelerating time-to-production and optimizing performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
4 months ago

Principal AI Solutions Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Lead design and deployment of production-grade Kubernetes-based AI infrastructure on AMD GPUs, enabling large-scale LLM training and inference. Architect multi-node, multi-datacenter GPU clusters with topology-aware scheduling, SLURM integration, and Kubernetes controllers (Kubeflow, MPI Operator, Volcano, Kueue). Translate advanced inference frameworks (vLLM, SGLang) into customer-ready solutions, accelerating time-to-production and optimizing performance.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design and deliver reference architectures for LLM training and inference on AMD GPUs, from single-node to multi-datacenter deployments using Kubernetes and SLURM
  • •Architect and validate Kubernetes-based distributed training stacks for large-scale LLM workloads on AMD GPUs
  • •Define and implement gang scheduling and topology-aware GPU placement for multi-node training workloads
  • •Enable Kubernetes-native training controllers including Kubeflow Training Operator, MPI Operator, Volcano, and Kueue
  • •Partner with enterprise customers and cloud providers to deploy and optimize production AMD GPU clusters for distributed inference and multi-tenant workloads

Key Requirements

  • •Strong experience with Kubernetes-based distributed training and inference on GPUs (Kubeflow Training Operator, MPI Operator, Volcano, Kueue)
  • •Expertise in Kubernetes GPU orchestration, including operators, device plugins, multi-tenancy, and observability
  • •Hands-on experience with distributed training on Kubernetes (Kubeflow, MPI Operator, Volcano, Kueue, Ray)
  • •Knowledge of gang scheduling and topology-aware GPU placement for multi-node workloads
  • •Experience with LLM frameworks and GPU interconnects (RCCL/NCCL, PCIe/NVMe like interconnects) and performance benchmarking
Experience:AI infrastructureGPU
Skills:CommunicationProblem-solvingCollaboration
Languages:English
Tech Stack:KubernetesSLURMKubeflowMPI OperatorVolcanoKueueRCCLNCCLVLLMSGLangRDMACNI

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn