AI Systems Engineer - HPC

AMD
San Jose
Workplace: OnsiteFull timeUSD 151,900 - 260,400 annuallyFunction: IT Operations (Systems/Network Admin)Skills: ["Organizational skills","Problem-solving","Troubleshooting","Communication"]

Design, develop, and administer GPU-focused HPC infrastructure for AMD’s AI compute platforms. Build and maintain GPU-based clusters, manage deployments, and administer distributed ML services including LLMs and AI inferencing. Automate provisioning and cluster management, collaborate with cross-functional teams on AI infrastructure needs, and monitor performance to ensure best-practice reliability and efficiency. Use AI/ML to improve internal delivery tools and processes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 days ago

AI Systems Engineer - HPC

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design, develop, and administer GPU-focused HPC infrastructure for AMD’s AI compute platforms. Build and maintain GPU-based clusters, manage deployments, and administer distributed ML services including LLMs and AI inferencing. Automate provisioning and cluster management, collaborate with cross-functional teams on AI infrastructure needs, and monitor performance to ensure best-practice reliability and efficiency. Use AI/ML to improve internal delivery tools and processes.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Develop, implement, and maintain GPU-based clusters to ensure optimal performance
  • •Administer ML/AI platforms including distributed ML services, LLMs, and AI inferencing (deployments, resource allocation, monitoring, security)
  • •Automate system provisioning and manage cluster operations end to end
  • •Collaborate with cross-functional teams to meet AI infrastructure requirements and provide technical expertise
  • •Monitor and evaluate AI system and cluster performance against industry best practices and company standards

Pay and Benefits

Salary: USD 151,900 - 260,400 annually

Key Requirements

  • •Experience developing Python-based AI apps (and UI)
  • •HPC infrastructure engineering experience in an AI/HPC domain
  • •Hands-on SLURM and Kubernetes management for AI workload environments
  • •Experience managing GPU clusters and optimizing GPU-based services, tools, or software
  • •Proficiency with cluster networking and systems tooling (e.g., RoCEv2, KVM, Ubuntu, GPU drivers, and 400G interconnect)
Experience:AIHPCDistributed computingGPU clusters
Education:
Skills:Organizational skillsProblem-solvingTroubleshootingCommunication
Languages:English
Tech Stack:PythonShellUbuntuKubernetesK8sSLURMRoCEv2KVMGPU driversGPU clusters400G networkingAnsibleSaltstackTerraformPrometheusGrafanaWeb servicesLLMs

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn