HPC Systems Engineer - AI Workloads

AMD
San Jose
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Education: bachelorsSkills: ["Organizational skills","Problem-solving","Troubleshooting","Verbal communication","Written communication"]

Design, develop, and administer HPC infrastructure for AI workloads, focusing on GPU clusters and AI workload schedulers. Build and optimize GPU-based cluster performance, administer distributed ML/LLM and AI inferencing platforms, and automate system provisioning and cluster management. Partner cross-functionally to meet AI infrastructure needs, while monitoring and improving performance using best practices and tools such as Prometheus/Grafana and automation frameworks.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 week ago

HPC Systems Engineer - AI Workloads

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Design, develop, and administer HPC infrastructure for AI workloads, focusing on GPU clusters and AI workload schedulers. Build and optimize GPU-based cluster performance, administer distributed ML/LLM and AI inferencing platforms, and automate system provisioning and cluster management. Partner cross-functionally to meet AI infrastructure needs, while monitoring and improving performance using best practices and tools such as Prometheus/Grafana and automation frameworks.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Develop, implement, and maintain GPU-based clusters to ensure optimal performance.
  • •Administer ML/AI platforms including distributed ML services, LLMs, and AI inferencing with deployments, resource allocation, monitoring, and security.
  • •Automate system provisioning and perform cluster management end to end.
  • •Collaborate with cross-functional teams to address AI infrastructure requirements and provide technical expertise for AI-related projects.
  • •Monitor and evaluate performance of AI systems and clusters, ensuring alignment with best practices and company standards.

Key Requirements

  • •Significant experience with large-scale distributed computing for AI and HPC workloads and end-to-end ownership of outcomes.
  • •Experience developing and administering GPU-based cluster infrastructure for AI/HPC workloads.
  • •Proficiency with ML/AI platform administration, including distributed ML services, LLMs, AI inferencing, deployments, resource allocation, monitoring, and security.
  • •Experience with HPC scheduling and orchestration such as SLURM and Kubernetes, including GPU cluster optimization.
  • •Bachelor's or master's degree in computer science or computer engineering (preferred).
Education:Bachelor's in computer science or computer engineering
Skills:Organizational skillsProblem-solvingTroubleshootingVerbal communicationWritten communication
Languages:English
Tech Stack:PythonSLURMKubernetesRoCEv2KVMUbuntuShellGPU drivers400G networkingAnsibleSaltstackTerraformPrometheusGrafanaLLMsAI inferencingWeb services

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn