Principal / Senior GPU SW Performance Engineer — Post‑Training

AMD
San Jose
Workplace: HybridFull timeFunction: Education & TrainingEducation: bachelorsSkills: ["Python","C++","PyTorch","ROCm","HIP","Distributed training","Torch.distributed","FSDP","ZeRO"]

Lead performance for post-training workloads on AMD Instinct GPUs, driving fast, stable, and reproducible training pipelines across kernels, distributed training, and framework integrations. Collaborate with AMD teams (framework, compiler, kernel, model) to improve throughput, memory efficiency, and scalability for deep learning workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
5 months ago

Principal / Senior GPU SW Performance Engineer — Post‑Training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Lead performance for post-training workloads on AMD Instinct GPUs, driving fast, stable, and reproducible training pipelines across kernels, distributed training, and framework integrations. Collaborate with AMD teams (framework, compiler, kernel, model) to improve throughput, memory efficiency, and scalability for deep learning workloads.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Education & Training
Seniority: Manager level

Key Responsibilities

  • •Lead performance for finetuning and RL training solutions on AMD GPUs.
  • •Improve throughput, memory efficiency, and stability across data, model, and optimizer steps.
  • •Optimize multi-GPU/multi-node training and communication patterns.
  • •Profile, diagnose, and resolve bottlenecks using standard tooling; prevent regressions in CI.
  • •Ship reproducible pipelines and documentation and collaborate with framework, compiler, and model teams to land durable improvements.

Key Requirements

  • •BS/MS/PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
  • •Proven GPU performance engineering for deep learning (ROCm/HIP, Triton, or similar)
  • •Proficient in Python and C++
  • •Distributed training experience with multi-GPU/multi-node setups (torch.distributed, FSDP/ZeRO)
  • •Hands-on experience with PyTorch and RL/finetuning workflows
Experience:Deep learningGPUAI
Education:Bachelor's in Computer Science
Skills:PythonC++PyTorchROCmHIPDistributed trainingTorch.distributedFSDPZeRO
Languages:English
Tech Stack:PythonC++PyTorchROCmHIPDistributed trainingTorch.distributedFSDPZeRO

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn