Principal / Senior GPU SW Performance Engineer — Post‑Training
AMD
San Jose
Workplace: HybridFull timeFunction: Education & TrainingEducation: bachelorsSkills: ["Python","C++","PyTorch","ROCm","HIP","Distributed training","Torch.distributed","FSDP","ZeRO"]Lead performance for post-training workloads on AMD Instinct GPUs, driving fast, stable, and reproducible training pipelines across kernels, distributed training, and framework integrations. Collaborate with AMD teams (framework, compiler, kernel, model) to improve throughput, memory efficiency, and scalability for deep learning workloads.

