Principal Software Engineer – PyTorch Training Frameworks

AMD
San Jose
Workplace: HybridFull timeFunction: Software EngineeringEducation: bachelorsSkills: ["Communication","Problem-solving","Leadership"]

Seeking a principal-level PyTorch training framework expert to drive performance, scalability, and correctness of large-scale AI training on AMD Instinct accelerators. You’ll lead PyTorch internals work, optimize single-node and multi-node workloads, and collaborate across compiler, kernel, driver, and architecture teams to deliver industry-leading training performance and developer experience.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
6 months ago

Principal Software Engineer – PyTorch Training Frameworks

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Seeking a principal-level PyTorch training framework expert to drive performance, scalability, and correctness of large-scale AI training on AMD Instinct accelerators. You’ll lead PyTorch internals work, optimize single-node and multi-node workloads, and collaborate across compiler, kernel, driver, and architecture teams to deliver industry-leading training performance and developer experience.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Act as a technical authority for PyTorch training at AMD, setting direction for performance, scalability, and reliability
  • •Drive optimization of key PyTorch training workloads (LLMs/foundation models) across single-node and multi-node systems
  • •Improve and debug training performance in areas such as DDP/FSDP, gradient checkpointing, mixed precision, memory planning, and communication/overlap
  • •Partner with ROCm compiler/runtime, kernel, and driver teams to resolve performance bottlenecks and correctness issues across the full stack
  • •Contribute to upstream PyTorch discussions, design, and performance fixes; develop benchmarks, profiling workflows, and perform performance regression detection

Key Requirements

  • •Deep experience with PyTorch training systems (Autograd, optimizers, dataloading, compilation paths, runtime behavior)
  • •Strong distributed training expertise: DDP, FSDP, tensor/pipeline parallel concepts, NCCL/RCCL, multi-node debugging
  • •Strong programming skills in Python and C/C++ (ability to land clean, maintainable changes in large codebases)
  • •Familiarity with PyTorch ecosystem components such as TorchInductor / torch.compile, Triton, CUDA/HIP-style programming models, and performance tooling
  • •Experience working across OS/hardware boundaries in Linux-based environments (containers, CI, drivers/runtimes)
Experience:OpenAIDistributed trainingPyTorch
Education:Bachelor's
Skills:CommunicationProblem-solvingLeadership
Languages:English
Tech Stack:PyTorchAutogradDataloadingTorchInductorTorch.compileTritonCUDAHIPPythonC/C++NCCLRCCLDDPFSDPMemory optimizationProfilingCIContainersROCm

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn