Fellow GPU Performance Optimization Engineer

AMD
San Jose
Workplace: HybridFull timeFunction: Software EngineeringEducation: phdSkills: ["Leadership","Influence","Communication","Collaboration","Problem-solving"]

A senior-level role focused on maximizing performance and efficiency of large-scale AI training workloads on AMD GPU platforms. You will optimize across the full software-hardware stack, drive distributed training at scale, and influence architecture, software ecosystems, and best practices for generative AI workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
4 months ago

Fellow GPU Performance Optimization Engineer

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

A senior-level role focused on maximizing performance and efficiency of large-scale AI training workloads on AMD GPU platforms. You will optimize across the full software-hardware stack, drive distributed training at scale, and influence architecture, software ecosystems, and best practices for generative AI workloads.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Lead performance optimization of large-scale AI training workloads on AMD GPU platforms across single-node and multi-node environments.
  • •Identify and eliminate system bottlenecks across compute, memory, and communication (e.g., kernel efficiency, memory bandwidth, network utilization).
  • •Optimize distributed training strategies (Data, Tensor, Pipeline Parallelism, ZeRO, etc.) for scalability and efficiency on AMD hardware.
  • •Drive cross-stack optimizations spanning kernels, compilers, runtimes, communication libraries, and ML frameworks.
  • •Develop and apply advanced profiling, benchmarking, and performance modeling methodologies.

Key Requirements

  • •Deep expertise in GPU performance analysis, large-scale distributed training, and ML workloads
  • •Strong understanding of GPU architecture, interconnects, memory hierarchies, and communication patterns with ability to translate into measurable training efficiency
  • •Experience operating across layers—from kernels and runtimes to frameworks and distributed strategies; proven track record of optimization and influencing technical direction
  • •Proficiency in Python and at least one systems language (C++/CUDA/HIP); experience with compiler stacks, kernel optimization, or graph-level optimization is a strong plus
  • •Academic credentials: Ph.D. in Computer Science, Computer Engineering, or related field preferred, or equivalent industry experience with significant technical impact
Experience:GPUDistributed systemsML workloads
Education:PhD / Doctorate
Skills:LeadershipInfluenceCommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PythonC++CUDAHIPPyTorchJAXTensorFlowNsightROCm

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn