Fellow Software Development Engineer, GEMM Optimization

AMD
San Jose
Workplace: OnsiteFull timeFunction: Product ManagementEducation: phdSkills: ["Communication","Collaboration","Problem-solving","Root-cause analysis","High-performance software development"]

Own deep, hardware-aware optimization for AMD GPU GEMM and GEMM+X (including attention). Analyze GPU specifications down to the instruction and pipeline level, develop code-generator features for new ISA and HW/SW optimizations, and innovate algorithms that improve matrix-multiply throughput. Profile and root-cause performance bottlenecks, then partner with hardware architects and stakeholders to deliver targeted gains across the full stack.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
6 days ago

Fellow Software Development Engineer, GEMM Optimization

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own deep, hardware-aware optimization for AMD GPU GEMM and GEMM+X (including attention). Analyze GPU specifications down to the instruction and pipeline level, develop code-generator features for new ISA and HW/SW optimizations, and innovate algorithms that improve matrix-multiply throughput. Profile and root-cause performance bottlenecks, then partner with hardware architects and stakeholders to deliver targeted gains across the full stack.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Product Management

Key Responsibilities

  • •Analyze AMD GPU hardware specifications and work with hardware architects to align kernel design with current silicon and influence future requirements.
  • •Develop support for new ISA and HW/SW optimization features in the code-generator.
  • •Innovate and implement new algorithms for GEMM and GEMM+X (including attention).
  • •Profile and root-cause GEMM performance bottlenecks across the software and hardware stack.
  • •Partner with customers and internal stakeholders to understand real-world workloads, reproduce issues, and deliver targeted performance improvements.

Key Requirements

  • •Strong command of GPU computer architecture (compute units, register files, cache/LDS, memory bandwidth, matrix cores/WMMA-style instructions, and instruction scheduling/latency hiding).
  • •Deep expertise in GEMM algorithms and mapping onto GPU hardware (tiling/blocking, register/LDS allocation, wave scheduling, and memory hierarchy utilization).
  • •Demonstrated experience optimizing GPU compute kernels, ideally GEMM and attention.
  • •Proficiency in C++ and assembly-level GPU programming; working knowledge of Python for tooling/automation.
  • •Production-quality high-performance software engineering fundamentals with clear written and verbal communication to collaborate across hardware, customers, and engineering teams.
Experience:GPU computingAIGPU microarchitectureParallel computing
Education:PhD / Doctorate in CS/CE or related field
Skills:CommunicationCollaborationProblem-solvingRoot-cause analysisHigh-performance software development
Tech Stack:C++AssemblyPythonGEMMAttentionGPU compiler backendCDNARDNANVIDIA HopperNVIDIA BlackwellTilingBlockingRegister/LDS allocationWave schedulingMemory hierarchy utilization

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn