Triton Compiler/GPU Kernel Performance Engineer

AMD
Shanghai
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringSkills: ["Communication","Leadership","Problem-solving"]

Kernel Performance Architect responsible for defining, analyzing, and optimizing performance across the full stack—from GPU microarchitecture and compiler behavior to runtime systems and deep learning frameworks—for AI workloads on AMD GPUs. Lead cross-team efforts on cross-architecture optimization, performance modeling, and portability, while guiding implementation engineers and communicating tradeoffs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
4 months ago

Triton Compiler/GPU Kernel Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Kernel Performance Architect responsible for defining, analyzing, and optimizing performance across the full stack—from GPU microarchitecture and compiler behavior to runtime systems and deep learning frameworks—for AI workloads on AMD GPUs. Lead cross-team efforts on cross-architecture optimization, performance modeling, and portability, while guiding implementation engineers and communicating tradeoffs.
Location: Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Lead performance architecture for key AI workloads across the full stack (kernel, memory, compiler, runtime, and frameworks).
  • •Define cross-architecture optimization strategies and diagnose bottlenecks across kernel, memory, compiler, and runtime layers.
  • •Design and implement performance methodologies, benchmarking frameworks, and regression metrics to ensure portability and repeatability.
  • •Collaborate with compiler, runtime, and hardware teams (LLVM, ROCm) to analyze generated code and guide optimization directions.
  • •Provide technical leadership by mentoring kernel engineers and communicating performance tradeoffs to hardware and software teams.

Pay and Benefits

Perks:Remote WorkHealth Insurance401kEquity

Key Requirements

  • •Strong experience in GPU kernel development and optimization (HIP, CUDA, or similar).
  • •Deep understanding of GPU microarchitecture concepts (memory hierarchy, registers, scheduling, occupancy).
  • •Excellent C++ experience in Linux environments; comfortable reading disassembly and compiler IR.
  • •Proven ability to diagnose performance bottlenecks using profiling tools and to design experiments to isolate bottlenecks.
  • •Ability to predict memory-bound vs compute-bound behavior and to design optimization strategies across kernel, memory, compiler, and runtime layers.
Experience:GpuAiHpcGpu kernel
Skills:CommunicationLeadershipProblem-solving
Languages:English
Tech Stack:HIPCUDAROCmLLVMC++LinuxGPUMemory hierarchyDisassemblyIR

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn