ML Kernel Development and Optimization Engineer on GPU

AMD
Belgrade
Workplace: HybridFull timeFunction: Product ManagementEducation: bachelorsSkills: ["Collaboration","Problem-solving","Communication","Cross-team collaboration","Learning mindset"]

Design and develop high-performance GPU kernels for machine learning and core workloads like Convolution, GEMM, and Attention. Work with GPU architects to evaluate emerging hardware features, then analyze, profile, and optimize kernel performance. Create technical documentation and support adoption of new capabilities within ROCm, partnering with engineers, researchers, customers, and the ecosystem to improve ROCm applications, libraries, tools, and hardware.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

ML Kernel Development and Optimization Engineer on GPU

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Design and develop high-performance GPU kernels for machine learning and core workloads like Convolution, GEMM, and Attention. Work with GPU architects to evaluate emerging hardware features, then analyze, profile, and optimize kernel performance. Create technical documentation and support adoption of new capabilities within ROCm, partnering with engineers, researchers, customers, and the ecosystem to improve ROCm applications, libraries, tools, and hardware.
Location: Belgrade
Workplace: Hybrid
Employment Type: Full time
Job Function: Product Management
Seniority: Mid level

Key Responsibilities

  • •Design and develop high-performance GPU kernels for machine learning and Convolution workloads using HIP.
  • •Collaborate with GPU architects to explore, evaluate, and leverage emerging hardware features.
  • •Analyze, profile, and optimize GPU kernel performance to drive continuous improvement.
  • •Create technical documentation, share best practices, and support adoption of new ROCm capabilities.
  • •Partner with engineers, researchers, customers, and ecosystem partners to enhance ROCm applications, libraries, tools, and hardware.

Key Requirements

  • •Bachelor’s degree in computer science, Computer Engineering, or a related technical field.
  • •Significant experience developing GPU software for machine learning or HPC.
  • •Experience with GPU architecture, performance analysis, algorithm development, and machine learning workloads.
  • •Hands-on experience developing and optimizing GPU software using HIP, CUDA, or similar technologies.
  • •Understanding of GPU runtimes and compilers, machine learning frameworks, HPC libraries, and performance optimization techniques.
Experience:Machine learningHPCGPU computing
Education:Bachelor's in Computer science, Computer Engineering, or a related technical field
Skills:CollaborationProblem-solvingCommunicationCross-team collaborationLearning mindset
Languages:English
Tech Stack:HIPCUDAROCm

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn