AI Software and GPU Kernel Development Eng.

AMD
Shanghai
Workplace: OnsiteFull timeCNY 411,670 - 588,100 annuallyFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Problem-solving","Collaboration","Analytical thinking","Independent work","Proactive mindset"]

Optimize and develop deep learning frameworks for AMD GPUs, improving GPU kernel performance and accelerating training and inference at scale across multi-GPU and multi-node systems. Collaborate with internal GPU software, compiler, and GPU math library teams, and engage with open-source communities to upstream compiler and framework optimizations. Work on performance tuning for frameworks like TensorFlow and PyTorch, plus SGLang scaling and graph compiler integration (e.g., XLA, TorchDynamo).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 week ago

AI Software and GPU Kernel Development Eng.

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Optimize and develop deep learning frameworks for AMD GPUs, improving GPU kernel performance and accelerating training and inference at scale across multi-GPU and multi-node systems. Collaborate with internal GPU software, compiler, and GPU math library teams, and engage with open-source communities to upstream compiler and framework optimizations. Work on performance tuning for frameworks like TensorFlow and PyTorch, plus SGLang scaling and graph compiler integration (e.g., XLA, TorchDynamo).
Location: Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Optimize deep learning frameworks (e.g., TensorFlow, PyTorch, SGLang) on AMD GPUs through upstream open-source contributions.
  • •Develop and tune large-scale training and inference models for optimal performance on AMD hardware.
  • •Design, implement, and optimize high-performance GPU kernels using HIP, Triton, or related tools.
  • •Collaborate with internal compiler and GPU library teams to align kernel-level optimizations with end-to-end performance goals.
  • •Tune and scale performance across multi-GPU and multi-node environments, including inference parallelism and graph compiler integration (e.g., XLA, TorchDynamo).

Pay and Benefits

Salary: CNY 411,670 - 588,100 annually

Key Requirements

  • •Strong C++ development experience in Linux environments, with the ability to define goals and deliver high-quality solutions.
  • •Experience debugging, profiling, and optimizing performance-critical code in C++ and/or Python.
  • •Hands-on experience optimizing SGLang or similar LLM inference frameworks (preferred).
  • •Knowledge of compiler design or familiarity with LLVM, MLIR, or ROCm (plus).
  • •Strong experience running and scaling distributed workloads on heterogeneous CPU+GPU clusters for training or inference (plus).
Experience:AIDeep learningLLM inferenceDistributed trainingOpen-source
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or a related field
Skills:Problem-solvingCollaborationAnalytical thinkingIndependent workProactive mindset
Languages:En-us
Tech Stack:C++PythonLinuxTensorFlowPyTorchSGLangHIPTritonXLATorchDynamoLLVMMLIRROCmCUDAGraph compilersDistributed trainingDistributed inferenceMulti-GPUMulti-nodeCollective communication

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn