AI Framework Engineer

AMD
Shanghai
Full timeCNY 581,350 - 830,500 annuallyFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Problem-solving","Independent work","Collaboration","Self-motivation","Software engineering best practices"]

Optimize and develop deep learning frameworks for AMD GPUs, improving GPU kernels, LLM performance, and training/inference efficiency across multi-GPU and multi-node systems. Collaborate with internal GPU library teams and open-source maintainers to upstream optimizations. Lead end-to-end performance engineering through profiling and bottleneck diagnosis, leveraging compiler and graph technologies to accelerate deep learning pipelines and integrate advanced techniques like FlashAttention and quantization.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

AI Framework Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Optimize and develop deep learning frameworks for AMD GPUs, improving GPU kernels, LLM performance, and training/inference efficiency across multi-GPU and multi-node systems. Collaborate with internal GPU library teams and open-source maintainers to upstream optimizations. Lead end-to-end performance engineering through profiling and bottleneck diagnosis, leveraging compiler and graph technologies to accelerate deep learning pipelines and integrate advanced techniques like FlashAttention and quantization.
Location: Shanghai
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Optimize major deep learning and LLM frameworks for AMD GPUs and contribute upstream improvements.
  • •Develop and tune GPU kernels and performance-critical operators to maximize throughput and minimize latency.
  • •Adapt and optimize LLM architectures and apply advanced techniques such as FlashAttention, PagedAttention, and quantization.
  • •Perform end-to-end performance engineering by profiling and addressing system, memory, and communication bottlenecks across multi-GPU and multi-node setups.
  • •Leverage compiler and pipeline acceleration, and prototype advanced optimization methods for production inference performance.

Pay and Benefits

Salary: CNY 581,350 - 830,500 annually

Key Requirements

  • •Strong practical experience with vLLM or SGLang and mastery of modern LLMs (e.g., DeepSeek, Qwen).
  • •Theoretical grounding in Transformer concepts including Attention, MoE, and KV Cache, plus hands-on inference optimizations (FlashAttention, PagedAttention, continuous batching, quantization like INT8/INT4/GPTQ/AWQ).
  • •Demonstrated ability to profile, diagnose, and optimize compute, memory, and communication bottlenecks in multi-GPU and multi-node environments.
  • •Experience integrating optimized GPU kernels into TensorFlow and/or PyTorch for scalable training and inference with strong throughput.
  • •Expert Python/C++ coding skills and debugging/testing practices, plus a track record of open-source contributions.
Experience:Deep learningLLMsHigh-performance computingOpen source
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or a related field
Skills:Problem-solvingIndependent workCollaborationSelf-motivationSoftware engineering best practices
Languages:English
Tech Stack:C++LinuxPythonTensorFlowPyTorchVLLMSGLangFlashAttentionPagedAttentionContinuous batchingQuantizationINT8INT4GPTQAWQTransformerAttentionMoEKV CacheMulti-GPU

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn