AI Framework Engineer

AMD
Shanghai
Workplace: OnsiteFull timeCNY 411,670 - 588,100 annuallyFunction: Data Science & Machine LearningEducation: mastersSkills: ["Problem-solving","Collaboration","Independent work","Debugging","Testing"]

Optimize and develop deep learning frameworks for AMD GPUs, improving GPU kernels, model performance, and multi-GPU/multi-node training and inference. Collaborate with internal GPU library teams and open-source maintainers to upstream optimizations, using profiling and performance engineering across compute, memory, and communication. Build acceleration via compiler technologies and graph compilers, and prototype advanced inference techniques like speculative decoding and weight-only quantization.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

AI Framework Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Optimize and develop deep learning frameworks for AMD GPUs, improving GPU kernels, model performance, and multi-GPU/multi-node training and inference. Collaborate with internal GPU library teams and open-source maintainers to upstream optimizations, using profiling and performance engineering across compute, memory, and communication. Build acceleration via compiler technologies and graph compilers, and prototype advanced inference techniques like speculative decoding and weight-only quantization.
Location: Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Optimize major DL/LLM frameworks (TensorFlow, PyTorch, vLLM, SGLang) for AMD GPUs and contribute upstream improvements.
  • •Develop and tune GPU kernels and performance-critical operators to maximize throughput and minimize latency.
  • •Adapt and optimize LLM architectures (e.g., Llama, Qwen, DeepSeek) using techniques like FlashAttention, PagedAttention, and quantization.
  • •Deliver end-to-end performance engineering by profiling and addressing system, memory, and communication bottlenecks across multi-GPU and multi-node setups.
  • •Leverage compiler and pipeline acceleration (graph compilers) and prototype advanced techniques such as speculative decoding and weight-only quantization.

Pay and Benefits

Salary: CNY 411,670 - 588,100 annually

Key Requirements

  • •Deep practical experience with vLLM or SGLang, modern LLMs (e.g., DeepSeek, Qwen), and Transformer/Attention/MoE/KV Cache concepts.
  • •Hands-on inference optimization experience including FlashAttention, PagedAttention, continuous batching, and quantization (INT8/INT4/GPTQ/AWQ).
  • •Proven ability to profile, diagnose, and optimize compute, memory, and communication bottlenecks in multi-GPU and multi-node environments.
  • •Experience integrating optimized GPU kernels into TensorFlow and/or PyTorch to accelerate training and inference at scale.
  • •Strong Python/C++ coding skills with effective debugging/testing, plus a track record of open-source contributions.
Experience:Deep learningLLMHigh-performance computingOpen source
Education:Master's in Computer Science, Computer Engineering, Electrical Engineering, or related field
Skills:Problem-solvingCollaborationIndependent workDebuggingTesting
Tech Stack:C++LinuxPythonCUDAHIPROCmLLVMMLIRTVMTensorFlowPyTorchVLLMSGLangFlashAttentionPagedAttentionLlamaQwenDeepSeekTransformerAttention

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn