AI Software Development Eng.

AMD
Shanghai
Full timeCNY 814,450 - 1,163,500 annuallyFunction: Data Science & Machine LearningSkills: ["Debugging","Performance analysis","Communication","Cross-functional collaboration"]

Build and optimize AMD’s large-scale training infrastructure on AMD GPUs. Work on distributed training performance, stability, and scalability by improving training frameworks, parallelism strategies, and communication scheduling to reduce latency and maximize utilization. Tune low-level core operators with HIP/CUDA and profiling tools, integrate open-source training frameworks (e.g., Megatron-LM, TorchTitan, DeepSpeed), and partner with model and platform teams to resolve system bottlenecks.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

AI Software Development Eng.

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Build and optimize AMD’s large-scale training infrastructure on AMD GPUs. Work on distributed training performance, stability, and scalability by improving training frameworks, parallelism strategies, and communication scheduling to reduce latency and maximize utilization. Tune low-level core operators with HIP/CUDA and profiling tools, integrate open-source training frameworks (e.g., Megatron-LM, TorchTitan, DeepSpeed), and partner with model and platform teams to resolve system bottlenecks.
Location: Shanghai
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Develop and maintain AMD’s internal training framework for large-scale training on AMD GPUs.
  • •Optimize distributed training pipelines and parallelism strategies including data, tensor, and pipeline parallelism, plus ZeRO.
  • •Improve communication scheduling and kernel overlap to reduce training latency and maximize GPU utilization.
  • •Tune performance of core operators using HIP/CUDA and low-level profiling tools.
  • •Integrate and adapt open-source training frameworks such as Megatron-LM, TorchTitan, and DeepSpeed; resolve system-level bottlenecks.

Pay and Benefits

Salary: CNY 814,450 - 1,163,500 annually

Key Requirements

  • •Solid engineering background with familiarity with end-to-end deep learning training workflows.
  • •Hands-on experience with training framework internals such as Megatron-LM, TorchTitan, DeepSpeed, or FairScale.
  • •Strong debugging and performance analysis skills using profiling/tracing tools.
  • •Understanding of distributed training techniques including data parallelism, tensor parallelism, pipeline parallelism, and ZeRO optimization.
  • •Excellent communication and cross-functional collaboration skills.
Skills:DebuggingPerformance analysisCommunicationCross-functional collaboration
Languages:En-us
Tech Stack:AMD GPUsHIPCUDAMegatron-LMTorchTitanDeepSpeedFairScaleZeRONCCLRCCLDistributed trainingData parallelismTensor parallelismPipeline parallelismProfilingTracingKernel overlap

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn