Principal ML Engineer - Large Scale Training Performance Optimization

AMD
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: mastersSkills: ["Communication","Problem-solving","Collaboration"]

Lead efforts to train large AI models at scale on AMD GPUs, optimizing distributed training pipelines and end-to-end performance. Collaborate across teams to push the AMD AI platform forward, contribute to open source, and stay at the forefront of training algorithms and techniques for large-scale models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
4 months ago

Principal ML Engineer - Large Scale Training Performance Optimization

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Lead efforts to train large AI models at scale on AMD GPUs, optimizing distributed training pipelines and end-to-end performance. Collaborate across teams to push the AMD AI platform forward, contribute to open source, and stay at the forefront of training algorithms and techniques for large-scale models.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Train large models to convergence on AMD GPUs at scale.
  • •Improve the end-to-end training pipeline performance.
  • •Optimize the distributed training pipeline and algorithm to scale out.
  • •Contribute your changes to open source.
  • •Stay up-to-date with the latest training algorithms and influence the AMD AI platform.

Key Requirements

  • •Must have experience with distributed training of large models on GPUs and hands-on expertise in distributed training algorithms (Data Parallel, Tensor Parallel, Pipeline Parallel, Expert Parallel ZeRO).
  • •Strong experience with ML/DL frameworks (e.g., PyTorch, JAX, TensorFlow) and distributed training frameworks (e.g., Megatron-LM, MaxText, TorchTitan).
  • •Excellent Python or C++ programming skills, with debugging, profiling, and performance analysis at scale.
  • •Experience with ML infra at kernel, framework, or system level is a plus.
  • •Academic credentials include a master's degree or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field.
Experience:AI/MLHigh Performance ComputingGPU computing
Education:Master's
Skills:CommunicationProblem-solvingCollaboration
Languages:English
Tech Stack:PythonC++PyTorchJAXTensorFlowMegatron-LMGPUCUDADistributed training

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn