Software Engineer- GPU/AI/ML

AMD
Santa Clara
Workplace: HybridFull timeFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["C++","HIP/CUDA","PyTorch","TensorFlow","JAX","NCCL","RocBLAS","HipDNN","CUDA","CUDA paths","Nsight","ROCm"]

Senior software engineer to lead performance-critical AI software from low-level GPU kernels to distributed training and inference. Drive ROCm ecosystem improvements, optimize transformer/LLM workloads, and co-design hardware-software solutions. Mentors others and shapes AMD's AI software strategy across ROCm and FPS platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 months ago

Software Engineer- GPU/AI/ML

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Senior software engineer to lead performance-critical AI software from low-level GPU kernels to distributed training and inference. Drive ROCm ecosystem improvements, optimize transformer/LLM workloads, and co-design hardware-software solutions. Mentors others and shapes AMD's AI software strategy across ROCm and FPS platforms.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Own the AI software stack: Establish best practices and drive performance from low-level GPU kernels to large-scale distributed systems; use modern LLMs and agent-based tooling to accelerate ROCm development.
  • •Accelerate foundation models and agents: Improve training, post-training, and inference for LLMs and autonomous AI workloads so AMD is the default platform for demanding use cases.
  • •Co-design hardware and software: Partner across the full lifecycle—from GPU architecture input to software for new accelerators and engage with the AI community to stay at the forefront.
  • •Mentor others and influence across teams as the AI software strategy scales.
  • •Deliver high-impact solutions in a hands-on role, owning hard technical problems and shaping AMD’s AI software direction.

Key Requirements

  • •Expert-level modern C++ and design of large, performance-critical systems; strong grasp of GPU architecture and kernel optimization (HIP/CUDA).
  • •Hands-on delivery on large-scale C++/HIP/CUDA codebases and ROCm components (rocBLAS, hipDNN, Composable Kernel, AITemplate) and CUDA ecosystem (cuBLAS, cuDNN, CUTLASS, Thrust, CUB, NCCL).
  • •Experience with ML frameworks and their C++/HIP/CUDA paths (PyTorch, TensorFlow, or JAX).
  • •Proficiency with profilers (ROCm Profiler, Nsight) in multi-GPU distributed settings.
  • •Deep understanding of transformers, model lifecycles, and AI post-training/inference optimization
Experience:AIGPU computingHPCSoftware engineering
Education:Bachelor's
Skills:C++HIP/CUDAPyTorchTensorFlowJAXNCCLRocBLASHipDNNCUDACUDA pathsNsightROCm
Languages:English
Tech Stack:C++HIPCUDAPyTorchTensorFlowJAXNCCLRocBLASHipDNNCUTLASSThrustCUB

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn