Staff Software Development Engineer: GPU, AI/ML Ops & Quality Engineering

AMD
Santa Clara
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Technical ownership","Problem-solving","Communication","Mentoring"]

Build high-performance AI/ML software that makes key applications and benchmarks run faster on GPUs. Work on a core team to architect the AI software stack, optimizing from low-level GPU kernels to distributed systems. Accelerate foundation models and autonomous agent workloads while contributing across a co-design lifecycle of AMD hardware and software. Bridge C++ kernel engineering with LLM post-training and inference optimizations, and mentor others.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 month ago

Staff Software Development Engineer: GPU, AI/ML Ops & Quality Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build high-performance AI/ML software that makes key applications and benchmarks run faster on GPUs. Work on a core team to architect the AI software stack, optimizing from low-level GPU kernels to distributed systems. Accelerate foundation models and autonomous agent workloads while contributing across a co-design lifecycle of AMD hardware and software. Bridge C++ kernel engineering with LLM post-training and inference optimizations, and mentor others.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Architect and optimize an AI software stack, from GPU kernels to distributed systems, to improve performance on AMD hardware.
  • •Accelerate foundational model and autonomous AI agent workloads for demanding use cases.
  • •Contribute across the co-design lifecycle, influencing future GPU architectures and developing software for new accelerators.
  • •Bridge high-performance C++ and low-level GPU programming with LLM post-training (reinforcement learning) to improve efficiency.
  • •Mentor others and communicate technical ideas to influence direction across teams.

Key Requirements

  • •Extensive professional software development experience in performance-critical environments.
  • •Deep hands-on GPU programming and kernel optimization using HIP/CUDA, including optimizing deep learning kernels and operators.
  • •Expertise in computer vision and strong understanding of GPU architecture and memory hierarchy to diagnose performance bottlenecks.
  • •Expert-level proficiency in modern C++ and object-oriented design for complex, scalable systems.
  • •Deep experience with LLM architectures and AI post-training, including alignment and post-training techniques such as SFT and reinforcement learning (e.g., RLHF, GRPO).
Experience:AI/MLComputer visionHigh-performance computingGPU computingDistributed systemsGenerative AILLMs
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
Skills:Technical ownershipProblem-solvingCommunicationMentoring
Languages:English
Tech Stack:C++HIPCUDAGPU programmingROCmRocm ecosystemRppMIVisionXRocALRocdecodeRocjpegCV-CUDACuDNNNCCLPyTorchTensorFlowJAXLLMsTransformer architecturesAttention mechanisms

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn