Principal GenAI Inference Optimization Engineer

AMD
San Jose
Workplace: HybridFull timeFunction: Communications, PR & CommunitySkills: ["Communication","Collaboration","Problem-solving","Teamwork"]

Lead optimization of GenAI inference workloads on AMD GPU platforms, improving latency, throughput, and cost efficiency across single-node and distributed environments. Collaborate across kernels, runtimes, and serving frameworks to push performance of large-scale models while balancing hardware-software constraints and cross-functional partnerships.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
4 months ago

Principal GenAI Inference Optimization Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Lead optimization of GenAI inference workloads on AMD GPU platforms, improving latency, throughput, and cost efficiency across single-node and distributed environments. Collaborate across kernels, runtimes, and serving frameworks to push performance of large-scale models while balancing hardware-software constraints and cross-functional partnerships.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Communications, PR & Community
Seniority: Sr. Manager level

Key Responsibilities

  • •Optimize performance of GenAI inference workloads on AMD GPU platforms across single-node and distributed environments.
  • •Improve latency, throughput, and cost efficiency for LLM and multimodal model serving in production.
  • •Analyze and resolve bottlenecks across compute, memory, and communication.
  • •Contribute to cross-stack optimizations spanning kernels, runtimes, communication libraries, and inference/serving frameworks.
  • •Document best practices and contribute to performance guidelines for GenAI deployment.

Key Requirements

  • •Must-have experience in GenAI inference optimization and GPU performance with hands-on work on large-scale serving systems.
  • •Strong understanding of GPU architecture, memory systems, and interconnects; ability to optimize kernels, runtimes, and communication libraries.
Experience:GPU computingGenAIAI inference
Skills:CommunicationCollaborationProblem-solvingTeamwork
Languages:English
Tech Stack:PythonC++CUDAHIPPyTorchJAXTensorFlowTritonVLLMKV-cacheQuantizationCUDA/HIP

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn