Senior GPU Inference Performance Engineer

AMD
Santa Clara
Workplace: HybridFull timeFunction: Solutions Engineering & Sales EngineeringEducation: bachelorsSkills: ["Evidence-driven","Rigorous","Communication","Written reporting","Collaboration"]

Own end-to-end performance analysis for GPU-accelerated AI inference workloads. Profile and diagnose bottlenecks across GPU hardware and software runtimes, optimize inference engines and LLM serving frameworks, and run head-to-head benchmark comparisons. Analyze distributed inference networking and Kubernetes/GPU operator overhead, then automate trace collection and performance regression dashboards to present evidence-backed findings to product and executive stakeholders.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 months ago

Senior GPU Inference Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Own end-to-end performance analysis for GPU-accelerated AI inference workloads. Profile and diagnose bottlenecks across GPU hardware and software runtimes, optimize inference engines and LLM serving frameworks, and run head-to-head benchmark comparisons. Analyze distributed inference networking and Kubernetes/GPU operator overhead, then automate trace collection and performance regression dashboards to present evidence-backed findings to product and executive stakeholders.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Own full-stack GPU profiling for inference workloads across AMD Instinct (ROCm, rocProfiler, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) to identify performance bottlenecks.
  • •Profile and optimize AI inference engines including vLLM and SGLang, focusing on KV-cache management, continuous batching, PagedAttention, speculative decoding, and quantization effects.
  • •Design and execute head-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads and explain why performance differs using hardware/software evidence.
  • •Profile and optimize distributed inference topologies, analyze network-level bottlenecks with RDMA/RoCE and NCCL/RCCL profiling, and quantify impact on end-to-end inference SLAs.
  • •Build reproducible benchmarking and profiling automation, including trace collection and performance regression dashboards for continuous validation.

Key Requirements

  • •Background in GPU performance engineering, HPC, or systems performance analysis.
  • •Hands-on proficiency with AMD (ROCm, rocProfiler, Omniperf/Omnitrace) or NVIDIA (CUDA, Nsight Systems/Compute, NCU) profiling toolchains, with deep GPU architecture understanding.
  • •Experience profiling vLLM, SGLang, or equivalent LLM serving frameworks, including quantization workflows and their impact on throughput/latency.
  • •Experience with multi-GPU and multi-node inference (tensor parallelism, pipeline parallelism, or PD disaggregation) including RCCL/NCCL profiling and network tools.
  • •Strong Python and C/C++ skills, comfortable reading GPU kernel code (HIP/CUDA).
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or related technical field
Skills:Evidence-drivenRigorousCommunicationWritten reportingCollaboration
Languages:En-us
Tech Stack:GPU profilingROCmRocProfilerOmniperfOmnitraceCUDANsight SystemsNsight ComputeDCGMVLLMSGLangKV-cacheContinuous batchingPagedAttentionSpeculative decodingQuantizationFP8MXFP4INT4AWQ

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn