Principal AI Performance Engineer

AMD
San Jose
Workplace: HybridFull timeFunction: Data Science & Machine LearningExperience: 7+ yearsEducation: bachelorsSkills: ["Communication","Leadership","Problem-solving","Customer-facing","Presentation"]

Performance-focused AI engineer to optimize AI inference on AMD GPUs. Lead a small, highly technical team end-to-end across profiling, diagnosing, and optimizing models for customer-serving configurations. Drive kernel- and system-level optimizations, engage with customers, and push for measurable uplifts and reusable methodologies across multi-node deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
4 months ago

Principal AI Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Performance-focused AI engineer to optimize AI inference on AMD GPUs. Lead a small, highly technical team end-to-end across profiling, diagnosing, and optimizing models for customer-serving configurations. Drive kernel- and system-level optimizations, engage with customers, and push for measurable uplifts and reusable methodologies across multi-node deployments.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Sr. Manager level

Key Responsibilities

  • •Drive end-to-end performance optimization across the stack for leading AI models and customer-relevant serving configurations, closing gaps through kernel- and systems-level work
  • •Profile, diagnose, and resolve cross-stack performance bottlenecks from GPU kernels to framework-level scheduling and multi-node communication
  • •Lead customer-facing technical engagements: present findings, recommend optimizations, and deliver measurable performance uplifts
  • •Integrate and optimize custom kernels within serving frameworks and understand dispatch paths, shape extraction, and backend selection
  • •Develop and refine shared optimization methodologies and promote AI-assisted development across the team

Pay and Benefits

Perks:Health InsuranceDental401kRemote WorkLearning Budget

Key Requirements

  • •7+ years of software development experience in GPU computing, AI systems, or high-performance computing
  • •Deep hands-on experience with AI serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM, or similar) and their internals
  • •Strong background in end-to-end workload profiling and bottleneck diagnosis across the stack
  • •Understanding of GPU kernel performance characteristics (occupancy, register/LDS pressure, memory coalescing, cache utilization, scheduling)
  • •Bachelor's degree in Computer Science/Engineering or equivalent (advanced degree preferred)
Experience:7+ yearsAI serving frameworksGPU computingHigh-performance computing
Education:Bachelor's
Skills:CommunicationLeadershipProblem-solvingCustomer-facingPresentation
Languages:English
Tech Stack:PythonC++CUDAHIPTritonCKPyDSLASMAITERGluonTensorRT-LLMVLLMSGLangPyTorchNCCLRCCLCUDA profiling toolsCPU/GPU profiling

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn