Principal GenAI Inference Optimization Engineer
San Jose
Workplace: HybridFull timeFunction: Communications, PR & CommunitySkills: ["Communication","Collaboration","Problem-solving","Teamwork"]Lead optimization of GenAI inference workloads on AMD GPU platforms, improving latency, throughput, and cost efficiency across single-node and distributed environments. Collaborate across kernels, runtimes, and serving frameworks to push performance of large-scale models while balancing hardware-software constraints and cross-functional partnerships.
Loading
Loading job details...
Preparing the role view and application actions.

