Senior AI Performance Engineer - LLM Inference (vLLM)
AMD
Helsinki, Stockholm
Workplace: HybridFull timeFunction: Data Science & Machine LearningExperience: 5+ yearsEducation: mastersSkills: ["Technical communication","Problem-solving","Performance optimization","Collaboration","Work under pressure"]Own end-to-end AI inference performance optimization on AMD GPUs using vLLM. Profile, diagnose, and resolve cross-stack bottlenecks spanning GPU kernels, operator dispatch, and the vLLM scheduler/KV cache to improve throughput and latency for customer-relevant configurations like long-context and agentic coding. Lead kernel and systems-level optimizations, integrate custom kernels, support multi-node distributed inference, and deliver measurable uplifts while building reusable performance methodology.

