Principal AI Performance Engineer - LLM Inference (SGLang)
AMD
Helsinki, Stockholm, Cambridge
Workplace: HybridFull timeEUR 99,050 - 141,500 annuallyFunction: Data Science & Machine LearningExperience: 7+ yearsEducation: mastersSkills: ["Technical leadership","Written and verbal communication","Problem-solving","Performance optimization mindset","Customer-facing presentation"]Lead end-to-end AI inference performance optimization on AMD GPUs using SGLang. Work with a small, highly technical team to profile, diagnose, and optimize cross-stack bottlenecks—from GPU kernels through SGLang scheduling and RadixAttention/prefix caching—across customer-relevant serving configurations. Deliver measurable uplifts, integrate custom kernels, advance shared methodology, and contribute upstream improvements to SGLang and related open-source frameworks.

