Senior Software Engineer, Inference Engine (Platform Software)
Seoul
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: bachelorsSkills: ["Problem-solving","Data analysis","Communication","Collaboration"]Build and optimize a high-performance inference engine for LLMs and multimodal LLMs running on FuriosaAI NPUs. Research and apply state-of-the-art inference optimization techniques for throughput, latency, and memory efficiency, including speculative decoding, KV-cache management, parallelism, and scheduling. Collaborate with compiler and hardware teams to co-design execution and enable distributed, scalable inference using disaggregation strategies and hierarchical/external KV-cache storage.
Loading
Loading job details...
Preparing the role view and application actions.

