Senior Software Engineer, Inference Engine (Platform Software)

FuriosaAI
Seoul
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: bachelorsSkills: ["Problem-solving","Data analysis","Communication","Collaboration"]

Build and optimize a high-performance inference engine for LLMs and multimodal LLMs running on FuriosaAI NPUs. Research and apply state-of-the-art inference optimization techniques for throughput, latency, and memory efficiency, including speculative decoding, KV-cache management, parallelism, and scheduling. Collaborate with compiler and hardware teams to co-design execution and enable distributed, scalable inference using disaggregation strategies and hierarchical/external KV-cache storage.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
2 weeks ago

Senior Software Engineer, Inference Engine (Platform Software)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 39 minutes agoStatus: Live
Reposted: similar role first listed 4 weeks ago

Job Summary

Build and optimize a high-performance inference engine for LLMs and multimodal LLMs running on FuriosaAI NPUs. Research and apply state-of-the-art inference optimization techniques for throughput, latency, and memory efficiency, including speculative decoding, KV-cache management, parallelism, and scheduling. Collaborate with compiler and hardware teams to co-design execution and enable distributed, scalable inference using disaggregation strategies and hierarchical/external KV-cache storage.
Location: Seoul
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and implement the next-generation inference engine for large and multimodal language models optimized for throughput, latency, and memory efficiency.
  • •Implement advanced inference optimizations such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling.
  • •Develop distributed, scalable inference capabilities including PD/Decode and EPD disaggregation plus hierarchical/external KV-cache storage (e.g., HiCache, Mooncake).
  • •Collaborate with the Compiler team to co-design and optimize execution for FuriosaAI NPUs to improve system-level performance.
  • •Research, evaluate, and integrate state-of-the-art inference optimization techniques and serving-framework features into production.

Key Requirements

  • •BS in Computer Science/Engineering (or related) with at least 3 years of relevant experience, or equivalent practical experience.
  • •Proficiency in Rust or C++.
  • •Knowledge and passion for deep learning, LLMs, and/or generative AI models.
  • •Excellent problem-solving and data analysis skills.
  • •Strong communication and collaboration skills.
Experience:Large language models
Education:Bachelor's in Computer Science, Engineering
Skills:Problem-solvingData analysisCommunicationCollaboration
Languages:English
Tech Stack:RustC++CUDATritonVLLMSGLangTensorRT-LLMHiCacheMooncake

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor