Senior Software Engineer, Inference Engine (Platform Software)

FuriosaAI
Seoul
Full timeFunction: Software EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["Problem-solving","Data analysis","Communication","Collaboration"]

Develop and optimize a high-performance inference engine for large and multimodal LLMs running on FuriosaAI NPUs. Research and implement state-of-the-art inference optimizations and serving features (throughput, latency, memory efficiency), including speculative decoding, KV-cache management, and tensor/model parallelism. Build distributed and scalable inference capabilities such as prefill–decode disaggregation and hierarchical KV-cache storage, partnering with compiler and hardware teams to maximize system performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
1 week ago

Senior Software Engineer, Inference Engine (Platform Software)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 minutes agoStatus: Live
Reposted: similar role first listed 2 weeks ago

Job Summary

Develop and optimize a high-performance inference engine for large and multimodal LLMs running on FuriosaAI NPUs. Research and implement state-of-the-art inference optimizations and serving features (throughput, latency, memory efficiency), including speculative decoding, KV-cache management, and tensor/model parallelism. Build distributed and scalable inference capabilities such as prefill–decode disaggregation and hierarchical KV-cache storage, partnering with compiler and hardware teams to maximize system performance.
Location: Seoul
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and implement a next-generation inference engine for large and multimodal language models, optimized for throughput, latency, and memory efficiency.
  • •Implement advanced inference optimizations such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling.
  • •Build distributed and scalable inference capabilities including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation and disaggregated speculative decoding.
  • •Develop hierarchical/external KV-cache storage capabilities (e.g., HiCache and Mooncake).
  • •Collaborate with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization.

Key Requirements

  • •BS in Computer Science, Engineering, or related field with at least 3 years of relevant industry experience (or equivalent practical experience).
  • •Proficiency in Rust or C++.
  • •Knowledge and passion for deep learning, LLMs, and/or generative AI models.
  • •Excellent problem-solving and data analysis skills.
  • •Strong communication and collaboration skills.
Experience:3+ yearsDeep learningLLMsGenerative AI
Education:Bachelor's in Computer Science, Engineering, or a related field
Skills:Problem-solvingData analysisCommunicationCollaboration
Languages:English
Tech Stack:RustC++CUDATritonVLLMSGLangTensorRT-LLMHiCacheMooncake

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor