Sr. Software Engineer - Inference Engine (Platform Software)

FuriosaAI
Seoul
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["Problem-solving","Data analysis","Communication","Collaboration"]

Develop and optimize a high-performance inference engine for large and multimodal LLMs running on FuriosaAI NPUs. Research and implement state-of-the-art inference optimizations (speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling) to improve throughput, latency, and memory efficiency. Collaborate with the Compiler and hardware teams on distributed, scalable inference features like PD/EPD disaggregation and hierarchical KV-cache storage.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
1 day ago

Sr. Software Engineer - Inference Engine (Platform Software)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Develop and optimize a high-performance inference engine for large and multimodal LLMs running on FuriosaAI NPUs. Research and implement state-of-the-art inference optimizations (speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling) to improve throughput, latency, and memory efficiency. Collaborate with the Compiler and hardware teams on distributed, scalable inference features like PD/EPD disaggregation and hierarchical KV-cache storage.
Location: Seoul
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and implement the next-generation inference engine for large and multimodal language models optimized for throughput, latency, and memory efficiency
  • •Implement advanced inference optimizations including speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling
  • •Build distributed and scalable inference capabilities such as PD/EPD disaggregation, disaggregated speculative decoding, and hierarchical/external KV-cache storage
  • •Collaborate with the Compiler team to co-design and optimize execution for FuriosaAI NPUs to improve system-level throughput, latency, and memory utilization
  • •Research, evaluate, and integrate state-of-the-art inference optimization techniques and LLM serving framework features into production

Key Requirements

  • •BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience
  • •Proficiency in Rust or C++
  • •Knowledge and passion of deep learning, LLM, and/or generative AI models
  • •Excellent problem-solving and data analysis skills
  • •Strong communication and collaboration skills
Experience:3+ yearsDeep learningLLMsGenerative AIInference serving systems
Education:Bachelor's in Computer Science, Engineering, or a related field
Skills:Problem-solvingData analysisCommunicationCollaboration
Tech Stack:RustC++Deep learningLLMsGenerative AI modelsCUDATritonVLLMSGLangTensorRT-LLMHiCacheMooncakeSpeculative decodingKV-cacheTensor/model parallelism

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor