Algorithm - Serving System Engineer

FuriosaAI
Seoul
Workplace: HybridFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 3+ yearsSkills: ["Technical communication","Cross-team collaboration","Problem-solving","Experimental debugging"]

Develop and prove core concepts for next-generation NPU-based serving systems, including AFD and KV cache reuse, through hands-on POC and prototyping. Explore and implement state-of-the-art compression methods such as quantization and KV cache compression in real NPU environments. Use insights from implementation and verification to propose new research topics and optimization ideas, with the implementation/validation loop as the main focus.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
1 week ago

Algorithm - Serving System Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Develop and prove core concepts for next-generation NPU-based serving systems, including AFD and KV cache reuse, through hands-on POC and prototyping. Explore and implement state-of-the-art compression methods such as quantization and KV cache compression in real NPU environments. Use insights from implementation and verification to propose new research topics and optimization ideas, with the implementation/validation loop as the main focus.
Location: Seoul
Workplace: Hybrid
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Prove key elements of next-generation serving system concepts (e.g., AFD, KV cache reuse) by implementing POC on an NPU.
  • •Explore, implement, and validate latest compression techniques (e.g., quantization and KV cache compression) in an NPU environment.
  • •Identify and propose new research tasks and optimization ideas based on insights gained from implementation.

Key Requirements

  • •3+ years of relevant hands-on experience (or comparable ability including research/project experience).
  • •Understanding of LLM inference fundamentals (attention, KV cache, prefill/decode, batching).
  • •Knowledge or hands-on experience with AI inference systems such as vLLM, SGLang, and TensorRT-LLM.
  • •Experience with CUDA, Triton, or accelerator programming.
  • •Experience experimenting and debugging to solve problems without predefined answers; ability to clearly communicate technical complexity and lead cross-team collaboration.
Experience:3+ years
Skills:Technical communicationCross-team collaborationProblem-solvingExperimental debugging
Tech Stack:NPULLM inferenceAFDKV cachePrefill/decodeBatchingVLLMSGLangTensorRT-LLMCUDATritonQuantizationKV cache compressionKernel programmingLow-level programming interfaceMulti-turn sessionsMemory hierarchySimulationPrototyping

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor