Algorithm - AI System Engineer

FuriosaAI
Seoul
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Experimentation","Debugging","Communication","Cross-team collaboration","Problem-solving"]

Build and validate next-generation NPU-based serving system concepts by implementing challenging POCs and prototypes, then partnering with SW teams to productize. Research and implement compression approaches such as quantization and KV cache compression directly in NPU environments, and use real-device bottlenecks and optimization lessons to propose new research ideas. Focus on hardware-level performance using a kernel programming stack and low-level interfaces, verified on real systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
1 week ago

Algorithm - AI System Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 minutes agoStatus: Live
Reposted: similar role first listed 4 weeks ago

Job Summary

Build and validate next-generation NPU-based serving system concepts by implementing challenging POCs and prototypes, then partnering with SW teams to productize. Research and implement compression approaches such as quantization and KV cache compression directly in NPU environments, and use real-device bottlenecks and optimization lessons to propose new research ideas. Focus on hardware-level performance using a kernel programming stack and low-level interfaces, verified on real systems.
Location: Seoul
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Prove core elements of next-generation serving system concepts (e.g., AFD, KV cache reuse) through NPU-based implementations (POC).
  • •Explore latest compression techniques such as quantization and KV cache compression, and implement/validate them in NPU environments.
  • •Use insights from implementation to discover and propose new research topics and optimization ideas.
  • •Implement next-generation serving system concepts directly on the company’s NPU, using a kernel programming stack and low-level programming interfaces.
  • •Perform a full validation cycle from hypothesis to real-device proof, with an emphasis on implementation and verification.

Key Requirements

  • •3+ years of hands-on experience in a related field, or equivalent capability (including research or project experience).
  • •Understanding of LLM inference fundamentals such as attention, KV cache, prefill/decode, and batching.
  • •Knowledge or usage experience with AI inference systems such as vLLM, SGLang, and TensorRT-LLM.
  • •Experience with CUDA, Triton, or accelerator programming.
  • •Experience breaking through problems without clear answers via experimentation and debugging, and communicating technical complexity while leading cross-team collaboration.
Experience:3+ yearsLLM inferenceAI systemsInference platform
Skills:ExperimentationDebuggingCommunicationCross-team collaborationProblem-solving
Tech Stack:CUDATritonVLLMSGLangTensorRT-LLMNPU

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor