Algorithm - AI System Engineer

FuriosaAI
Seoul
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Communication","Cross-team collaboration","Problem-solving","Debugging","Technical leadership"]

Improve and validate next-generation NPU inference serving concepts by implementing proof-of-concepts on real hardware. Explore compression approaches like quantization and KV cache compression in an NPU environment, then use implementation insights to propose new research topics and optimization ideas. You will work across the cycle from modeling and simulation to on-device debugging and verification, partnering with SW teams to turn findings into product-ready systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
2 weeks ago

Algorithm - AI System Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 39 minutes agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Improve and validate next-generation NPU inference serving concepts by implementing proof-of-concepts on real hardware. Explore compression approaches like quantization and KV cache compression in an NPU environment, then use implementation insights to propose new research topics and optimization ideas. You will work across the cycle from modeling and simulation to on-device debugging and verification, partnering with SW teams to turn findings into product-ready systems.
Location: Seoul
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Prove key components of next-generation NPU-based serving concepts (e.g., AFD, KV cache reuse) via on-device POCs.
  • •Explore state-of-the-art compression methods (quantization, KV cache compression) and implement/validate them in the NPU environment.
  • •Generate new research topics and optimization ideas based on lessons learned from implementation.
  • •Work across a verification cycle that may include modeling, simulation, analysis, and real-hardware implementation.
  • •Collaborate with SW teams during productization to complete the serving system based on on-device findings.

Key Requirements

  • •3+ years of relevant hands-on experience (including research/project experience).
  • •Understanding of LLM inference fundamentals (attention, KV cache, prefill/decode, batching).
  • •Knowledge or hands-on experience with AI inference systems such as vLLM, SGLang, or TensorRT-LLM.
  • •Experience with CUDA or Triton, or accelerator/low-level programming.
  • •Experience tackling problems without known solutions through experimentation and debugging, while communicating technical complexity and leading cross-team collaboration.
Experience:3+ yearsLLM inferenceAI inference systemsAccelerator programming
Skills:CommunicationCross-team collaborationProblem-solvingDebuggingTechnical leadership
Tech Stack:NPULLM inferenceVLLMSGLangTensorRT-LLMCUDATritonKV cachePrefillDecodeBatchingQuantizationCompressionMulti-threadingMemory managementPerformance optimizationTensor Contraction LanguageVirtual ISARoofline modelMulti-chip parallelism

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor