Software Engineer, AI Model Enablement & Inference

FuriosaAI
Seoul
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["Technical communication","Cross-team collaboration","Analysis","Debugging"]

Build production-ready inference support for frontier LLMs on FuriosaAI’s Tensor Contraction Processor (TCP) architecture. You’ll develop and optimize TCL (Tensor Contraction Language) model kernels (including attention and mixture-of-experts), integrate new models into Furiosa-LLM, and validate correctness on NPUs with reference comparisons and regression tests. You’ll also improve onboarding via reusable analysis/integration tools and evaluate approaches from vLLM and SGLang to enhance kernel integration and performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
10 hours ago

Software Engineer, AI Model Enablement & Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build production-ready inference support for frontier LLMs on FuriosaAI’s Tensor Contraction Processor (TCP) architecture. You’ll develop and optimize TCL (Tensor Contraction Language) model kernels (including attention and mixture-of-experts), integrate new models into Furiosa-LLM, and validate correctness on NPUs with reference comparisons and regression tests. You’ll also improve onboarding via reusable analysis/integration tools and evaluate approaches from vLLM and SGLang to enhance kernel integration and performance.
Location: Seoul
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Analyze model architectures, algorithms, and reference implementations to define implementation requirements and trade-offs, partnering with Inference Engine and Compiler teams.
  • •Design, implement, and optimize model-specific TCL kernels for high-performance, efficient NPU resource usage (e.g., attention and mixture-of-experts).
  • •Integrate new models and TCL kernels into Furiosa-LLM to enable correct, efficient inference.
  • •Improve onboarding for new models by building reusable analysis/integration tools, automating validation and benchmarking, and documenting repeatable workflows.
  • •Validate kernel and model correctness on NPUs vs reference implementations, investigate numerical differences, and build regression tests; evaluate techniques from vLLM and SGLang and document findings.

Key Requirements

  • •Deep understanding of transformer-based LLMs, including attention variants, mixture-of-experts architectures, and KV-cache behavior.
  • •Strong Python skills with hands-on ability to read, modify, and debug model implementations in PyTorch (or similar).
  • •Hands-on experience implementing, debugging, and optimizing tensor operations or accelerator kernels, with attention to compute, memory, and numerical correctness.
  • •Understanding of LLM inference performance (prefill/decode, batching, latency–throughput trade-offs) and experience with quantitative performance evaluation.
  • •Ability to read and reason about Rust or C++ code when working on inference systems.
Experience:LLM inferenceTransformersMixture-of-expertsAccelerator kernels
Skills:Technical communicationCross-team collaborationAnalysisDebugging
Languages:English
Tech Stack:PythonPyTorchRustC++Tensor Contraction Language (TCL)TCLTensor Contraction Processor (TCP)Furiosa-LLMNPUsVLLMSGLangTensorRT-LLMCUDATritonML compilersKV-cacheAttentionMixture-of-experts

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor