Software Engineer, LLM Performance & Evaluation

FuriosaAI
Seoul
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["Communication","Collaboration","Experimental design","Data analysis","Statistical analysis"]

Own measurement and continuous tracking of LLM inference performance and model accuracy on NPUs. Build benchmarks, evaluation pipelines, and analysis tools to isolate inference-feature impacts, quantify performance–accuracy trade-offs, and detect regressions as Furiosa-LLM evolves. Partner with MLSys, Inference Engine, Compiler, and LLM Serving teams to establish reproducible baselines, automate scheduled/triggered evaluations, and integrate evaluation checks into CI and release workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
10 hours ago

Software Engineer, LLM Performance & Evaluation

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Own measurement and continuous tracking of LLM inference performance and model accuracy on NPUs. Build benchmarks, evaluation pipelines, and analysis tools to isolate inference-feature impacts, quantify performance–accuracy trade-offs, and detect regressions as Furiosa-LLM evolves. Partner with MLSys, Inference Engine, Compiler, and LLM Serving teams to establish reproducible baselines, automate scheduled/triggered evaluations, and integrate evaluation checks into CI and release workflows.
Location: Seoul
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and maintain LLM inference benchmarks to isolate impacts of inference features like continuous batching, prefix caching, speculative decoding, and distributed inference.
  • •Measure and report time to first token (TTFT), latency distributions, throughput, memory use, and accelerator utilization across models and configurations.
  • •Own model-level accuracy evaluations using representative datasets and task-appropriate metrics, comparing NPU results to trusted references and prior releases.
  • •Automate benchmark and accuracy runs with scheduled and change-triggered checks, versioning models, datasets, prompts, scoring logic, and execution configurations for reproducibility.
  • •Build dashboards and regression workflows, defining thresholds and release acceptance criteria, investigating regressions with profiling and controlled experiments, and communicating trade-offs through technical reports.

Key Requirements

  • •Strong Python skills building maintainable automation, data pipelines, or engineering tools.
  • •Hands-on experience measuring and analyzing ML inference or complex systems performance, including profiling and latency/throughput analysis.
  • •Understanding of transformer-based LLM inference (prefill/decode, batching, KV-cache behavior, and effects of workload and numerical precision).
  • •Practical experience evaluating model accuracy or numerical correctness using PyTorch, Hugging Face Transformers, or comparable tools.
  • •Sound experimental design and data analysis skills for controlled comparisons and reproducible reporting.
Experience:LLM inferenceAI infrastructureMachine learningBenchmarking
Skills:CommunicationCollaborationExperimental designData analysisStatistical analysis
Languages:English
Tech Stack:PythonPyTorchHugging Face TransformersVLLMSGLangTensorRT-LLMLm-evaluation-harnessLinuxCIContainersClustersRustC++KV-cacheTransformersQuantizationMixed precisionSpeculative decodingDistributed inferenceGPU

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor