LLM Engineer (Optimization)
42dot
South Korea
Workplace: HybridFull timeFunction: Software EngineeringExperience: 3+ yearsSkills: ["Performance analysis","Problem-solving","Collaboration"]Optimize large language model inference for real-world production: improve latency, throughput, and memory efficiency across long-context and multi-turn conversations. Build and accelerate GPU/accelerator-based inference engines and runtimes using modern frameworks (vLLM, TensorRT-LLM, SGLang, llama.cpp, ONNX Runtime, MLX). Advance model compression and compiler/kernel optimizations (quantization, CUDA/Triton/TVM/MLIR) and benchmark performance across hardware backends.

