LLM Engineer (Optimization)
South Korea
Workplace: HybridFull timeFunction: Software EngineeringExperience: 3+ yearsSkills: ["Collaboration","Problem-solving","Performance analysis"]Develop and optimize large language model inference for real-world services, improving latency, throughput, and memory efficiency across GPU clusters as well as edge and on-device environments. Build and enhance inference engines and runtimes, leveraging modern serving frameworks and techniques like speculative decoding and prefill/decode disaggregation. Apply model compression and compiler/kernel optimizations, and use profiling/benchmarking tools to tune trade-offs among quality, performance, and cost.
Loading
Loading job details...
Preparing the role view and application actions.

