LLM Engineer (Optimization)

42dot
South Korea
Workplace: HybridFull timeFunction: Software EngineeringExperience: 3+ yearsSkills: ["Collaboration","Problem-solving","Performance analysis"]

Develop and optimize large language model inference for real-world services, improving latency, throughput, and memory efficiency across GPU clusters as well as edge and on-device environments. Build and enhance inference engines and runtimes, leveraging modern serving frameworks and techniques like speculative decoding and prefill/decode disaggregation. Apply model compression and compiler/kernel optimizations, and use profiling/benchmarking tools to tune trade-offs among quality, performance, and cost.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
42dot
42dot
2 months ago

LLM Engineer (Optimization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live
Reposted: similar role first listed 2 months ago

Job Summary

Develop and optimize large language model inference for real-world services, improving latency, throughput, and memory efficiency across GPU clusters as well as edge and on-device environments. Build and enhance inference engines and runtimes, leveraging modern serving frameworks and techniques like speculative decoding and prefill/decode disaggregation. Apply model compression and compiler/kernel optimizations, and use profiling/benchmarking tools to tune trade-offs among quality, performance, and cost.
Location: South Korea
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Optimize LLM inference performance (latency, throughput, memory efficiency)
  • •Research and apply optimization methods for different model structures and inference environments, including long-context and multi-turn use cases
  • •Develop and optimize GPU/accelerator-based LLM inference engines and runtimes using modern inference frameworks
  • •Apply model compression and compiler/kernel optimizations (quantization, pruning, distillation, and kernel optimization)
  • •Perform profiling and benchmarking across hardware/inference backends and design serving architectures balancing quality, performance, and cost

Key Requirements

  • •3+ years of experience in LLM, machine learning infrastructure, or inference optimization
  • •Experience developing an LLM inference engine or AI runtime
  • •Understanding of GPU architecture, CUDA programming, or parallel computing
  • •Understanding model optimization techniques including quantization and compiler optimization
  • •Experience using deep learning frameworks such as PyTorch, ONNX, and TensorRT, and strong software engineering skills in Python and/or C/C++
Experience:3+ yearsLLMMachine learning infrastructureInference optimization
Skills:CollaborationProblem-solvingPerformance analysis
Tech Stack:VLLMTensorRT-LLMSGLangLlama.cppONNX RuntimeMLXCUDATritonTensorRTTVMMLIRPyTorchONNXSpeculative DecodingPrefill-Decode DisaggregationQuantizationModel CompressionCompiler OptimizationNVIDIA Nsight SystemsNVIDIA Nsight Compute

Company Brief

42dot
Develops autonomous driving and mobility software platforms, including AI-based perception, mapping, routing, and connected-vehicle technologies. It works on next-generation transportation systems and self-driving vehicle capabilities for automotive applications.
Industry: Autonomous Vehicles
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Seoul, South Korea
Founded: 2019
WebsiteLinkedIn