LLM Engineer (Reinforcement Learning)
42dot
South Korea
Workplace: HybridFull timeFunction: Education & TrainingExperience: 3+ yearsSkills: ["Collaboration"]Design and improve LLM training pipelines that power generation models used in real services. Drive efficiency and enhance accuracy and stability by applying alignment approaches (e.g., PPO/GRPO/DPO) and building training structures that mitigate reward hacking and enable self-refinement. Develop foundational models that integrate with external knowledge and APIs, including learning how to select the right tools based on instruction type.

