LLM Engineer (LLM Evaluation)
42dot
South Korea
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Collaboration"]Design and build an end-to-end evaluation system for large language models, including benchmark datasets, human/LLM-based metrics, and evaluation protocols focused on reproducibility. Develop evaluation automation using Argo Workflows and MLflow, integrate it with ML pipelines, and add regression detection with alerting for deployment validation. Operate and continuously improve model quality validation workflows for large-scale, stable model releases in Kubernetes-based environments.

