LLM Engineer (Data Generation)

42dot
South Korea
Workplace: HybridFull timeFunction: Software EngineeringExperience: 3+ yearsSkills: ["Problem-solving","Collaboration","Communication"]

Design, generate, evaluate, and iteratively improve training data to enhance next-generation generative AI model performance. Analyze model bottlenecks and failure cases with Research and Model Training teams to define data requirements across instruction, preference, reasoning, and domain-specific datasets. Build synthetic data generation pipelines, establish data quality and evaluation criteria (e.g., LLM-as-a-judge), and develop iterative filtering strategies to systematically improve dataset quality at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
42dot
42dot
2 months ago

LLM Engineer (Data Generation)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Design, generate, evaluate, and iteratively improve training data to enhance next-generation generative AI model performance. Analyze model bottlenecks and failure cases with Research and Model Training teams to define data requirements across instruction, preference, reasoning, and domain-specific datasets. Build synthetic data generation pipelines, establish data quality and evaluation criteria (e.g., LLM-as-a-judge), and develop iterative filtering strategies to systematically improve dataset quality at scale.
Location: South Korea
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and generate data to improve model performance, collaborating with Research and Model Training to define data requirements.
  • •Design, generate, and refine instruction, preference, reasoning, and domain-specific training datasets.
  • •Experimentally analyze how generated data affects model performance and iteratively improve data generation strategies.
  • •Build and operate data generation pipelines, including synthetic data automation and large-scale workflow execution.
  • •Define data quality and evaluation criteria and validate quality using LLM-as-a-judge, rule-based validation, and human feedback, then iterate via generation and filtering.

Key Requirements

  • •3+ years of experience in LLM, Machine Learning, or Data Generation-related work.
  • •General understanding of deep learning, machine learning, and natural language processing.
  • •Understanding of how training data is composed, preprocessed, evaluated for quality, and reflected into training.
  • •Strong capability in Python-based data processing and automation development.
  • •Experience handling, cleaning, filtering, and quality managing large-scale training datasets (or equivalent).
Experience:3+ yearsMachine learningData generationLLM
Skills:Problem-solvingCollaborationCommunication
Tech Stack:PythonLLMMachine LearningNatural Language ProcessingSynthetic DataLLM-as-a-JudgeRule-based ValidationHuman FeedbackPromptingDPORLHFRLAIFOpenAI EvalsLM Evaluation HarnessDeepEvalAirflowRaySparkTool CallingAgent

Company Brief

42dot
Develops autonomous driving and mobility software platforms, including AI-based perception, mapping, routing, and connected-vehicle technologies. It works on next-generation transportation systems and self-driving vehicle capabilities for automotive applications.
Industry: Autonomous Vehicles
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Seoul, South Korea
Founded: 2019
WebsiteLinkedIn