Senior AI Engineer (Driving VLM/VLA)

42dot
South Korea
Workplace: HybridFull timeFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Independent problem definition","Experiment design","Proactive iteration"]

Develop driving foundation models using VLM/VLA with large-scale real autonomous driving data to improve scene understanding, multimodal reasoning, and planning performance. Build and evaluate multimodal sequential models trained on camera/video, map, trajectory, actions, and language; develop applications like captioning, VQA, auto-labeling, and retrieval. Design dataset pipelines, training recipes, benchmarks, and ablations, and partner with Simulation/Data/ML Platform teams to integrate models into autonomous driving systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
42dot
42dot
1 month ago

Senior AI Engineer (Driving VLM/VLA)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Develop driving foundation models using VLM/VLA with large-scale real autonomous driving data to improve scene understanding, multimodal reasoning, and planning performance. Build and evaluate multimodal sequential models trained on camera/video, map, trajectory, actions, and language; develop applications like captioning, VQA, auto-labeling, and retrieval. Design dataset pipelines, training recipes, benchmarks, and ablations, and partner with Simulation/Data/ML Platform teams to integrate models into autonomous driving systems.
Location: South Korea
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design and develop driving foundation models based on Vision-Language-Action using autonomous driving data.
  • •Train and evaluate multimodal sequential models using camera/video, map, trajectory, action, and language modalities.
  • •Develop models for driving scene understanding, temporal reasoning, agent interaction modeling, and risk/event understanding.
  • •Build VLM-based application features such as scene captioning, visual question answering, auto-labeling, data mining, and retrieval.
  • •Create large-scale multimodal datasets and plan training recipes, evaluation benchmarks, and ablation experiments, collaborating to integrate with autonomous driving systems.

Key Requirements

  • •5+ years of research/development experience in Computer Vision, Machine Learning, Robotics, Autonomous Driving, or Multimodal AI, or equivalent proven achievements in Foundation Model/Multimodal/Generative AI.
  • •Experience building and training deep learning models using PyTorch.
  • •Deep understanding and implementation experience with one or more of: vision/video models, VLM, VLA, world model, trajectory prediction, imitation learning.
  • •Experience processing multimodal or sequential data such as images, video, sensors, trajectory, language, and actions.
  • •Understanding of one or more of Transformer, diffusion, autoregressive models, and representation learning.
Experience:5+ yearsAutonomous drivingRoboticsMultimodal AIFoundation modelGenerative AI
Skills:Independent problem definitionExperiment designProactive iteration
Tech Stack:PyTorchTransformerDiffusionAutoregressive model

Company Brief

42dot
Develops autonomous driving and mobility software platforms, including AI-based perception, mapping, routing, and connected-vehicle technologies. It works on next-generation transportation systems and self-driving vehicle capabilities for automotive applications.
Industry: Autonomous Vehicles
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Seoul, South Korea
Founded: 2019
WebsiteLinkedIn