Research Engineer, Multimodal Data

Eventual
San Francisco
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Skills: ["Communication","Collaboration","Problem-solving"]

Research engineer role focusing on building and deploying multimodal data models for visual understanding. You’ll train and evaluate vision-language models and perception models on large-scale video/sensor datasets, design taxonomies, and ship datasets and pipelines to production, enabling efficient data curation for customer training iterations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Eventual
Eventual
4 months ago

Research Engineer, Multimodal Data

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Research engineer role focusing on building and deploying multimodal data models for visual understanding. You’ll train and evaluate vision-language models and perception models on large-scale video/sensor datasets, design taxonomies, and ship datasets and pipelines to production, enabling efficient data curation for customer training iterations.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Own the visual understanding roadmap end-to-end: from picking the model family for a customer’s taxonomy to landing it in production inference at corpus scale.
  • •Train, fine-tune, and evaluate VLMs, VQA models, embedding models, and convolutional perception models against customer datasets and benchmarks.
  • •Drive down per-clip annotation cost — model selection, distillation, batching, decode pipelining — so annotation of large corpora stays economical.
  • •Build the rich, queryable datasets that customers train on: design taxonomies with researchers, instrument quality, version outputs.
  • •Partner with the dataloading and storage teams so visual understanding outputs flow into the index and onto the GPU without re-engineering.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kEquityMeal AllowancePaid Leave

Key Requirements

  • •Strong familiarity with modern vision and multimodal models (convolution nets, VLMs, VQA, embeddings) and a sense for deployable SOTA.
  • •Experience running these models at scale on real video and sensor data for perception tasks (detection, tracking, segmentation, retrieval, captioning).
  • •Background from a perception team at a self-driving, robotics, or visual-data company, or equivalent depth from a research lab.
  • •Comfortable with cloud infrastructure and large-scale data processing; shipped jobs that ran on thousands of GPU-hours of video.
  • •Bias toward data and infrastructure: prioritize annotating the whole corpus over fine-tuning another model.
Experience:MultimodalVisionRoboticsPerceptionVideo processing
Skills:CommunicationCollaborationProblem-solving
Tech Stack:VLMsVQAEmbeddingsConvolutional netsGPUCloudDaftSparkRay

Eligibility

Visa:US citizen
Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Eventual
Eventual (daft.ai) develops Daft, an AI platform that helps developers build, simulate, and deploy autonomous agents and decision-making systems. It offers tooling for agent orchestration, environment simulation, and integrations to accelerate production AI-driven applications.
Industry: Developer Tools
Website