Multimodal ML Engineer

White Circle
Paris, London
Workplace: HybridFull timeUSD 120,000 - 250,000 annuallyFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Experimentation","Engineering fundamentals","Shipping to production","Distributed training","Documentation"]

Train and fine-tune large-scale multimodal models (vision-language, audio, and speech) for an AI safety platform that runs at 100M+ API calls per month. You’ll design experiments, build multimodal data pipelines (including synthetic data), and develop alignment across modalities (SFT/DPO/GRPO and reward modeling). Own end-to-end delivery, optimizing MoE inference and deploying models for low-latency production serving with meaningful evaluation benchmarks.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
White Circle
White Circle
2 months ago

Multimodal ML Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Train and fine-tune large-scale multimodal models (vision-language, audio, and speech) for an AI safety platform that runs at 100M+ API calls per month. You’ll design experiments, build multimodal data pipelines (including synthetic data), and develop alignment across modalities (SFT/DPO/GRPO and reward modeling). Own end-to-end delivery, optimizing MoE inference and deploying models for low-latency production serving with meaningful evaluation benchmarks.
Location: Paris, London
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Train and fine-tune large-scale multimodal models from scratch and from pretrained checkpoints.
  • •Extend and improve models across modalities, including image understanding, video temporal modeling, long-context processing, and streaming audio.
  • •Design and run training experiments, including architecture changes, data mixes, and training recipes.
  • •Build and maintain multimodal data pipelines to create training-ready datasets, including synthetic data generation.
  • •Deploy models end-to-end for production and define evaluation metrics and benchmarks that matter for the product.

Pay and Benefits

Salary: USD 120,000 - 250,000 annually
Perks:Paid LeaveHealth InsuranceRelocationSubscriptionsOff-sites

Key Requirements

  • •3+ years training large-scale deep learning models in multimodal domains (vision-language, audio, speech, or acoustic).
  • •Strong PyTorch skills with hands-on distributed training experience (DeepSpeed, FSDP, or similar).
  • •Deep experience with multimodal architectures and how vision/audio encoders, projectors, and LLMs fit together (e.g., LLaVA, Qwen-VL, InternVL, Whisper, HuBERT, Conformer).
  • •Hands-on with RLHF/alignment for multimodal: GRPO, DPO, and reward modeling across modalities.
  • •Experience shipping models to production, optimizing inference for latency and deploying end-to-end with solid engineering fundamentals.
Experience:3+ yearsAI safetyMultimodalVision-languageAudioSpeech
Skills:ExperimentationEngineering fundamentalsShipping to productionDistributed trainingDocumentation
Tech Stack:PyTorchDeepSpeedFSDPLLaVAQwen-VLInternVLAudio FlamingoOmni QwenAudio QwenWhisperHuBERTConformerMoESFTDPOGRPOReward modelingQuantizationDistillationBatching

Company Brief

White Circle
Builds AI-driven products and services to help businesses automate workflows, extract insights from data, and improve decision-making using machine learning and natural language processing technologies.
Industry: AI & Machine Learning
Website