Staff Research Engineer - Multimodal Generative Modelling

Synthesia
London
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Skills: ["Prototype quickly","Iterate efficiently"]

Build and ship multimodal generative models for real-time interactive voice-video experiences. Join the Voice team within a 40+ person R&D org to define a research roadmap, propose novel text-and-voice architectures, and develop low-latency streaming conversational systems. Work end-to-end from pretraining and post-training (e.g., DPO, fine-tuning, distillation) through dataset curation, evaluation metrics, integration/testing (neural codecs, diffusion, flow-matching), and production deployment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Synthesia
Synthesia
1 month ago

Staff Research Engineer - Multimodal Generative Modelling

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Build and ship multimodal generative models for real-time interactive voice-video experiences. Join the Voice team within a 40+ person R&D org to define a research roadmap, propose novel text-and-voice architectures, and develop low-latency streaming conversational systems. Work end-to-end from pretraining and post-training (e.g., DPO, fine-tuning, distillation) through dataset curation, evaluation metrics, integration/testing (neural codecs, diffusion, flow-matching), and production deployment.
Location: London
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Shape the roadmap to create new model capabilities and unlock customer-facing functionality across short and long time horizons.
  • •Propose and design novel multimodal system architectures (especially text and voice) and develop streaming/conversational systems for low-latency voice-video synthesis.
  • •Implement and integrate architectures (neural codecs, diffusion, flow-matching), and define evaluation metrics for latency-aware, interaction-based conversational performance.
  • •Track and incorporate state-of-the-art research across audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.
  • •Curate datasets and lead post-training initiatives (DPO, fine-tuning, distillation) to bring models to shipping quality and deploy to production with optimized runtime.

Key Requirements

  • •Strong understanding of generative modelling for sequential or multimodal data.
  • •Hands-on experience with large language models or transformer-based architectures.
  • •High proficiency in PyTorch, including distributed training and model optimization.
  • •Solid grasp of time-series modeling and tokenization, preferably for audio, speech, or video.
  • •Proven experience training deep learning models end-to-end, from data preparation through evaluation.
Experience:Generative modelsMultimodal systemsLLMs
Skills:Prototype quicklyIterate efficiently
Tech Stack:PyTorchDistributed trainingTokenizationLarge language modelsTransformer-based architecturesNeural codecsDiffusion modelsFlow-matching modelsAutoregressive decodersAudio-visual diffusionMultimodal LLMs

Company Brief

Synthesia
Provides an AI video generation platform that creates realistic synthetic presenters and video content from text, enabling enterprises to produce scalable training, marketing, and communications videos without cameras or actors.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: London, United Kingdom
Founded: 2017
WebsiteLinkedIn