AI Researcher (Multimodal Audio/Video Generation)

Tavus
San Francisco
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Experience: 2-3 yearsEducation: phdSkills: ["PyTorch","GPU-accelerated workflows","Communication","Mentoring"]

Lead research in audio-visual avatar generation, pushing multimodal generative models toward realistic, conversational AI. Design diffusion-based, long-video and audio-visual systems synced with conversation flow, translating research into production with Applied ML and engineering. Mentor researchers, publish impactful work, and shape how humans interact with AI Humans at Tavus.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tavus
Tavus
4 months ago

AI Researcher (Multimodal Audio/Video Generation)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live
Reposted: similar role first listed 11 months ago

Job Summary

Lead research in audio-visual avatar generation, pushing multimodal generative models toward realistic, conversational AI. Design diffusion-based, long-video and audio-visual systems synced with conversation flow, translating research into production with Applied ML and engineering. Mentor researchers, publish impactful work, and shape how humans interact with AI Humans at Tavus.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Lead research efforts on audio-visual generation for avatars (Neural Avatars, Talking-Heads), with a focus on conversational settings.
  • •Design models that are coupled with conversation flow — capturing and generating verbal + non-verbal signals in sync.
  • •Drive innovation in diffusion models, long-video generation, and audio-visual modeling.
  • •Translate research into production by partnering with Applied ML and engineering.
  • •Mentor researchers, set research directions, and publish impactful work.

Key Requirements

  • •PhD or equivalent research experience
  • •2–3+ years hands-on experience applying generative models at scale
  • •Expertise in diffusion models and awareness of efficiency techniques
  • •Experience in multimodal generation—video, audio, and language
  • •Proven track record of publications in top-tier venues and experience leading research activities
Experience:2-3 yearsMultimodalAI researchGenerative models
Education:PhD / Doctorate
Skills:PyTorchGPU-accelerated workflowsCommunicationMentoring
Languages:English
Tech Stack:PyTorchGPUDiffusion modelsMultimodal generationVideo generationAudio generation

Company Brief

Tavus
Builds an AI-powered platform for creating personalized video at scale, enabling businesses to generate custom video messages tailored to individual recipients for marketing, sales outreach, and customer engagement.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2022
WebsiteLinkedIn