Director of Research, Text to Speech

Deepgram
San Francisco, Ann Arbor
Workplace: RemoteFull timeUSD 213,000 - 328,300 annuallyFunction: Research & Scientific (R&D)Skills: []

Own the end-to-end Text-to-Speech research program, from roadmap and technical bets to experiments, training strategy, and evaluation. Lead a hands-on team that advances neural audio modeling, prosody/expressiveness, controllability, multilingual and voice consistency, and improves inference performance. Build benchmarking that combines automated metrics with human perceptual assessment, and partner with engineering and product to drive ship-ready production gains.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Deepgram
Deepgram
3 days ago

Director of Research, Text to Speech

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Own the end-to-end Text-to-Speech research program, from roadmap and technical bets to experiments, training strategy, and evaluation. Lead a hands-on team that advances neural audio modeling, prosody/expressiveness, controllability, multilingual and voice consistency, and improves inference performance. Build benchmarking that combines automated metrics with human perceptual assessment, and partner with engineering and product to drive ship-ready production gains.
Location: San Francisco, Ann Arbor
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Director level

Key Responsibilities

  • •Own the TTS research and model roadmap, deciding technical directions that can materially improve speech-generation quality.
  • •Drive advances across neural audio modeling, prosody/expressiveness, controllability, multilingual speech, voice identity/consistency, data & training strategy, post-training, and inference performance.
  • •Stay deeply technical by reviewing research, challenging assumptions, designing experiments, diagnosing failures, and tackling the highest-leverage problems.
  • •Build evaluation and benchmarking that explains why models improve using automated metrics plus human perceptual assessment.
  • •Lead a mix of individual contributors and tech lead managers, hiring/developing talent while maintaining an exceptionally high technical bar and pushing decisions down to teams.

Pay and Benefits

Salary: USD 213,000 - 328,300 annually
Equity and Bonus:Equity

Key Requirements

  • •Deep expertise in modern TTS, speech generation, or audio generative modeling, including personally training and improving large-scale neural models.
  • •Command of the speech-generation stack and core challenges in naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost.
  • •Experience setting research direction under uncertainty—prioritizing experiments, allocating compute/researcher time, and stopping approaches that don’t work.
  • •Experience leading researchers/research engineers through technical leaders while staying technically influential yourself.
  • •Ability to make complex technical tradeoffs legible to product, engineering, and executive audiences.
Experience:Voice AISpeech generationText to speechGenerative audioAI research

Company Brief

Deepgram
Deepgram builds a real-time Voice AI platform delivering speech-to-text, text-to-speech, and voice-agent APIs for developers and enterprises, focusing on low-latency, high-accuracy voice models and scalable deployment options.
Industry: API Platforms
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.5
WebsiteLinkedInGlassdoor