Research Staff, Voice AI Foundations

Deepgram
California, Ann Arbor, San Francisco
Workplace: RemoteFull timeUSD 150,000 - 250,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Communication","Collaboration","Problem-solving","Critical thinking","Creative"]

Research Staff role focused on pioneering Latent Space Models for advanced Voice AI. Expected to develop next-gen neural codecs, steerable generative models for speech, and scalable embedding systems, while designing hardware-efficient training and inference for billions of conversations. You’ll work in a fast-paced, AI-first environment shaping foundational research and contributing to practical, scalable AI systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Deepgram
Deepgram
7 months ago

Research Staff, Voice AI Foundations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Research Staff role focused on pioneering Latent Space Models for advanced Voice AI. Expected to develop next-gen neural codecs, steerable generative models for speech, and scalable embedding systems, while designing hardware-efficient training and inference for billions of conversations. You’ll work in a fast-paced, AI-first environment shaping foundational research and contributing to practical, scalable AI systems.
Location: California, Ann Arbor, San Francisco
Workplace: Remote
Employment Type: Full time · Permanent
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Pioneer development of Latent Space Models (LSMs) to address data, scale, and cost challenges in robust voice AI systems.
  • •Research and develop solutions for problems including next-generation neural audio codecs, steerable generative models for diverse speech, and embedding systems with interpretable latent dimensions.
  • •Design model architectures, training schemes, and inference algorithms optimized for hardware at bare metal, enabling cost-efficient training on billion-hour datasets and real-time inference for hundreds of millions of conversations.
  • •Build and manage data pipelines to curate massive, diverse audio datasets suitable for training, evaluation, and benchmarking of new models.
  • •Collaborate across teams to validate theories with rigorous experiments, publish findings, and contribute to open-source efforts when applicable.

Pay and Benefits

Salary: USD 150,000 - 250,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kPaid LeaveLearning BudgetRemote WorkHome Office

Key Requirements

  • •Strong mathematical foundation in statistical learning theory, particularly in areas relevant to self-supervised and multimodal learning
  • •Deep expertise in foundation model architectures, with an understanding of how to scale training across multiple modalities
  • •Proven ability to bridge theory and practice—someone who can both derive novel mathematical formulations and implement them efficiently
  • •Demonstrated ability to build data pipelines that can process and curate massive datasets while maintaining quality and diversity
  • •Track record of designing controlled experiments that isolate the impact of architectural innovations and validate theoretical insights
Experience:Voice AISpeechMachine learning
Skills:CommunicationCollaborationProblem-solvingCritical thinkingCreative
Languages:English
Tech Stack:PythonMachine learningDeep learningLatent space modelsNeural networks

Company Brief

Deepgram
Deepgram builds a real-time Voice AI platform delivering speech-to-text, text-to-speech, and voice-agent APIs for developers and enterprises, focusing on low-latency, high-accuracy voice models and scalable deployment options.
Industry: API Platforms
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.5
WebsiteLinkedInGlassdoor