Research Engineer, Audio and Speech

Decagon
San Francisco, New York
Workplace: OnsiteFull timeUSD 200,000 - 400,000 annuallyFunction: Research & Scientific (R&D)Experience: 2+ yearsSkills: ["End-to-end ownership","Shipping measurable improvements","High-impact technical decision-making"]

Build real-time audio and speech systems for conversational AI, taking models and agent harnesses from idea to production. Design streaming agent components for turn-taking and interruptions, train multimodal full-duplex models, and improve speech recognition, voice activity detection, endpointing, and speech generation. Create evaluations tied to accuracy, latency, naturalness, and task outcomes, and optimize inference for responsiveness, throughput, stability, and cost at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Decagon
Decagon
1 day ago

Research Engineer, Audio and Speech

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Build real-time audio and speech systems for conversational AI, taking models and agent harnesses from idea to production. Design streaming agent components for turn-taking and interruptions, train multimodal full-duplex models, and improve speech recognition, voice activity detection, endpointing, and speech generation. Create evaluations tied to accuracy, latency, naturalness, and task outcomes, and optimize inference for responsiveness, throughput, stability, and cost at scale.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design and build agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction.
  • •Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech.
  • •Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages.
  • •Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes.
  • •Optimize end-to-end inference for responsiveness, throughput, stability, and cost in partnership with Voice Platform and Infrastructure teams.

Pay and Benefits

Salary: USD 200,000 - 400,000 annually
Perks:Health InsuranceDentalVisionLife InsuranceRetirementParental LeaveWellness Stipend

Key Requirements

  • •2+ years of experience in speech, audio ML, multimodal ML, or production machine learning.
  • •Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models.
  • •Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio.
  • •Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in machine learning and signal processing.
  • •Demonstrated ability to take research ideas from prototype to reliable, measurable production impact.
Experience:2+ yearsSpeechAudio MLMultimodal MLProduction machine learningStreaming systemsReal-time audio
Skills:End-to-end ownershipShipping measurable improvementsHigh-impact technical decision-making
Tech Stack:PythonPyTorchAutoregressiveDiffusionFlow-matchingCodec-based models

Company Brief

Decagon
Builds conversational AI agents and a platform that automates customer support across chat, email, and voice, enabling brands to deliver concierge-level customer experiences at scale.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2023
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor