Research Engineer - Voice Intelligence

Pocket
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 300,000 annuallyFunction: Research & Scientific (R&D)Experience: 6-10 yearsSkills: ["Signal processing","Speech","Machine learning","Production engineering","Real-time systems"]

Open-ended research-minded software engineer to advance Pocket’s voice stack end-to-end—from signal capture to real-time experiences—focusing on diarization, VAD, noise robustness, transcription quality, and speaker identification. Collaborates with product/design to translate voice capabilities into user-facing features, and ships data-driven improvements with reliable latency at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pocket
Pocket
2 months ago

Research Engineer - Voice Intelligence

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 minutes agoStatus: Live

Job Summary

Open-ended research-minded software engineer to advance Pocket’s voice stack end-to-end—from signal capture to real-time experiences—focusing on diarization, VAD, noise robustness, transcription quality, and speaker identification. Collaborates with product/design to translate voice capabilities into user-facing features, and ships data-driven improvements with reliable latency at scale.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Invent and ship improvements to our voice stack: diarization, VAD, noise robustness, transcription quality, speaker ID, and post-processing.
  • •Build new voice-powered product capabilities: better memory, better meeting/voice summaries, better action extraction, better personalization.
  • •Run tight research loops: define metrics, build eval sets, iterate on models/algorithms, and productionize what works.
  • •Improve real-time performance and reliability: streaming pipelines, latency budgets, fallbacks, and graceful degradation.
  • •Partner closely with product + design to translate voice capabilities into features people feel immediately.

Pay and Benefits

Salary: USD 200,000 - 300,000 annually

Key Requirements

  • •6–10+ years of experience shipping production systems with strong technical ownership.
  • •Strong in applied ML/research engineering and can turn prototypes into robust product.
  • •Understand audio/voice fundamentals (signal processing basics helpful) and modern model/eval workflows.
  • •Care deeply about performance and reliability.
  • •Nice-to-haves: experience with streaming audio pipelines and real-time systems; experience with LLM post-processing, structured extraction, and evaluation; experience building tooling for labeling, dataset curation, and QA.
Experience:6-10 yearsVoiceMLAmbient intelligence
Skills:Signal processingSpeechMachine learningProduction engineeringReal-time systems
Tech Stack:MLSignal processingTranscriptionDiarizationVADSpeaker IDReal-time pipelines

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Pocket
Offers a consumer-facing mobile service called Pocket aimed at helping users manage money and digital wallets through a simple app and card, focusing on everyday spending and personal finance convenience.
Industry: Neobanking
Website