Senior Software Engineer - Model Evaluation & AI Systems

Deepgram
California
Workplace: RemoteFull timeUSD 180,000 - 240,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Analytical skills","Communication","Taking charge","Quality-focused engineering"]

Build and scale evaluation and quality-assurance systems that validate speech-to-text, text-to-speech, and LLM/RAG/agent/multimodal models before they ship. Define evaluation methodology, pass/fail gates, and monitoring; design automated pipelines for batch and streaming (e.g., WER and hallucination detection); and integrate quality checks into CI/CD. Partner with Research, model training, inference, and product teams to deliver trusted signals that protect customer experience.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Deepgram
Deepgram
1 month ago

Senior Software Engineer - Model Evaluation & AI Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Build and scale evaluation and quality-assurance systems that validate speech-to-text, text-to-speech, and LLM/RAG/agent/multimodal models before they ship. Define evaluation methodology, pass/fail gates, and monitoring; design automated pipelines for batch and streaming (e.g., WER and hallucination detection); and integrate quality checks into CI/CD. Partner with Research, model training, inference, and product teams to deliver trusted signals that protect customer experience.
Location: California
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Define and build evaluation methodologies across speech-to-text, text-to-speech, and emerging LLM/RAG/agent/multimodal systems.
  • •Design, build, and maintain automated evaluation pipelines for batch and streaming with focus on correctness and reproducibility.
  • •Build scalable evaluation infrastructure (harnesses, orchestration, result aggregation), including runs against production models and GPU clusters when needed.
  • •Translate research benchmarks and expected metrics into automated, enforceable pass/fail gates.
  • •Operate canaries and continuous monitoring to detect production quality regressions before they impact customers.

Pay and Benefits

Salary: USD 180,000 - 240,000 annually
Equity and Bonus:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, AI, Applied Math, or related field, or equivalent experience.
  • •5+ years of professional software or QA engineering experience shipping test infrastructure or evaluation systems.
  • •Backend/scripting experience in Python, Rust, Go, or a similar language.
  • •Experience designing and building automated test pipelines, evaluation frameworks, or data-processing systems.
  • •Strong analytical skills to reason about metrics, thresholds, and statistical variation to distinguish real regressions from noise.
Experience:AIMachine learningSpeech recognitionVoice AI
Education:Bachelor's in Computer Science, AI, Applied Math, or related field
Skills:Analytical skillsCommunicationTaking chargeQuality-focused engineering
Tech Stack:PythonRustGoReact NativeLLMsRAGCI/CDGrafanaGPU clusters

Company Brief

Deepgram
Deepgram builds a real-time Voice AI platform delivering speech-to-text, text-to-speech, and voice-agent APIs for developers and enterprises, focusing on low-latency, high-accuracy voice models and scalable deployment options.
Industry: API Platforms
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.5
WebsiteLinkedInGlassdoor