Applied ML Engineer

Deepgram
United States
Workplace: RemoteFull timeUSD 150,000 - 220,000 annuallyFunction: Data Science & Machine LearningSkills: ["Builder mindset","Collaboration","Adaptability","Continuous learning"]

Own and streamline the research-to-production pipeline for voice AI models. Partner with research scientists to convert experimental training and evaluation code into robust, reproducible workflows, then build tooling, model release gates, and benchmarking/validation to ship at scale. Optimize training and inference for production latency, throughput, and resource efficiency across hybrid GPU and cloud infrastructure, while instrumenting a feedback loop that accelerates the next model iteration.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Deepgram
Deepgram
2 months ago

Applied ML Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Own and streamline the research-to-production pipeline for voice AI models. Partner with research scientists to convert experimental training and evaluation code into robust, reproducible workflows, then build tooling, model release gates, and benchmarking/validation to ship at scale. Optimize training and inference for production latency, throughput, and resource efficiency across hybrid GPU and cloud infrastructure, while instrumenting a feedback loop that accelerates the next model iteration.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Own the research-to-production pipeline by turning research checkpoints into deployed, monitored, scalable production models.
  • •Partner with research scientists to productionize new models by translating experimental training/evaluation code into robust, reproducible workflows.
  • •Build and extend tooling and abstractions for moving models through training, evaluation, packaging, and deployment with reproducibility.
  • •Design and own automated model release gates using evaluation, regression detection, and quality/latency/throughput checks.
  • •Optimize model serving and operational readiness across hybrid infrastructure, including instrumentation for a feedback loop back to research.

Pay and Benefits

Salary: USD 150,000 - 220,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Strong software engineering fundamentals with proficiency in Python and experience writing production-quality, well-tested ML code.
  • •Hands-on experience taking ML models from research/prototype stage into production at scale, including shipping and operating them.
  • •Working understanding of the modern deep learning stack (e.g., PyTorch) and the realities of training, evaluating, and serving large models.
  • •Experience building ML pipelines and tooling (training orchestration, evaluation harnesses, model packaging, deployment, or CI/CD for models).
  • •Familiarity with serving and inference optimization (latency, throughput, batching, and resource efficiency) and operating across distributed systems and GPU compute.
Experience:Voice AISpeechAudioReal-time/streaming MLDeep learning
Skills:Builder mindsetCollaborationAdaptabilityContinuous learning
Tech Stack:PythonPyTorchCI/CDTraining orchestrationEvaluation harnessesModel packagingDeploymentDistributed systemsGPU computeQuantizationDistillationCompilationRuntime tuning

Company Brief

Deepgram
Deepgram builds a real-time Voice AI platform delivering speech-to-text, text-to-speech, and voice-agent APIs for developers and enterprises, focusing on low-latency, high-accuracy voice models and scalable deployment options.
Industry: API Platforms
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.5
WebsiteLinkedInGlassdoor