Embedded AI Engineer, On-Device Models
Deepgram
California
Workplace: RemoteFull timeUSD 219,300 - 274,100 annuallyFunction: Data Science & Machine LearningSkills: ["Communication","Builder mindset"]Build and optimize Deepgram’s speech and conversational models for on-device, real-time inference on resource-constrained hardware. You’ll define on-device inference architecture across diverse processors and accelerators, apply techniques like quantization and pruning, write performance-critical C/C++/Rust runtime code for embedded/RTOS environments, and integrate with edge runtimes and vendor NPU/DSP toolchains. You’ll also own benchmarking, OTA deployment pipelines, and work with silicon partners and research to keep models edge-friendly.

