Take Deepgram's speech AI models and make them run fast, accurately, and efficiently on resource-constrained embedded and edge platforms — phones, earbuds, wearables, and appliances.
Responsibilities
- ▹Get Deepgram's speech and conversational models running on embedded and low-power consumer hardware, defining the architecture for on-device, real-time inference
- ▹Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation
- ▹Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and RTOS (FreeRTOS, Zephyr)
- ▹Integrate with edge inference runtimes and vendor NPU/DSP toolchains
- ▹Build on-device runtime plumbing: model packaging, deployment pipelines, OTA update mechanisms, lightweight telemetry
- ▹Establish repeatable benchmarking and validation across target hardware, measuring latency, accuracy, power, and memory footprint
- ▹Partner with silicon and device vendors on SDK integration and performance tuning
- ▹Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs
Requirements
- ▹Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices
- ▹Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments
- ▹Hands-on experience with model optimization for on-device deployment (quantization, pruning, knowledge distillation, architecture-specific compilation)
- ▹Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor NPU/DSP toolchains
- ▹Strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point arithmetic, power management
- ▹Experience working close to the metal: bare-metal or RTOS environments (FreeRTOS, Zephyr), embedded Linux, or microcontroller/edge SoC development
- ▹Strong communication skills and a builder mindset
Nice to have
- ▹Real-time audio processing on embedded platforms — DSP pipelines, codec optimization, wake-word detection
- ▹Depth in ML optimization techniques — custom quantization schemes, mixed-precision inference, neural architecture search for edge targets
- ▹Background in hardware evaluation and benchmarking
- ▹Experience shipping AI features in consumer products at scale
- ▹Familiarity with model compilation and optimization toolchains
- ▹Experience with secure, robust on-device deployment practices (code signing, encrypted model storage)
Soft skills
Builder mindsetCommunicationScoping ambiguous optimization problems independently
About the company
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time speech-to-text (STT) and text-to-speech (TTS) APIs and production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build on Deepgram, backed by a recent Series C.
Similar jobs

Job
Machine Learning Engineer
Qualcomm
AI/MLCppLlm
+4
$122,800–$184,200/yr
gross
🏢 On-site
San Diego
🗣️ EN

Job
Machine Learning Engineer - H/F/NB
Ubisoft
AI/MLCppData Science
+6
💰 Salary: not specified
🏢 On-site
Paris

Job
Machine Learning Engineer Automatisiertes Fahren (m/w/d)
ARRK Engineering GmbH
AI/MLCppData Science
+11
💰 Salary: not specified
🔀 Hybrid
München
🗣️ German

Job
Computer Vision Engineer, Space
Anduril
AI/MLCpp
+1
$166,000–$220,000/yr
gross
🏢 On-site
Washington
🗣️ EN

Job
Computer Vision Engineer, Space
Anduril
AI/MLCpp
+1
$166,000–$220,000/yr
gross
🏢 On-site
Costa Mesa
🗣️ EN

Job
Machine Learning Engineer, Detection and Tracking
Helsing
AI/MLCppMlflow
+1
💰 Salary: not specified
🏢 On-site
Washington
🗣️ EN
