← Zurück zur Liste
Stelle

Embedded AI Engineer, On-Device Models

AI / ML Engineer • Remote • Vollzeit • Vereinigte Staaten USA

Take Deepgram's speech AI models and make them run fast, accurately, and efficiently on resource-constrained embedded and edge platforms — phones, earbuds, wearables, and appliances.

Responsibilities

  • ▹Get Deepgram's speech and conversational models running on embedded and low-power consumer hardware, defining the architecture for on-device, real-time inference
  • ▹Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation
  • ▹Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and RTOS (FreeRTOS, Zephyr)
  • ▹Integrate with edge inference runtimes and vendor NPU/DSP toolchains
  • ▹Build on-device runtime plumbing: model packaging, deployment pipelines, OTA update mechanisms, lightweight telemetry
  • ▹Establish repeatable benchmarking and validation across target hardware, measuring latency, accuracy, power, and memory footprint
  • ▹Partner with silicon and device vendors on SDK integration and performance tuning
  • ▹Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs

Requirements

  • ▹Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices
  • ▹Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments
  • ▹Hands-on experience with model optimization for on-device deployment (quantization, pruning, knowledge distillation, architecture-specific compilation)
  • ▹Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor NPU/DSP toolchains
  • ▹Strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point arithmetic, power management
  • ▹Experience working close to the metal: bare-metal or RTOS environments (FreeRTOS, Zephyr), embedded Linux, or microcontroller/edge SoC development
  • ▹Strong communication skills and a builder mindset

Nice to have

  • ▹Real-time audio processing on embedded platforms — DSP pipelines, codec optimization, wake-word detection
  • ▹Depth in ML optimization techniques — custom quantization schemes, mixed-precision inference, neural architecture search for edge targets
  • ▹Background in hardware evaluation and benchmarking
  • ▹Experience shipping AI features in consumer products at scale
  • ▹Familiarity with model compilation and optimization toolchains
  • ▹Experience with secure, robust on-device deployment practices (code signing, encrypted model storage)

Soft skills

Builder mindsetCommunicationScoping ambiguous optimization problems independently

About the company

Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time speech-to-text (STT) and text-to-speech (TTS) APIs and production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build on Deepgram, backed by a recent Series C.

Ähnliche Stellen