← Zurück zur Liste
Stelle

Applied ML Engineer

AI / ML Engineer • Remote • Vollzeit • Vereinigte Staaten USA

Deepgram is hiring an Applied ML Engineer to own the research-to-production pipeline — hardening training/evaluation workflows and building the deployment paths that turn research checkpoints into reliable production models.

Responsibilities

  • ▹Own the research-to-production pipeline, defining a repeatable path from a working result to a deployed, monitored service
  • ▹Partner with research scientists to productionize new models, turning experimental code into robust, tested workflows
  • ▹Build tooling and abstractions that let models move through training, evaluation, packaging, and deployment with minimal friction
  • ▹Design and own model release gates: automated evaluation, regression detection, quality/latency/throughput checks
  • ▹Optimize models and serving for production: efficient inference, batching, memory and latency tuning
  • ▹Strengthen the build/delivery layer across GPU compute and cloud environments
  • ▹Establish benchmarking and validation running consistently from development through production
  • ▹Build the feedback loop: instrument production model behavior and feed insights back to research

Requirements

  • ▹Strong software engineering fundamentals, proficiency in Python, production-quality tested ML code
  • ▹Hands-on experience taking ML models from research/prototype into production at scale
  • ▹Working understanding of the modern deep learning stack (e.g., PyTorch) and training/evaluating/serving large models
  • ▹Experience building ML pipelines and tooling: training orchestration, evaluation harnesses, model packaging, CI/CD
  • ▹Familiarity with serving and inference optimization: latency, throughput, batching, resource efficiency
  • ▹Comfort operating across distributed systems and GPU compute (cloud, bare metal, or both)
  • ▹Collaborative, builder mindset able to scope ambiguous problems and drive results

Nice to have

  • ▹Experience with the research-to-production handoff specifically
  • ▹Background in speech, audio, or other real-time/streaming ML domains
  • ▹Experience designing automated model evaluation and release-gating systems
  • ▹Familiarity with hybrid infrastructure spanning on-prem GPU clusters and cloud
  • ▹Experience with inference optimization (quantization, distillation, compilation, runtime tuning)
  • ▹Track record building internal platforms/tooling that measurably improved how a team ships models

Soft skills

Cares about reproducibility and evaluation rigor, not just shippingFluent bridge between research and engineeringTreats infrastructure and tooling as a product

About the company

Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT) and text-to-speech (TTS), and powering production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings on Deepgram's technology, including Twilio, Cloudflare, and Sierra. Backed by a recent Series C, Deepgram has processed over 50,000 years of audio and transcribed more than a trillion words.

Ähnliche Stellen