Deepgram is hiring an Applied ML Engineer to own the research-to-production pipeline — hardening training/evaluation workflows and building the deployment paths that turn research checkpoints into reliable production models.
Responsibilities
- ▹Own the research-to-production pipeline, defining a repeatable path from a working result to a deployed, monitored service
- ▹Partner with research scientists to productionize new models, turning experimental code into robust, tested workflows
- ▹Build tooling and abstractions that let models move through training, evaluation, packaging, and deployment with minimal friction
- ▹Design and own model release gates: automated evaluation, regression detection, quality/latency/throughput checks
- ▹Optimize models and serving for production: efficient inference, batching, memory and latency tuning
- ▹Strengthen the build/delivery layer across GPU compute and cloud environments
- ▹Establish benchmarking and validation running consistently from development through production
- ▹Build the feedback loop: instrument production model behavior and feed insights back to research
Requirements
- ▹Strong software engineering fundamentals, proficiency in Python, production-quality tested ML code
- ▹Hands-on experience taking ML models from research/prototype into production at scale
- ▹Working understanding of the modern deep learning stack (e.g., PyTorch) and training/evaluating/serving large models
- ▹Experience building ML pipelines and tooling: training orchestration, evaluation harnesses, model packaging, CI/CD
- ▹Familiarity with serving and inference optimization: latency, throughput, batching, resource efficiency
- ▹Comfort operating across distributed systems and GPU compute (cloud, bare metal, or both)
- ▹Collaborative, builder mindset able to scope ambiguous problems and drive results
Nice to have
- ▹Experience with the research-to-production handoff specifically
- ▹Background in speech, audio, or other real-time/streaming ML domains
- ▹Experience designing automated model evaluation and release-gating systems
- ▹Familiarity with hybrid infrastructure spanning on-prem GPU clusters and cloud
- ▹Experience with inference optimization (quantization, distillation, compilation, runtime tuning)
- ▹Track record building internal platforms/tooling that measurably improved how a team ships models
Soft skills
Cares about reproducibility and evaluation rigor, not just shippingFluent bridge between research and engineeringTreats infrastructure and tooling as a product
About the company
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT) and text-to-speech (TTS), and powering production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings on Deepgram's technology, including Twilio, Cloudflare, and Sierra. Backed by a recent Series C, Deepgram has processed over 50,000 years of audio and transcribed more than a trillion words.
Similar jobs

Job
Applied Scientist / Machine Learning Engineer
Wayve
AI/MLSQL
💰 Salary: not specified
🔀 Hybrid
Sunnyvale
🗣️ EN

Job
Founding AI/ML Engineer
Clera
AI/MLLlm
+3
$130,000–$170,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
Machine Learning Engineer, Application Software
Wayve
AI/MLCpp
$311,850–$350,625/yr
gross
🔀 Hybrid
Sunnyvale
🗣️ EN

Job
Computer Vision Systems Engineer
Qualcomm
AI/MLCpp
$155,400–$233,200/yr
gross
🏢 On-site
San Diego
🗣️ EN
Job
AI/ML Engineer
Cinder
AI/MLDatabricksData Science
+6
$220,000–$260,000/yr
gross
🔀 Hybrid
New York
🗣️ EN

Job
AI / ML Engineer
Accenture Federal Services
AI/MLBigqueryData Science
+9
$103,200–$196,400/yr
gross
🏢 On-site
Arlington
🗣️ EN
