Deepgram is hiring an ML Ops Infrastructure Engineer to own the bridge between research and production — building the CI/CD pipelines, deployment systems, and monitoring that ship models safely at scale.
Responsibilities
- ▹Design and build CI/CD pipelines tailored for ML model development, validation, and deployment
- ▹Architect and maintain model deployment pipelines from research to staging to production
- ▹Build A/B testing infrastructure for controlled rollouts and real-world performance measurement
- ▹Implement monitoring for model performance: accuracy metrics, latency, drift detection, regression alerts
- ▹Develop automated retraining pipelines triggered by data changes or performance degradation
- ▹Create build/test environments that mirror production for high-fidelity researcher feedback
- ▹Establish model versioning, artifact management, and rollback capabilities
- ▹Collaborate with research engineers to define and enforce model quality gates
- ▹Build observability dashboards for real-time model health insight
- ▹Optimize model serving infrastructure for latency, throughput, and cost efficiency
Requirements
- ▹4+ years of experience in MLOps, DevOps, or infrastructure engineering focused on ML systems
- ▹Strong proficiency in Python, building automation and tooling for ML workflows
- ▹Deep experience with CI/CD systems and pipelines for software and model delivery
- ▹Hands-on experience with Docker and Kubernetes for containerized workload management
- ▹Practical experience deploying and serving ML models in production
- ▹Familiarity with model evaluation, validation, and QA processes
- ▹Understanding of monitoring and observability principles for ML systems
- ▹Strong problem-solving skills with a bias toward automation
Nice to have
- ▹Experience with model serving frameworks such as NVIDIA Triton, TensorRT, or ONNX Runtime
- ▹Background in speech, audio, or real-time media ML systems
- ▹Experience with IaC tools such as Terraform or Pulumi
- ▹Hands-on experience with monitoring/observability stacks (Prometheus, Grafana, Datadog)
- ▹Familiarity with GPU-accelerated inference optimization and profiling
- ▹Experience with feature stores, data versioning, or ML metadata management
- ▹Knowledge of canary deployment strategies and progressive delivery for ML models
Soft skills
Bias toward automation over manual processesBelieves great infrastructure turns research into customer valueEnjoys designing automated, self-healing systems
About the company
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT) and text-to-speech (TTS), and powering production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings on Deepgram's technology, including Twilio, Cloudflare, and Sierra. Backed by a recent Series C, Deepgram has processed over 50,000 years of audio and transcribed more than a trillion words.
Ähnliche Stellen

Stelle
Application/DevOps Platform Engineer
Qualcomm
ArgocdCloudformationDatadog
+10
💰 Gehalt: keine Angabe
🏢 Vor Ort
San Diego
🗣️ EN
Stelle
Platform Engineer - Commercial Services Tools
Veeva
Cloudformation
+10
79 854–133 090 €/Jahr
brutto
🏢 Vor Ort
Massachusetts - Boston
🗣️ EN

Stelle
Platform Infrastructure Engineer
Menlosecurity
+5
💰 Gehalt: keine Angabe
🌍 Remote
US - Distributed
🗣️ EN

Stelle
AI Native Founding Product and Platform Engineer
SGS
Github Actions
+9
💰 Gehalt: keine Angabe
🔀 Hybrid
Bloomfield
🗣️ EN

Stelle
AI Native Founding Product and Platform Engineer
SGS
Github Actions
+9
💰 Gehalt: keine Angabe
🔀 Hybrid
Seattle
🗣️ EN

Stelle
AI Native Founding Product and Platform Engineer
SGS
Github Actions
+9
💰 Gehalt: keine Angabe
🔀 Hybrid
Hayward
🗣️ EN
