← Back to list

Job
Site Reliability Engineer - AI & ML Infrastructure (Kubernetes, AWS & Terraform)
DevOps / SRE
• Remote
• Full-time
•
USA
Deepgram is looking for an experienced Site Reliability Engineer to build and operate its hybrid AI/ML infrastructure spanning AWS and on-premise bare metal data centers.
Responsibilities
- ▹Architect and maintain the core Kubernetes compute platform on AWS and on-premise
- ▹Develop and manage infrastructure using Infrastructure-as-Code principles with Terraform
- ▹Design and optimize AI/ML job scheduling by integrating Slurm with Kubernetes for GPU resources
- ▹Provision, manage, and maintain on-premise bare metal server infrastructure for GPU computing
- ▹Implement networking (CNI, service mesh) and storage (CSI, S3) solutions for hybrid workloads
- ▹Develop a comprehensive observability stack and build automation for operational tasks
- ▹Collaborate with AI researchers and ML engineers to understand infrastructure needs
- ▹Automate the lifecycle of single-tenant, managed deployments
Requirements
- ▹5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE)
- ▹Proven, hands-on experience building and managing production infrastructure with Terraform
- ▹Expert-level knowledge of Kubernetes architecture and operations at a large scale
- ▹Experience with HPC job schedulers, specifically Slurm, for managing GPU-intensive AI workloads
- ▹Experience managing bare metal infrastructure, including provisioning (PXE boot, MAAS) and lifecycle management
- ▹Strong scripting and automation skills (e.g. Python, Go, Bash)
Nice to have
- ▹Experience with CI/CD systems (e.g. GitLab CI, Jenkins, ArgoCD) and building developer tooling
- ▹Familiarity with FinOps principles and cloud cost optimization strategies
- ▹Knowledge of Kubernetes networking (Calico, Cilium) and storage (Ceph, Rook) solutions
- ▹Experience in a multi-region or hybrid cloud environment
Soft skills
Passionate about building platforms that empower developers and researchersEnjoys creating elegant, automated solutions for complex infrastructure challengesThrives on optimizing hybrid infrastructure for performance, cost, and reliability
About the company
Deepgram is the leading platform for the Voice AI economy, providing real-time speech-to-text and text-to-speech APIs and production-grade voice agents. Over 200,000 developers and 1,300+ organizations - including Twilio, Cloudflare, and Sierra - build on Deepgram, which is backed by a recent Series C round from leading global investors.
Similar jobs

Job
DevOps Engineer (5431)
Shield Ai
Azure Devops
+10
$150,000–$220,000/yr
gross
🏢 On-site
Dallas
🗣️ EN

Job
DevOps Engineer
Accenture Federal Services
Cloudformation
+6
$118,700–$225,100/yr
gross
🏢 On-site
Chantilly
🗣️ EN
Job
Director, Engineering – Release Engineering, DevOps & SRE
Ddn
AI/MLGithub Actions
+8
$250,000–$300,000/yr
gross
🌍 Remote
North Carolina
🗣️ EN

Job
Site Reliability Engineer III (DBA)
Backblaze
+11
$125,000–$150,000/yr
gross
🌍 Remote
🗣️ EN

Job
Cloud DevOps Engineer
Accenture Federal Services
+10
$103,200–$196,400/yr
gross
🏢 On-site
Huntsville
🗣️ EN

Job
Cloud DevOps Engineer
Accenture Federal Services
Argocd
+6
$100,200–$203,400/yr
gross
🏢 On-site
Los Angeles
🗣️ EN