← Zurück zur Liste

Stelle
Site Reliability Engineer - AI & ML Infrastructure (Kubernetes, AWS & Terraform)
DevOps / SRE
• Remote
• Vollzeit
•
USA
Deepgram is looking for an experienced Site Reliability Engineer to build and operate its hybrid AI/ML infrastructure spanning AWS and on-premise bare metal data centers.
Responsibilities
- ▹Architect and maintain the core Kubernetes compute platform on AWS and on-premise
- ▹Develop and manage infrastructure using Infrastructure-as-Code principles with Terraform
- ▹Design and optimize AI/ML job scheduling by integrating Slurm with Kubernetes for GPU resources
- ▹Provision, manage, and maintain on-premise bare metal server infrastructure for GPU computing
- ▹Implement networking (CNI, service mesh) and storage (CSI, S3) solutions for hybrid workloads
- ▹Develop a comprehensive observability stack and build automation for operational tasks
- ▹Collaborate with AI researchers and ML engineers to understand infrastructure needs
- ▹Automate the lifecycle of single-tenant, managed deployments
Requirements
- ▹5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE)
- ▹Proven, hands-on experience building and managing production infrastructure with Terraform
- ▹Expert-level knowledge of Kubernetes architecture and operations at a large scale
- ▹Experience with HPC job schedulers, specifically Slurm, for managing GPU-intensive AI workloads
- ▹Experience managing bare metal infrastructure, including provisioning (PXE boot, MAAS) and lifecycle management
- ▹Strong scripting and automation skills (e.g. Python, Go, Bash)
Nice to have
- ▹Experience with CI/CD systems (e.g. GitLab CI, Jenkins, ArgoCD) and building developer tooling
- ▹Familiarity with FinOps principles and cloud cost optimization strategies
- ▹Knowledge of Kubernetes networking (Calico, Cilium) and storage (Ceph, Rook) solutions
- ▹Experience in a multi-region or hybrid cloud environment
Soft skills
Passionate about building platforms that empower developers and researchersEnjoys creating elegant, automated solutions for complex infrastructure challengesThrives on optimizing hybrid infrastructure for performance, cost, and reliability
About the company
Deepgram is the leading platform for the Voice AI economy, providing real-time speech-to-text and text-to-speech APIs and production-grade voice agents. Over 200,000 developers and 1,300+ organizations - including Twilio, Cloudflare, and Sierra - build on Deepgram, which is backed by a recent Series C round from leading global investors.
Ähnliche Stellen

Stelle
DevOps Engineer (5431)
Shield Ai
Azure Devops
+10
133 090–195 199 €/Jahr
brutto
🏢 Vor Ort
Dallas
🗣️ EN

Stelle
DevOps Engineer
Accenture Federal Services
Cloudformation
+6
105 319–199 724 €/Jahr
brutto
🏢 Vor Ort
Chantilly
🗣️ EN
Stelle
Director, Engineering – Release Engineering, DevOps & SRE
Ddn
AI/MLGithub Actions
+8
221 817–266 181 €/Jahr
brutto
🌍 Remote
North Carolina
🗣️ EN

Stelle
Site Reliability Engineer III (DBA)
Backblaze
+11
110 909–133 090 €/Jahr
brutto
🌍 Remote
🗣️ EN

Stelle
Cloud DevOps Engineer
Accenture Federal Services
+10
91 566–174 260 €/Jahr
brutto
🏢 Vor Ort
Huntsville
🗣️ EN

Stelle
Cloud DevOps Engineer
Accenture Federal Services
Argocd
+6
88 904–180 470 €/Jahr
brutto
🏢 Vor Ort
Los Angeles
🗣️ EN