← Back to list
Job · Architect

Systems Architect AI/ML Infrastructure

MLOps Engineer • Architect • Remote • Full-time • United States USA

Deepgram, the leading Voice AI platform, is hiring a senior Systems Architect to own the end-to-end infrastructure spanning bare metal GPU clusters, multi-cloud deployments, and global edge presence, serving both real-time inference and large-scale model training.

Responsibilities

  • ▹Define and drive end-to-end infrastructure architecture for AI/ML workloads across production inference and research training
  • ▹Design multi-cloud and hybrid infrastructure strategies balancing performance, reliability, cost, and vendor flexibility
  • ▹Architect compute orchestration systems that efficiently schedule and manage GPU and CPU workloads
  • ▹Design storage architectures for massive datasets, from high-throughput training pipelines to low-latency model serving
  • ▹Lead capacity planning across all infrastructure dimensions, modeling growth
  • ▹Drive cost optimization and FinOps practices, identifying opportunities to reduce infrastructure spend
  • ▹Design burstable, elastic training infrastructure that scales up for large training runs and down to minimize idle cost
  • ▹Establish architectural standards, design review processes, and technical documentation practices
  • ▹Collaborate with engineering leadership to align infrastructure strategy with business objectives

Requirements

  • ▹7+ years of experience in infrastructure engineering, systems architecture, or a senior technical role focused on large-scale infrastructure
  • ▹Proven experience designing multi-cloud architectures spanning AWS and at least one other major cloud provider or on-premises environment
  • ▹Deep expertise in storage system design — block, object, and file storage — with performance tuning for large-scale data workloads
  • ▹Strong experience with compute orchestration using Kubernetes
  • ▹Hands-on experience with GPU infrastructure — procurement, cluster design, driver and runtime management
  • ▹Track record of capacity planning and infrastructure scaling for high-growth environments
  • ▹Ability to communicate complex architectural decisions clearly to technical and non-technical stakeholders
  • ▹Strong understanding of networking fundamentals as they relate to infrastructure architecture

Nice to have

  • ▹Direct experience architecting infrastructure for ML training workloads (distributed training, large dataset management, experiment infrastructure)
  • ▹Background in cost optimization and FinOps for large-scale cloud and bare metal infrastructure
  • ▹Experience operating and managing bare metal infrastructure in colocation facilities
  • ▹Expertise in network architecture design, including high-bandwidth GPU interconnects and global traffic routing
  • ▹Experience with infrastructure modeling and simulation for capacity planning
  • ▹Familiarity with Slurm, Ray, or other HPC/ML job scheduling systems
  • ▹Understanding of power, cooling, and physical infrastructure considerations for GPU-dense deployments

Soft skills

Systems thinking — naturally seeing the connections between compute, storage, and networkAbility to operate at a strategic level while staying technically deep enough to validate designs and debug issuesOpenness to experimentation and rapid adaptation in an AI-first, fast-changing environment

About the company

Deepgram is the leading platform in the fast-growing, trillion-dollar Voice AI economy, providing real-time speech-to-text (STT) and text-to-speech (TTS) APIs and production-grade voice agents. More than 200,000 developers and 1,300+ organizations build on it, including Twilio, Cloudflare, and Sierra; backed by a recent Series C round, the company has processed over 50,000 years of audio to date.

Similar jobs