← Back to list
Job

Director of Infrastructure Engineering

Platform Engineer • Remote • Full-time United States USA

A leadership role owning and scaling Runpod's core cloud and bare-metal environments — SRE, global networking, HPC networks, and distributed storage — in close partnership with Product Engineering and GTM leadership.

Responsibilities

  • Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage
  • Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation
  • Oversee the design, scaling, and operation of Runpod's global network backbone and ultra-low-latency HPC cluster networks
  • Drive implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) for massive, multi-node GPU training workloads
  • Direct the architecture and performance tuning of highly scalable, distributed storage systems
  • Hire, mentor, and grow highly technical engineering managers and senior ICs
  • Partner with Program Management and Product to forecast capacity requirements and shape technical roadmaps
  • Drive measurable improvements in deployment frequency, MTTR, IaC coverage, and system uptime
  • Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters
  • Coordinate cleanly with product delivery and platform teams

Requirements

  • 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments
  • 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms
  • Proven hands-on background or strong architectural understanding of ultra-low latency networking: InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing (BGP)
  • Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF)
  • Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks
  • Experience building culture, accountability, and momentum across distributed remote-first technical teams
  • Clear written and verbal communication, strong stakeholder management, calm and decisive leadership during high-stakes incidents
  • Successful completion of a background check

Nice to have

  • Direct experience architecting infrastructure optimized for massive GPU clusters and AI/ML workloads
  • Deep understanding of hardware architectures, GPU interconnects (NVLink), and datacenter topology
  • Track record of scaling infrastructure teams in hyper-growth startup environments
  • Open-source contributions or active recognition within the infrastructure, networking, or Kubernetes communities

Soft skills

Calm, decisive leadership during high-stakes operational incidentsStrong stakeholder managementClear written and verbal communication

What we offer

  • Competitive base pay of $225,000–$325,000
  • Meaningful equity in a fast-growing company
  • Generous medical, dental & vision plans
  • Flexible PTO
  • $1,200 home office & equipment stipend
  • Remote-first, Slack-driven collaborative culture

About the company

Runpod is the AI Developer Cloud, used by more than one million developers to experiment, train, and scale AI. The company closed a $100M Series A in June 2026 and is building the platform the next generation of developers will depend on.

Similar jobs