A leadership role owning and scaling Runpod's core cloud and bare-metal environments — SRE, global networking, HPC networks, and distributed storage — in close partnership with Product Engineering and GTM leadership.
Responsibilities
- ▹Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage
- ▹Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation
- ▹Oversee the design, scaling, and operation of Runpod's global network backbone and ultra-low-latency HPC cluster networks
- ▹Drive implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) for massive, multi-node GPU training workloads
- ▹Direct the architecture and performance tuning of highly scalable, distributed storage systems
- ▹Hire, mentor, and grow highly technical engineering managers and senior ICs
- ▹Partner with Program Management and Product to forecast capacity requirements and shape technical roadmaps
- ▹Drive measurable improvements in deployment frequency, MTTR, IaC coverage, and system uptime
- ▹Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters
- ▹Coordinate cleanly with product delivery and platform teams
Requirements
- ▹7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments
- ▹8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms
- ▹Proven hands-on background or strong architectural understanding of ultra-low latency networking: InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing (BGP)
- ▹Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF)
- ▹Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks
- ▹Experience building culture, accountability, and momentum across distributed remote-first technical teams
- ▹Clear written and verbal communication, strong stakeholder management, calm and decisive leadership during high-stakes incidents
- ▹Successful completion of a background check
Nice to have
- ▹Direct experience architecting infrastructure optimized for massive GPU clusters and AI/ML workloads
- ▹Deep understanding of hardware architectures, GPU interconnects (NVLink), and datacenter topology
- ▹Track record of scaling infrastructure teams in hyper-growth startup environments
- ▹Open-source contributions or active recognition within the infrastructure, networking, or Kubernetes communities
Soft skills
Calm, decisive leadership during high-stakes operational incidentsStrong stakeholder managementClear written and verbal communication
What we offer
- ▹Competitive base pay of $225,000–$325,000
- ▹Meaningful equity in a fast-growing company
- ▹Generous medical, dental & vision plans
- ▹Flexible PTO
- ▹$1,200 home office & equipment stipend
- ▹Remote-first, Slack-driven collaborative culture
About the company
Runpod is the AI Developer Cloud, used by more than one million developers to experiment, train, and scale AI. The company closed a $100M Series A in June 2026 and is building the platform the next generation of developers will depend on.
Similar jobs

Job
Platform Engineer, Model Shaping
Together AI
AI/MLArgocd
+10
$200,000–$290,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
Founding Cloud Infrastructure Engineer
Clera
AI/MLCloudformationDatadog
+5
💰 Salary: not specified
🏢 On-site
San Mateo
🗣️ EN

Job
Platform Engineer
Clera
AI/ML
+1
$150,000–$250,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
Security Infrastructure Engineer
Tailscale
$163,000–$204,000/yr
gross
🌍 Remote
🗣️ EN

Job
AI Platform Engineer (m/w/d)
Sopra Steria
AI/ML
+12
💰 Salary: not specified
🏢 On-site
🗣️ German

Job
Platform Engineer (P4)
Appsilon
AI/MLArgocd
+10
💰 Salary: not specified
🏢 On-site
Wasaw
🗣️ EN
