Runpod is looking for a Datacenter Infrastructure Specialist to own the operational health and lifecycle of its global high-density GPU fleet, acting as the technical link between hardware partners and internal engineering. The role blends HPC systems engineering, network troubleshooting and process automation.
Responsibilities
- ▹Validate new hardware and ensure partner deployments meet Runpod's specifications for distributed AI/ML workloads
- ▹Monitor fleet health to identify performance degradation, audit downtime and provide the technical data needed to protect customer SLAs
- ▹Work with LLMs and AI agents to automate network triage and generate dynamic runbooks for the fleet
- ▹Coordinate technical incident communications and translate outages into actionable resolutions
- ▹Support the growth of infrastructure partners
Requirements
- ▹3-5 years of experience in infrastructure operations, systems reliability or datacenter engineering
- ▹Strong proficiency in standard datacenter networking and performance troubleshooting
- ▹Hands-on experience with the NVIDIA software stack (driver installation, performance utilities) and an understanding of multi-node performance tuning
- ▹Solid Linux system administration skills and experience with containerization (Docker)
Nice to have
- ▹Exposure to RDMA, InfiniBand or RoCE (highly preferred)
Soft skills
CommunicationOwnership
What we offer
- ▹Remote work (remote-first team)
About the company
Runpod is the AI Developer Cloud, used by more than one million developers to experiment, train, fine-tune, deploy and scale AI. The platform has processed more than 20 billion inference requests, and the company closed a $100M Series A in June 2026. A small, remote-first team.
Ähnliche Stellen

Stelle
Hardware Systems Engineer
Cloudflare
Bitbucket
+4
💰 Gehalt: keine Angabe
🏢 Vor Ort
In-Office
🗣️ EN

Stelle
Response Engineer - Cloudflare Managed Defense Center (CMDC)
Cloudflare
+3
💰 Gehalt: keine Angabe
🏢 Vor Ort
🗣️ EN

Stelle
DevSecOps Engineer
Raya
Cloudformation
+16
💰 Gehalt: keine Angabe
🌍 Remote
🗣️ EN

Stelle
Director of Production Engineering
Legion
CloudformationDatadog
+9
💰 Gehalt: keine Angabe
🌍 Remote
Anywhere in the World
🗣️ EN

Stelle
System Engineer - Network Systems
Cloudflare
Clickhouse
+3
💰 Gehalt: keine Angabe
🏢 Vor Ort
🗣️ EN

Stelle
Systems Engineer - Database Platform
Cloudflare
Cpp
+10
💰 Gehalt: keine Angabe
🏢 Vor Ort
🗣️ EN
