← Zurück zur Liste
Stelle

Datacenter Infrastructure Specialist

Sonstige • Remote • Vollzeit • Europäische Union EU/EMEA

Runpod is looking for a Datacenter Infrastructure Specialist to own the operational health and lifecycle of its global high-density GPU fleet, acting as the technical link between hardware partners and internal engineering. The role blends HPC systems engineering, network troubleshooting and process automation.

Responsibilities

  • ▹Validate new hardware and ensure partner deployments meet Runpod's specifications for distributed AI/ML workloads
  • ▹Monitor fleet health to identify performance degradation, audit downtime and provide the technical data needed to protect customer SLAs
  • ▹Work with LLMs and AI agents to automate network triage and generate dynamic runbooks for the fleet
  • ▹Coordinate technical incident communications and translate outages into actionable resolutions
  • ▹Support the growth of infrastructure partners

Requirements

  • ▹3-5 years of experience in infrastructure operations, systems reliability or datacenter engineering
  • ▹Strong proficiency in standard datacenter networking and performance troubleshooting
  • ▹Hands-on experience with the NVIDIA software stack (driver installation, performance utilities) and an understanding of multi-node performance tuning
  • ▹Solid Linux system administration skills and experience with containerization (Docker)

Nice to have

  • ▹Exposure to RDMA, InfiniBand or RoCE (highly preferred)

Soft skills

CommunicationOwnership

What we offer

  • ▹Remote work (remote-first team)

About the company

Runpod is the AI Developer Cloud, used by more than one million developers to experiment, train, fine-tune, deploy and scale AI. The platform has processed more than 20 billion inference requests, and the company closed a $100M Series A in June 2026. A small, remote-first team.

Ähnliche Stellen