← Back to list
Job · Senior

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

AI / ML Engineer • Senior • Remote • Full-time • Poland Poland

NVIDIA's Deep Learning Frameworks infrastructure team is looking for a senior HPC cluster administrator to lead the design, deployment and reliability of large-scale GPU clusters, from DGX/HGX platforms to Grace Blackwell systems.

Responsibilities

  • ▹Own the full lifecycle of GPU compute clusters (procurement, provisioning, configuration management, monitoring, deprecation) across heterogeneous Linux environments (DGX, HGX, embedded systems)
  • ▹Design and scale storage solutions (NFS, Lustre, WekaFS or equivalent) with a clear roadmap for capacity and performance growth
  • ▹Lead infrastructure automation using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab)
  • ▹Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management and MIG/GPU partitioning strategies
  • ▹Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents
  • ▹Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads
  • ▹Evaluate and introduce new technologies (networking fabrics such as InfiniBand, NVLink, EFA/RDMA; storage tiers; container runtimes) to improve performance and reliability
  • ▹Mentor junior engineers and contribute to team-wide engineering standards

Requirements

  • ▹BS/MS in CS, EE, CE, or equivalent hands-on experience
  • ▹5+ years of experience deploying and administering large-scale HPC or ML training clusters
  • ▹Deep expertise in Linux systems administration at scale
  • ▹Strong scripting and automation skills in Python and/or bash
  • ▹Hands-on experience with Slurm (scheduling, accounting, cgroup configuration)
  • ▹Proficiency with configuration management and IaC (Ansible required; Terraform a plus)
  • ▹Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes)
  • ▹Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA)
  • ▹Experience with distributed/parallel filesystems and storage architecture
  • ▹Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders

Nice to have

  • ▹Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)
  • ▹Familiarity with cluster management platforms (Colossus, Bright Cluster Manager, xCAT or similar)
  • ▹Experience supporting large-scale distributed deep learning workloads (PyTorch, JAX, Megatron)
  • ▹Knowledge of BMC/IPMI/Redfish for out-of-band management and hardware lifecycle
  • ▹Background in MLOps tooling or ML platform engineering

Soft skills

Owning problems end-to-endClear communication with engineering and management stakeholdersMentoring

What we offer

  • ▹Base salary for Poland: 221,250-383,500 PLN for Level 3, 292,500-507,000 PLN for Level 4
  • ▹Collaborative and inclusive environment

About the company

NVIDIA's Deep Learning Frameworks (DLFW) Infrastructure team runs the large-scale GPU compute clusters that handle the industry's most demanding deep learning training, inference and HPC workloads.

Education: BS/MS informatikából, villamosmérnöki vagy számítógép-mérnöki területen, vagy egyenértékű gyakorlati tapasztalat

Similar jobs