← Back to list

Job
· Senior
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
AI / ML Engineer
• Senior
• Remote
• Full-time
•
Poland
NVIDIA's Deep Learning Frameworks infrastructure team is looking for a senior HPC cluster administrator to lead the design, deployment and reliability of large-scale GPU clusters, from DGX/HGX platforms to Grace Blackwell systems.
Responsibilities
- ▹Own the full lifecycle of GPU compute clusters (procurement, provisioning, configuration management, monitoring, deprecation) across heterogeneous Linux environments (DGX, HGX, embedded systems)
- ▹Design and scale storage solutions (NFS, Lustre, WekaFS or equivalent) with a clear roadmap for capacity and performance growth
- ▹Lead infrastructure automation using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab)
- ▹Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management and MIG/GPU partitioning strategies
- ▹Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents
- ▹Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads
- ▹Evaluate and introduce new technologies (networking fabrics such as InfiniBand, NVLink, EFA/RDMA; storage tiers; container runtimes) to improve performance and reliability
- ▹Mentor junior engineers and contribute to team-wide engineering standards
Requirements
- ▹BS/MS in CS, EE, CE, or equivalent hands-on experience
- ▹5+ years of experience deploying and administering large-scale HPC or ML training clusters
- ▹Deep expertise in Linux systems administration at scale
- ▹Strong scripting and automation skills in Python and/or bash
- ▹Hands-on experience with Slurm (scheduling, accounting, cgroup configuration)
- ▹Proficiency with configuration management and IaC (Ansible required; Terraform a plus)
- ▹Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes)
- ▹Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA)
- ▹Experience with distributed/parallel filesystems and storage architecture
- ▹Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders
Nice to have
- ▹Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)
- ▹Familiarity with cluster management platforms (Colossus, Bright Cluster Manager, xCAT or similar)
- ▹Experience supporting large-scale distributed deep learning workloads (PyTorch, JAX, Megatron)
- ▹Knowledge of BMC/IPMI/Redfish for out-of-band management and hardware lifecycle
- ▹Background in MLOps tooling or ML platform engineering
Soft skills
Owning problems end-to-endClear communication with engineering and management stakeholdersMentoring
What we offer
- ▹Base salary for Poland: 221,250-383,500 PLN for Level 3, 292,500-507,000 PLN for Level 4
- ▹Collaborative and inclusive environment
About the company
NVIDIA's Deep Learning Frameworks (DLFW) Infrastructure team runs the large-scale GPU compute clusters that handle the industry's most demanding deep learning training, inference and HPC workloads.
Education: BS/MS informatikából, villamosmérnöki vagy számítógép-mérnöki területen, vagy egyenértékű gyakorlati tapasztalat
Similar jobs

Job
· Senior
Senior Machine Learning Engineer (AdTech)
Sigma Software
AI/MLBigqueryData Science
+8
💰 Salary: not specified
🌍 Remote
🗣️ EN
Himalayas

Job
· Senior
AI/ML Engineer
Pythian
AI/ML
+11
💰 Salary: not specified
🌍 Remote
🗣️ EN
Himalayas

Job
· Senior
AI Engineer
Huzzle
AI/ML
+9
💰 Salary: not specified
🌍 Remote
🗣️ EN
Himalayas

Job
· Senior
Senior GenAI Engineer
EY
AI/MLAzure DevopsData Science
+11
💰 Salary: not specified
🏢 On-site
Katowice
🗣️ EN

Job
· Senior
Senior AI Engineer
Procter & Gamble
AI/MLDatabricks
+9
💰 Salary: not specified
🏢 On-site
WARSAW
🗣️ EN

Job
· Senior
Senior AI Engineer
Bayer
AI/MLCosmosdbDatabricks
+16
💰 Salary: not specified
🏢 On-site
Warszawa
🗣️ EN