← Zurück zur Liste
Stelle · Senior

Senior HPC AI Cluster Engineer

Sonstige • Senior • Remote • Vollzeit • Schweiz Schweiz

NVIDIA is looking for an engineer to design and operate large-scale HPC/AI clusters on its Networking Clusters Solutions Infrastructure team, working on supercomputer and GPU-based systems. The role spans the full lifecycle from design to monitoring and troubleshooting.

Responsibilities

  • ▹Design and monitor large-scale HPC/AI clusters
  • ▹Manage Linux job scheduling and orchestration
  • ▹Develop and maintain CI/CD pipelines
  • ▹Build automation tooling for infrastructure
  • ▹Deploy monitoring for servers, network, storage
  • ▹Troubleshoot from bare metal to application layer
  • ▹Document standard methodologies for internal teams

Requirements

  • ▹CS/engineering degree, 8+ years experience
  • ▹Knowledge of HPC and AI solution technologies (CPU/GPU)
  • ▹Experience with Slurm, Kubernetes orchestration
  • ▹Deep Linux (RedHat/CentOS, Ubuntu) and networking knowledge
  • ▹Storage solutions: Lustre, GPFS, Weka.io
  • ▹Python programming and bash scripting
  • ▹Automation tools: Jenkins, Ansible, Puppet/Chef
  • ▹InfiniBand, Ethernet networking protocols
  • ▹Virtualization: VMware, Hyper-V, KVM, Citrix
  • ▹Cloud platforms: AWS, Azure, GCP

Nice to have

  • ▹Knowledge of CPU/GPU architecture
  • ▹Kubernetes and microservice technologies
  • ▹GPU hardware/software experience (DGX, CUDA)
  • ▹RDMA (InfiniBand or RoCE) experience

Soft skills

Collaboration with researchers, developers, customersStrong documentation and knowledge-sharing skills

About the company

NVIDIA's Networking Clusters Solutions Infrastructure team builds supercomputers and AI clusters using groundbreaking technologies.

Ähnliche Stellen