← Zurück zur Liste
Stelle · Senior

Senior HPC DevOps Engineer

DevOps / SRE • Senior • Remote • Vollzeit • Deutschland Deutschland

NVIDIA is looking for an experienced HPC DevOps and Network Engineer to help build the supercomputers and HPC clusters of the future, working in AI and GPU computing.

Responsibilities

  • ▹Design, implement and maintain large-scale HPC/AI clusters with state-of-the-art monitoring, logging and alerting
  • ▹Use and develop tools to manage infrastructure as code for scalable, repeatable deployments
  • ▹Develop and maintain CI/CD pipelines to automate deployment
  • ▹Develop automation scripts and tools for deployment, configuration management and operational monitoring
  • ▹Develop complex networking automations
  • ▹Perform comprehensive troubleshooting from bare metal to application level
  • ▹Serve as a technical resource and share best practices with internal teams
  • ▹Support R&D activities and proofs of concept and proofs of value

Requirements

  • ▹B.Sc. in Computer Science, Engineering or a related field with 5+ years of experience
  • ▹Deep knowledge of HPC and AI solution technologies, including CPUs, GPUs, high-speed interconnects and supporting software
  • ▹Advanced proficiency in programming and scripting languages and a solid understanding of object-oriented programming principles
  • ▹Familiarity with Jenkins, Ansible, Puppet/Chef
  • ▹Excellent knowledge of Windows and Linux (Red Hat/CentOS and Ubuntu), networking and OS-level security
  • ▹Deep understanding of networking protocols such as InfiniBand and Ethernet
  • ▹Experience with job scheduling and orchestration tools such as Slurm and Kubernetes
  • ▹Background with storage solutions such as Lustre, GPFS, ZFS and XFS
  • ▹Expertise with virtual systems (VMware, Hyper-V, KVM, Citrix)
  • ▹Familiarity with cloud platforms (AWS, Azure, Google Cloud)

Nice to have

  • ▹Proven networking experience or professional networking training
  • ▹Knowledge of CPU and/or GPU architecture
  • ▹Understanding of Kubernetes and container-related microservice technologies
  • ▹Experience with GPU-focused hardware/software (DGX, CUDA)
  • ▹Background with RDMA (InfiniBand or RoCE) fabrics

Soft skills

TeamworkKnowledge sharing

About the company

NVIDIA is the world leader in accelerated computing and artificial intelligence.

Education: Informatikai vagy mérnöki alapdiploma (B.Sc.)

Ähnliche Stellen