← Zurück zur Liste
Stelle

Industrial AI Cloud - Infrastructure Engineer

Platform Engineer • Vor Ort • Vollzeit • Ungarn Budapest, Ungarn

You build, automate and operate the compute, network and storage environment of the joint NVIDIA and Deutsche Telekom industrial AI cloud: large-scale GPU clusters, bare-metal servers, IaC automation and monitoring following ITIL processes.

Responsibilities

  • ▹Coordinate operations with data center teams: hardware lifecycle activities (installs, GPU upgrades, storage expansion, firmware updates) and server/network interconnections and related documentation (NetBox)
  • ▹Provision and maintain bare-metal servers and GPU nodes (PXE boot, OS installs, firmware updates)
  • ▹Design and operate the NVIDIA AI related infrastructure stack
  • ▹Develop and maintain Ansible and Terraform playbooks to automate provisioning, configuration and deployments
  • ▹Maintain Debian-based environments, apply patches and manage firmware upgrades at scale
  • ▹Integrate and maintain Keycloak, Entra ID / CAIMAN and AD for user authentication and authorization
  • ▹Support and operate distributed AI workloads on bare-metal hosts and in Kubernetes environments
  • ▹Operate Prometheus and Grafana stacks for proactive infrastructure monitoring and alerting
  • ▹Manage high-performance storage environments (WEKA by Hitachi)
  • ▹Follow and improve incident, problem and change management workflows; document runbooks and standard operating procedures; adhere to ZERO Outage guidelines
  • ▹Provide project deliverables with a focus on the NVIDIA technology stack

Requirements

  • ▹Experience in hardware installation, maintenance and operations
  • ▹Advanced proficiency with Linux (Debian preferred) in production environments
  • ▹Hands-on experience with Infrastructure-as-Code (Ansible, Terraform); Redfish desirable
  • ▹Knowledge of NVIDIA GPU-accelerated server platforms
  • ▹Knowledge of the NVIDIA AI software stack related to GPU orchestration
  • ▹Knowledge of GPU-based cloud platform software stacks including dependencies on the underlying layers
  • ▹Solid understanding of networking fundamentals (IP, routing, VLANs, DNS, firewalls, L1, L2)
  • ▹Experience with identity and access management systems (Keycloak, Entra ID, LDAP)
  • ▹Familiarity with monitoring stacks (Prometheus, Grafana)
  • ▹Knowledge of high-performance storage systems (WEKA by Hitachi advantageous)
  • ▹Working knowledge of ITIL processes (incident, problem, change)
  • ▹Strong troubleshooting and operational support skills in a 24/7 mission-critical environment

Nice to have

  • ▹Experience with large GPU clusters, HPC or data center environments
  • ▹Knowledge of sovereign cloud and data security/compliance requirements
  • ▹Familiarity with Terraform and GitOps for infrastructure changes

Soft skills

Coordination between multiple teams

What we offer

  • ▹Work on Europe's first industrial AI cloud with cutting-edge technologies
  • ▹Direct collaboration with NVIDIA and Deutsche Telekom experts
  • ▹Hybrid working model, training opportunities and career progression
  • ▹Remote work is only possible within Hungary

About the company

Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group with more than 5300 employees and four sites in Hungary: Budapest, Debrecen, Pécs and Szeged. The industrial AI cloud, developed jointly with NVIDIA, is an AI factory in Germany that will host 10,000 GPUs.

Ähnliche Stellen