← Zurück zur Liste
Stelle · Mid-level

AI Infrastructure & Platform Operations Engineer (remote in the EU)

Sonstige • Mid-level • Remote • Vollzeit • Polen Polen

Mirantis is building a European AI Infrastructure & Platform Operations team to run AI environments powered by NVIDIA GPUs, high-performance networking and Kubernetes across multiple datacenters. The role works in a shift-based operational environment.

Responsibilities

  • ▹Monitor, operate and support production AI infrastructure platforms
  • ▹Investigate and resolve infrastructure, networking, hardware and platform incidents
  • ▹Support NVIDIA GPU infrastructure and associated platform services
  • ▹Monitor and troubleshoot Kubernetes-based environments
  • ▹Investigate performance, availability and reliability issues
  • ▹Collaborate with engineering teams, hardware vendors, datacenter personnel and service delivery teams
  • ▹Participate in incident response, root cause analysis and operational improvement
  • ▹Improve monitoring, observability, automation and operational processes
  • ▹Maintain operational documentation, runbooks and knowledge articles

Requirements

  • ▹3+ years in infrastructure, platform, network, cloud or datacenter operations, SRE or related technical roles
  • ▹Strong Linux administration and troubleshooting skills
  • ▹Good understanding of networking and experience diagnosing infrastructure issues
  • ▹Working knowledge of Kubernetes in production
  • ▹Experience supporting production infrastructure and services
  • ▹Strong analytical and problem-solving skills
  • ▹Experience within structured operational and incident management processes
  • ▹Excellent communication and collaboration skills
  • ▹Ability to work in a shift-based operational environment

Nice to have

  • ▹NVIDIA GPU infrastructure and accelerated computing platforms
  • ▹InfiniBand networking and NVIDIA UFM
  • ▹Kubernetes platform operations
  • ▹AI infrastructure or HPC environments
  • ▹Site Reliability Engineering (SRE) or platform engineering
  • ▹Observability platforms such as Grafana, Prometheus, ELK or OpenTelemetry
  • ▹Infrastructure automation and Infrastructure-as-Code practices
  • ▹Large-scale distributed systems and production platforms

Soft skills

Analytical thinkingProblem solvingCommunicationCollaboration

What we offer

  • ▹Work with some of the most advanced AI infrastructure environments in production
  • ▹Exposure to NVIDIA GPU technologies, Kubernetes platforms and high-performance networking
  • ▹Help define how next-generation AI infrastructure is operated
  • ▹Growing organisation investing heavily in AI infrastructure

About the company

Mirantis is a Kubernetes-native AI infrastructure company providing scalable, secure and sovereign infrastructure for AI and machine learning workloads. Its customers include enterprises such as Adobe, DocuSign, PayPal and Volkswagen.

Ähnliche Stellen