← Back to list
Job · Senior

Senior Site Reliability Engineer (SRE)

DevOps / SRE • Senior • Remote • Full-time • Bulgaria Bulgaria

Senior SRE role at Mirantis, contributing to the design, development and operation of cloud-based AI solutions on the CNCF/Kubernetes ecosystem, focused on deploying AI infrastructure on NVIDIA-certified hardware.

Responsibilities

  • ▹Design, develop and operate cloud AI solutions on Kubernetes
  • ▹Deploy AI infrastructure on NVIDIA-certified hardware
  • ▹Optimize system performance, reliability and scalability
  • ▹Troubleshoot and debug complex technical issues
  • ▹Design and implement AI-driven DevOps automation
  • ▹Facilitate customer knowledge transfer during delivery

Requirements

  • ▹5+ years DevOps experience with cloud/infrastructure incl. Kubernetes and/or OpenStack
  • ▹Experience with high-performance data center processing, networking, storage
  • ▹Golang exposure plus working knowledge of Python/JavaScript
  • ▹Strong distributed systems, microservices and CI/CD knowledge
  • ▹Strong Linux and Kubernetes troubleshooting skills
  • ▹Excellent written and spoken English, customer-facing skills
  • ▹Willingness to travel up to 25%

Nice to have

  • ▹Network and/or storage architecture experience
  • ▹HPC/GPU infrastructure experience (GPU scheduling, MIG/vGPU, RDMA/RoCE, InfiniBand, NVLink, DCGM, NVIDIA AI Enterprise)
  • ▹Open source community presence, conference talks
  • ▹Experience with Rancher, OpenShift, VMware

Soft skills

Leads technical tasks, collaborates across diverse teamsIndependent judgment with customersCommitment to continuous learningStrong customer-facing communication

About the company

Mirantis builds cloud and AI infrastructure solutions based on open source software, including Kubernetes, OpenStack and the CNCF ecosystem.

Similar jobs