← Back to list
Job

Infrastructure Site Reliability Engineer

DevOps / SRE • Remote • Full-time • European Union Gloucestershire, EU/EMEA

Radiant is looking for an Infrastructure SRE to run and evolve its infrastructure stack across bare-metal, virtualization and orchestration layers. The role covers 24/7 stability, security, automation and mentorship for AI/HPC workloads.

Responsibilities

  • ▹Deploy resilient, scalable infrastructure for AI/HPC workloads
  • ▹Optimize Linux system configuration, BIOS/firmware and kernel
  • ▹Configure and manage bare-metal infrastructure (IPMI, Redfish)
  • ▹Build and maintain automation scripts and Infrastructure as Code
  • ▹Apply ITSM frameworks
  • ▹Maintain the observability stack (Prometheus, Grafana)
  • ▹Support 24/7 production environments, on-call rotation
  • ▹Mentor junior engineers

Requirements

  • ▹5+ years in globally scaled environments
  • ▹Expert-level Linux administration (Ubuntu)
  • ▹System tuning and disk I/O optimization experience
  • ▹Familiarity with out-of-band management tools (IPMI, Redfish, PXE)
  • ▹Strong networking fundamentals (TCP/IP, DNS, DHCP, VLANs)
  • ▹Infrastructure scripting (Bash, Python, Ansible)
  • ▹Hands-on experience with orchestration platforms (Kubernetes, MAAS, Tinkerbell)

Nice to have

  • ▹Knowledge of HPC workloads and GPU-based infrastructure
  • ▹Experience with InfiniBand networks and HPC performance tuning
  • ▹Bachelor's or Master's degree in CS/Engineering

Soft skills

Excellent communication and mentorship skillsCustomer-focused attitude

About the company

Radiant designs and operates AI-native cloud platforms engineered for sovereignty, performance and scale, powering GPU-native workloads.

Similar jobs