← Zurück zur Liste
Stelle

Platform Site Reliability Engineer

DevOps / SRE • Remote • Vollzeit • Europäische Union Gloucestershire, EU/EMEA

Radiant builds AI-native cloud platforms for GPU-native workloads and is looking for a Platform SRE to deploy and evolve Kubernetes clusters across bare-metal and partner infrastructure. The role includes automation, observability and 24/7 operational support.

Responsibilities

  • ▹Deploy and manage Kubernetes clusters
  • ▹Develop Kubernetes manifests and operators
  • ▹Optimize Linux system configuration
  • ▹Build and maintain automation scripts and Infrastructure as Code
  • ▹Apply ITSM frameworks (incident, change management)
  • ▹Maintain the observability stack (Prometheus, Grafana)
  • ▹Support 24/7 production environments, on-call rotation
  • ▹Mentor junior engineers

Requirements

  • ▹5+ years in globally scaled, performance-intensive environments
  • ▹3+ years operating Kubernetes
  • ▹Expert-level Linux administration (Ubuntu)
  • ▹System tuning and disk I/O optimization knowledge
  • ▹Strong networking fundamentals (TCP/IP, DNS, DHCP, VLANs)
  • ▹Infrastructure scripting (Bash, Python, Ansible)
  • ▹Knowledge of observability tools (Prometheus, Grafana)

Nice to have

  • ▹Knowledge of running AI workloads on orchestration platforms
  • ▹Bachelor's or Master's degree in CS/Engineering
  • ▹LPIC certification
  • ▹ITIL Foundation qualification
  • ▹Certified Kubernetes Administrator (CKA)

Soft skills

Excellent communication and mentorship skillsCustomer-focused attitudeCommitment to documentation and automation

About the company

Radiant designs and operates AI-native cloud platforms engineered for sovereignty, performance and scale, powering GPU-native workloads.

Ähnliche Stellen