← Back to list
Job

HPC Infrastructure Site Reliability Engineer

DevOps / SRE • Remote • Full-time • European Union Gloucestershire, EU/EMEA

A fast-growing GPU-as-a-Service provider is hiring a senior Infrastructure SRE with experience operating large-scale distributed systems and hands-on HPC/AI infrastructure. The role ensures reliability and performance of mission-critical infrastructure in a 24/7 on-call environment.

Responsibilities

  • ▹Operate mission-critical HPC/AI infrastructure in 24/7 on-call rotation
  • ▹Troubleshoot complex, cross-layer issues in GPU-based HPC systems
  • ▹Perform performance evaluation and acceptance testing of new HPC environments
  • ▹Improve observability and automation across the platform
  • ▹Collaborate with network, platform SRE and data centre teams
  • ▹Feed operational insight back into infrastructure design decisions

Requirements

  • ▹Experience operating large-scale, distributed infrastructure
  • ▹Recent hands-on HPC and AI infrastructure experience
  • ▹Strong Linux and distributed systems expertise
  • ▹Knowledge of NVIDIA GPU ecosystems and RDMA networking (RoCE, InfiniBand)
  • ▹Experience across bare metal, networking, storage, virtualisation, orchestration

Nice to have

  • ▹Experience with latest high-density AI compute platforms
  • ▹Performance validation and benchmarking experience
  • ▹Familiarity with Ansible, Prometheus, Grafana

Soft skills

Cross-functional collaboration across multiple teamsDeep technical, analytical problem-solving mindsetIndependent ownership in a fast-moving, critical environment

What we offer

  • ▹Hands-on access to cutting-edge GPU/CPU platforms
  • ▹Work on some of the world's most advanced HPC infrastructure

About the company

A fast-growing GPU-as-a-Service provider delivering scalable, high-performance compute infrastructure for AI and HPC workloads across global data centres.

Similar jobs