← Back to list

Job
HPC Infrastructure Site Reliability Engineer
DevOps / SRE
• Remote
• Full-time
•
Gloucestershire, EU/EMEA
A fast-growing GPU-as-a-Service provider is hiring a senior Infrastructure SRE with experience operating large-scale distributed systems and hands-on HPC/AI infrastructure. The role ensures reliability and performance of mission-critical infrastructure in a 24/7 on-call environment.
Responsibilities
- ▹Operate mission-critical HPC/AI infrastructure in 24/7 on-call rotation
- ▹Troubleshoot complex, cross-layer issues in GPU-based HPC systems
- ▹Perform performance evaluation and acceptance testing of new HPC environments
- ▹Improve observability and automation across the platform
- ▹Collaborate with network, platform SRE and data centre teams
- ▹Feed operational insight back into infrastructure design decisions
Requirements
- ▹Experience operating large-scale, distributed infrastructure
- ▹Recent hands-on HPC and AI infrastructure experience
- ▹Strong Linux and distributed systems expertise
- ▹Knowledge of NVIDIA GPU ecosystems and RDMA networking (RoCE, InfiniBand)
- ▹Experience across bare metal, networking, storage, virtualisation, orchestration
Nice to have
- ▹Experience with latest high-density AI compute platforms
- ▹Performance validation and benchmarking experience
- ▹Familiarity with Ansible, Prometheus, Grafana
Soft skills
Cross-functional collaboration across multiple teamsDeep technical, analytical problem-solving mindsetIndependent ownership in a fast-moving, critical environment
What we offer
- ▹Hands-on access to cutting-edge GPU/CPU platforms
- ▹Work on some of the world's most advanced HPC infrastructure
About the company
A fast-growing GPU-as-a-Service provider delivering scalable, high-performance compute infrastructure for AI and HPC workloads across global data centres.
Similar jobs
Job
Site Reliability Engineer
Incident IQ
Datadog
+10
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
DevOps Engineer mit Schwerpunkt Kubernetes (all genders)
dymatrix
+6
💰 Salary: not specified
🏢 On-site
Indien
🗣️ EN

Job
DevOps Engineer - Contractor
Coin Market Cap
Argocd
+7
💰 Salary: not specified
🌍 Remote
Anywhere in the World
🗣️ EN

Job
Site Reliability Engineer
Astera
+3
$100,000–$300,000/yr
gross
🌍 Remote
Emeryville HQ
🗣️ EN

Job
Platform Site Reliability Engineer
Radiant
+2
💰 Salary: not specified
🌍 Remote
Gloucestershire
🗣️ EN

Job
Infrastructure Site Reliability Engineer
Radiant
+2
💰 Salary: not specified
🌍 Remote
Gloucestershire
🗣️ EN