← Zurück zur Liste

Stelle
HPC Infrastructure Site Reliability Engineer
DevOps / SRE
• Remote
• Vollzeit
•
Gloucestershire, EU/EMEA
A fast-growing GPU-as-a-Service provider is hiring a senior Infrastructure SRE with experience operating large-scale distributed systems and hands-on HPC/AI infrastructure. The role ensures reliability and performance of mission-critical infrastructure in a 24/7 on-call environment.
Responsibilities
- ▹Operate mission-critical HPC/AI infrastructure in 24/7 on-call rotation
- ▹Troubleshoot complex, cross-layer issues in GPU-based HPC systems
- ▹Perform performance evaluation and acceptance testing of new HPC environments
- ▹Improve observability and automation across the platform
- ▹Collaborate with network, platform SRE and data centre teams
- ▹Feed operational insight back into infrastructure design decisions
Requirements
- ▹Experience operating large-scale, distributed infrastructure
- ▹Recent hands-on HPC and AI infrastructure experience
- ▹Strong Linux and distributed systems expertise
- ▹Knowledge of NVIDIA GPU ecosystems and RDMA networking (RoCE, InfiniBand)
- ▹Experience across bare metal, networking, storage, virtualisation, orchestration
Nice to have
- ▹Experience with latest high-density AI compute platforms
- ▹Performance validation and benchmarking experience
- ▹Familiarity with Ansible, Prometheus, Grafana
Soft skills
Cross-functional collaboration across multiple teamsDeep technical, analytical problem-solving mindsetIndependent ownership in a fast-moving, critical environment
What we offer
- ▹Hands-on access to cutting-edge GPU/CPU platforms
- ▹Work on some of the world's most advanced HPC infrastructure
About the company
A fast-growing GPU-as-a-Service provider delivering scalable, high-performance compute infrastructure for AI and HPC workloads across global data centres.
Ähnliche Stellen
Stelle
Site Reliability Engineer
Incident IQ
Datadog
+10
💰 Gehalt: keine Angabe
🌍 Remote
🗣️ EN

Stelle
DevOps Engineer mit Schwerpunkt Kubernetes (all genders)
dymatrix
+6
💰 Gehalt: keine Angabe
🏢 Vor Ort
Indien
🗣️ EN

Stelle
DevOps Engineer - Contractor
Coin Market Cap
Argocd
+7
💰 Gehalt: keine Angabe
🌍 Remote
Anywhere in the World
🗣️ EN

Stelle
Site Reliability Engineer
Astera
+3
89 130–267 391 €/Jahr
brutto
🌍 Remote
Emeryville HQ
🗣️ EN

Stelle
Platform Site Reliability Engineer
Radiant
+2
💰 Gehalt: keine Angabe
🌍 Remote
Gloucestershire
🗣️ EN

Stelle
Infrastructure Site Reliability Engineer
Radiant
+2
💰 Gehalt: keine Angabe
🌍 Remote
Gloucestershire
🗣️ EN