← Back to list

Job
· Principal
Member of Technical Staff (AI Infrastructure Engineer)
Platform Engineer
• Principal
• Hybrid
• Full-time
•
London, United Kingdom
We are looking for an AI Infrastructure Engineer to join our growing team, building, deploying, and optimizing large-scale AI training and inference clusters in close partnership with the Inference and Research teams.
Responsibilities
- ▹Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads
- ▹Manage and optimize Slurm-based HPC environments for distributed training of large language models
- ▹Develop robust APIs and orchestration systems for training pipelines and inference services
- ▹Implement resource scheduling and job management across heterogeneous compute environments
- ▹Benchmark system performance, diagnose bottlenecks, and improve both training and inference infrastructure
- ▹Build monitoring, alerting, and observability solutions for ML workloads on Kubernetes and Slurm
- ▹Respond quickly to outages and collaborate across teams to maintain high uptime for training and inference
- ▹Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands
Requirements
- ▹Strong Kubernetes administration expertise (CRDs, operators, cluster management)
- ▹Hands-on Slurm workload management (job scheduling, resource allocation, cluster optimization)
- ▹Experience deploying and managing distributed training systems at scale
- ▹Deep understanding of container orchestration and distributed systems architecture
- ▹High-level familiarity with LLM architecture and training (Multi-Head/Multi-Query Attention, distributed training strategies)
- ▹Experience managing GPU clusters and optimizing compute utilization
- ▹Expert-level Kubernetes administration and YAML configuration management
- ▹Python and C++ programming with a systems/infrastructure automation focus
- ▹Hands-on experience with ML frameworks (PyTorch) in distributed training contexts
- ▹Strong understanding of networking, storage, and compute resource management for ML workloads
- ▹Experience developing APIs and managing distributed systems for batch and real-time workloads
- ▹Solid debugging and monitoring skills for containerized environments
Nice to have
- ▹Kubernetes operators and custom controllers for ML workloads
- ▹Advanced Slurm administration including multi-cluster federation and scheduling policies
- ▹GPU cluster management and CUDA optimization
- ▹Familiarity with other ML frameworks (TensorFlow) or distributed training libraries
- ▹HPC, parallel computing, and high-performance networking background
- ▹Infrastructure as code (Terraform, Ansible) and GitOps practices
- ▹Container registries, image optimization, and multi-stage builds for ML workloads
Soft skills
Ability to respond quickly and stay composed under pressure during outagesStrong cross-team collaboration
Similar jobs

Job
· Principal
Principal Platform Infrastructure Engineer (SRE Enablement)
Menlosecurity
+6
💰 Salary: not specified
🌍 Remote
EMEA - Distributed (UK)
🗣️ EN

Job
· Principal
Principal Platform Engineer
Clarity Innovations
+6
$113,000–$300,000/yr
gross
🏢 On-site
Herndon
🗣️ EN

Job
· Principal
Senior Principal Platform Engineer
Clarity Innovations
AI/MLArgocd
+16
$123,000–$322,000/yr
gross
🏢 On-site
Jessup
🗣️ EN

Job
· Principal
Principal AI Platform Engineer
Arm
Llm
💰 Salary: not specified
🏢 On-site
Cambridge
🗣️ EN

Job
· Principal
Senior Staff Machine Learning Platform Engineer
Faire
AI/MLDatabricksDatadog
+16
$295,000–$405,500/yr
gross
🔀 Hybrid
San Francisco
🗣️ EN

Job
· Principal
Staff AI Infrastructure Engineer
Anduril
AI/MLCpp
+3
$220,000–$292,000/yr
gross
🏢 On-site
Costa Mesa
🗣️ EN