← Back to list

Job
· Principal
Principal ML Ops Engineer (EMEA Remote)
Other
• Principal
• Remote
• Full-time
•
Ukraine, Ukraine
Pragmatike is hiring on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup, seeking a Principal ML Ops Engineer to build and operate production-grade model serving infrastructure for real-time AI applications.
Responsibilities
- ▹Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent
- ▹Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models
- ▹Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers
- ▹Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance
- ▹Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health
- ▹Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments
- ▹Own the full lifecycle of ML systems from development through production, including on-call responsibilities
- ▹Define engineering best practices and contribute to platform scalability
Requirements
- ▹4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems
- ▹Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent
- ▹Strong background in container orchestration and operating GPU-based workloads in production
- ▹Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines
- ▹Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)
- ▹Strong understanding of distributed systems, performance tuning, and production reliability engineering
- ▹Ability to effectively use AI coding assistants
- ▹Ownership mindset with the ability to operate independently in a remote-first environment
Nice to have
- ▹Experience with ML platforms such as Kubeflow, MLflow, or KubeAI
- ▹Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems
- ▹Experience with cost optimization across different GPU types and inference workloads
- ▹Background in early-stage startups or greenfield infrastructure projects
- ▹Proven experience building production systems from scratch rather than maintaining legacy platforms
Soft skills
Ownership mindset with the ability to operate independently in a remote-first environment
About the company
Pragmatike is recruiting on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup building next-generation AI-native cloud services. The company provides GPU-powered infrastructure for AI/ML workloads, secure storage, and high-speed data transfer through a decentralized architecture that significantly reduces environmental impact compared to traditional cloud providers.
Languages: Angol: Felsőfok
Similar jobs

Job
· Principal
Sr. Staff Software Development Engineer - AI Platform
Zscaler
AI/MLArgocd
+10
💰 Salary: not specified
🏢 On-site
Bangalore
🗣️ EN

Job
· Principal
Member of Technical Staff, Cloud Infrastructure
Fireworks AI
AI/MLArgocdCpp
+8
💰 Salary: not specified
🏢 On-site
San Mateo
🗣️ EN

Job
· Principal
Member of Technical Staff, Cloud Infrastructure
Fireworks AI
AI/MLArgocd
+10
$175,000–$220,000/yr
gross
🏢 On-site
New York
🗣️ EN
Job
· Principal
Member of Technical Staff, Machine Learning - NomadicML
Pear Vc
AI/MLHuggingfaceKubeflowMlflow
+2
💰 Salary: not specified
🏢 On-site
San Francisco
🗣️ EN

Job
· Principal
Senior Member Technical Staff (MTS 3) - Machine learning
The Nielsen Company
AI/MLCosmosdb
+18
💰 Salary: not specified
🏢 On-site
Bengaluru
🗣️ EN

Job
· Principal
(US) Principal ML System Engineer
PointClickCare
AI/MLDatabricks
+7
$195,000–$217,000/yr
gross
🌍 Remote
🗣️ EN