← Back to list
Job · Mid-level

MLOps Engineer

MLOps Engineer • Mid-level • Remote • Full-time European Union EU/EMEA

This role seeks an MLOps Engineer to connect the Data and ML world with production infrastructure and operations, ensuring models and AI agents run reliably and efficiently across AWS, GCP, or Azure.

Responsibilities

  • Design, implement, and operate scalable cloud infrastructure on AWS, GCP, or Azure for data and ML workloads
  • Automate infrastructure and deployments using Terraform, Ansible, Helm, and Kubernetes
  • Build and maintain CI/CD pipelines for continuous deployment of ML models, AI agents, and observability platforms
  • Manage containers and orchestration with Docker and Kubernetes (EKS)
  • Build and maintain reproducible training, validation, and inference pipelines (Airflow, MLflow, DVC, Spark)
  • Implement and optimize monitoring, logging, and observability solutions (Prometheus, Grafana, ELK/EFK, OpenTelemetry)
  • Ensure the security, availability, and resilience of cloud platforms
  • Integrate the full Data -> ML -> Deployment -> Monitoring cycle, working closely with Data, ML, DevOps, and Backend teams
  • Collaborate with ML teams to package and deploy models and AI agents to production
  • Participate in designing hybrid pipelines combining traditional ML, generative AI, and LLM-based agents (LangChain, LangGraph, CrewAI)

Requirements

  • Degree in Computer Engineering, Software, Telecommunications, or a related field
  • 3+ years of experience in automation, deployment, or cloud infrastructure management
  • Solid experience with infrastructure as code and automation (Terraform, Helm, Ansible)
  • Advanced knowledge of Linux, networking, and distributed systems
  • Hands-on experience with CI/CD pipelines (GitLab CI, Jenkins, ArgoCD, FluxCD)
  • Proficiency with Docker and Kubernetes
  • Experience with model and data management tools such as MLflow, DVC, or Vertex AI
  • Knowledge of monitoring, logging, and observability
  • Experience with SQL and NoSQL databases (PostgreSQL, TimescaleDB, MongoDB)
  • Ability to diagnose and resolve incidents in high-traffic production environments, optimizing cost and performance

Nice to have

  • Experience with messaging and streaming systems such as Kafka, Redpanda, or Benthos
  • Knowledge of serverless and event-driven architectures (AWS Lambda, SNS/SQS)
  • Experience with advanced observability and model performance metrics
  • Knowledge of Site Reliability Engineering (SRE) practices
  • Experience with modern MLOps platforms (Kubeflow, MLflow)
  • Familiarity with automated deployments and DevOps best practices
  • Knowledge of generative AI infrastructure (GPU, optimized containers, RAG serving)
  • Advanced technical English, written and spoken

Soft skills

Good communication and teamwork skills in multidisciplinary environments

Similar jobs