← Zurück zur Liste
Stelle · Mid-level

ML Infrastructure Engineer

MLOps Engineer • Mid-level • Remote • Vollzeit • Europäische Union EU/EMEA

A seed-stage enterprise AI infrastructure company is hiring an ML Infrastructure Engineer to own and scale the inference and model-serving systems that keep AI agents running reliably and fast — a hands-on production engineering role, not research.

Responsibilities

  • ▹Own inference and model-serving infrastructure end to end — from initial design through production deployment and ongoing scaling
  • ▹Build and scale systems that enable AI agents to run reliably and efficiently under high and increasing concurrency
  • ▹Identify and resolve infrastructure bottlenecks in collaboration with ML and platform engineering teams
  • ▹Optimize systems for latency, throughput, and reliability across cloud-hosted production environments
  • ▹Drive observability, monitoring, and debugging practices across the production ML stack

Requirements

  • ▹5+ years of hands-on experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments
  • ▹Demonstrated experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems
  • ▹Proven ability to optimize production ML systems for latency, throughput, and reliability at scale
  • ▹Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads
  • ▹Background in distributed systems that handle high concurrency and dynamic resource allocation under load
  • ▹Proficiency with monitoring and observability tooling — e.g., Prometheus, Grafana, ELK stack, distributed tracing
  • ▹Experience deploying and managing ML systems on cloud platforms (AWS, GCP, or Azure)
  • ▹Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java

Nice to have

  • ▹Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune, or similar)
  • ▹Background in real-time inference or low-latency serving requirements
  • ▹Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines
  • ▹Experience with enterprise data infrastructure, data pipelines, or data integration platforms

What we offer

  • ▹Early-stage opportunity with significant ownership and impact
  • ▹Work on genuinely hard distributed systems problems in a production AI context
  • ▹Small, experienced team with deep ML and enterprise engineering backgrounds
  • ▹Well-funded at the seed stage with strong institutional backing

About the company

A seed-stage enterprise AI infrastructure company building the context layer that makes AI agents reliable, accurate, and secure for critical business operations — including highly regulated industries like insurance, banking, asset management, and healthcare.

Ähnliche Stellen