About the Role This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context layer that makes AI agents reliable, accurate, and secure for mission-critical business operations. You'll own the systems that keep those agents running fast and reliably in production — from design through deployment — working closely with ML and infrastructure teams to scale inference at increasing concurrency. What You'll Do Own inference and model-serving infrastructure end to end, from architecture design through production deployment. Build and scale systems that enable AI agents to run reliably and efficiently under high concurrency in production environments. Collaborate with ML and infrastructure teams to ensure seamless integration and drive performance optimization. Identify infrastructure bottlenecks and lead the engineering effort to resolve them. What We're Looking For 5+ years of experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments. Hands-on experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems. Demonstrated ability to optimize production ML systems for latency, throughput, and reliability at scale. Strong proficiency with containerization and orchestration technologies — Docker and Kubernetes — for deploying ML workloads. Experience building or maintaining distributed systems that handle concurrent requests and manage resource allocation under load. Solid command of monitoring, observability, and debugging tooling for production systems (e.g., Prometheus, Grafana, ELK, distributed tracing). Experience deploying and managing ML systems on cloud platforms such as AWS, GCP, or Azure. Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java. Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune) is a plus. Familiarity with real-time or low-latency inference systems, agentic AI pipelines, or enterprise data infrastructure is a plus. Location On-site in San Mateo, California, United States. Visa sponsorship is not available for this role.
Ähnliche Stellen

Stelle
CV/ML Platform Engineer
Allen Control Systems
AI/MLKubeflow
+3
💰 Gehalt: keine Angabe
🏢 Vor Ort
Austin
🗣️ EN

Stelle
CV/ML Platform Engineer
Allen Control Systems
AI/MLKubeflow
+3
💰 Gehalt: keine Angabe
🏢 Vor Ort
Austin
🗣️ EN

Stelle
Machine Learning Operations (MLOps) Engineer - VOIS
Vodafone
AI/ML
+16
💰 Gehalt: keine Angabe
🏢 Vor Ort
Pune
🗣️ EN

Stelle
MLOps Engineer (Machine Learning, MLFlow, Kubernetes, DVC)
Capgemini
AI/ML
+3
💰 Gehalt: keine Angabe
🏢 Vor Ort
Wrocław
🗣️ EN

Stelle
Plattformsingenjör (AI & MLOps)
Capgemini
AI/MLAzure DevopsBicep
+12
💰 Gehalt: keine Angabe
🏢 Vor Ort
Göteborg

Stelle
Cloud & MLOps Engineer - Financial Services
EY
AI/MLAzure DevopsBicep
+19
💰 Gehalt: keine Angabe
🏢 Vor Ort
Diegem
🗣️ EN
