← Back to list
Job · Principal

Staff Machine Learning Engineer - Ops

AI / ML Engineer • Principal • On-site • Full-time United Kingdom London, United Kingdom
The role Today at Wayve, our model development cycle is composed of multiple complex training phases, each building on the last. As a Staff Machine Learning Engineer (Ops/Release), you are expected to have a deep understanding of each of the training phases, understanding how our models are created from start to finish. You will ultimately be responsible for setting and enforcing the standard of each of the release gates along this journey, ensuring that each training phase is sufficiently validated before the next phase begins. You'll drive technical excellence across our ML delivery pipelines. You'll review release content to ensure it meets our standards, identify bottlenecks in the process, and partner with platform teams to make sure tooling meets our delivery needs. You'll work with CI/CD teams to adapt workflows and streamline model delivery, and with evaluation teams to keep our methods reliable — spotting gaps and driving new methodology for evaluating our models. This is a high-trust, high-visibility role: our release process directly protects our model baseline, and a mistake here has real consequences for how the product performs on-road and how it's perceived externally. Key responsibilities: Collaborate with ML engineers, data engineers and product teams to deliver features end to end. Review release content — model and metric changes, evaluation results — to confirm everything meets Wayve's quality and safety standards before it ships. Identify bottlenecks in the ML delivery pipeline and drive fixes that improve speed without compromising quality. Collaborate with AI Platform teams to ensure tooling meets our delivery needs, defining and building the checks and automation that catch issues earlier. Collaborate with CI/CD teams to adapt workflows and streamline model delivery. Collaborate with evaluation teams to ensure evaluation methods are reliable, identify gaps, and drive new methodology for evaluating our models. Stay up to date with the latest in MLOps practices and tools and bring improvements into the workflow. About you In order to set you up for success as a Staff Machine Learning - Ops at Wayve, we’re looking for the following skills and experience. Essential Full system thinker with experience of introducing operational processes to build engineering excellence. Strong ML Ops, model registry and ML lifecycle experience A deep technical depth in ML training A strong understanding of ML code infrastructure and best practices – experience with pytorch, tensor RT, quantisation and model deployment Strong CI/CD and Github Actions experience Strong communications skills with a collaborative mindset Desirable Experience with Pytorch, TensorRT, quantisation and model deployment Experience with Grafana monitoring and production observability This is a full-time role based in our office in London. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home.

Similar jobs