← Back to list
Job · Senior

Senior Machine Learning Engineer, LLM Inference Optimization

AI / ML Engineer • Senior • Remote • Full-time • European Union EU/EMEA

On Nebius's Applied AI team you will own model and endpoint optimization from model artifacts through production, improving latency, throughput and cost per token.

Responsibilities

  • ▹Own optimization work for specific model families, customer endpoints or serving backends
  • ▹Run engine comparisons and recommend practical serving configurations
  • ▹Debug model quality or performance regressions during production rollouts
  • ▹Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality and cost per token
  • ▹Deploy, configure, benchmark and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo
  • ▹Build and productionize model-compression workflows: quantization, quantization-aware training, distillation, low-bit serving
  • ▹Implement speculative decoding, KV-cache optimization, prefix caching, chunked prefill, continuous batching and disaggregated prefill/decode serving
  • ▹Build reproducible benchmark harnesses (TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, cost per token)
  • ▹Partner with GPU kernel and platform engineers to diagnose bottlenecks
  • ▹Write design docs, performance reports, rollout plans and customer-facing technical explanations

Requirements

  • ▹Strong Python and PyTorch engineering skills
  • ▹Hands-on experience deploying or optimizing LLM, VLM or high-throughput transformer inference systems
  • ▹Practical knowledge of at least one modern inference stack (vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe or equivalent)
  • ▹Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching and parallelism

About the company

Nebius builds a full-stack AI cloud platform, from data and model training to production deployment. Listed on Nasdaq and headquartered in Amsterdam, with a team of 1,500+.

Similar jobs