← Back to list
Job · Manager

Manager, ML Solutions Architecture - Token Factory

Solution / Enterprise Architect • Manager • Remote • Full-time • European Union EU/EMEA

You lead the EU regional Solutions Architect teams for Nebius Token Factory, a serverless platform for running and customizing open-source LLMs in production. Your scope is PoC delivery and post-sales technical support: the people, the standard and the operating cadence. You report to the Head of Solutions Architecture and can work remotely from any EU country.

Responsibilities

  • ▹Manage a team of 8 Solutions Architects, with continued growth planned: 1:1s, goal setting, performance reviews, promotion cases and individual growth plans
  • ▹Build an accurate picture of each SA's strengths, gaps and preferences, and allocate accounts and engagements against both expertise and interest
  • ▹Run the cadence that surfaces blockers early, take them to the development, product and business teams that can clear them, and stay on them until they do
  • ▹Hold people to outcomes: distinguish effort from delivered results, say so plainly when the two diverge and reflect it in ratings and compensation decisions
  • ▹Onboard new joiners through to their first independently delivered engagement
  • ▹Coach SAs into stronger engineers and stronger communicators, and make deliberate calls about who is ready for more scope
  • ▹Be accountable for your team's delivery outcomes: time from PoC kick-off to first optimized dedicated endpoint, success-criteria hit rate and the quality of the technical relationship after the customer goes to production
  • ▹Review technical work before it reaches the customer: benchmarking methodology, serving configurations, results, closure documents; catch the wrong conclusion drawn from a metrics artifact before a customer sees it
  • ▹Ensure staffing and escalation coverage across accounts and timezones, including post-sales request load that does not respect sprint boundaries
  • ▹Call infeasibility early and with evidence, rather than letting the team burn iterations against requirements the platform cannot meet today
  • ▹Maintain and extend the team's documentation: responsibilities, runbooks, guides, onboarding, definitions of done, engagement closure templates; keep it accurate and reachable, link don't copy
  • ▹Get documentation used, not just written, and keep the ticket tracker the system of record so PoC and production status is readable without asking anyone
  • ▹Instrument the work: define and report the metrics that show whether delivery is getting faster and more reliable over time
  • ▹With development and research teams: convert recurring customer pain into prioritized platform work and represent the customer's technical reality in roadmap discussions
  • ▹With pre-sales: hold the scoping-to-execution boundary, push back on under-scoped engagements and feed feasibility signal back upstream
  • ▹With account management: make production handoffs uneventful and keep post-sales technical requests moving
  • ▹With business and leadership: give a straight read on account health, technical feasibility and capacity needs

Requirements

  • ▹3+ years managing technical teams, including performance management and difficult conversations
  • ▹Experience managing a customer-facing team: running a team whose work is visible to customers, on customer timelines, with customer escalations
  • ▹Strong ML knowledge: LLM architectures, fine-tuning approaches (SFT/LoRA, RL-based), evaluation design, and a working command of inference internals (quantization, KV-cache management, batching, routing, speculative decoding) and of the frameworks the team works in (vLLM, SGLang, TensorRT-LLM)
  • ▹Enough technical judgment to review someone else's benchmark and find the flaw in the methodology, not just in the conclusion
  • ▹Python strong enough to read and review your team's code
  • ▹Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences, from engineers to executives, including in front of enterprise customers under pressure
  • ▹Genuine tolerance for operational work: documentation, process design, reporting and the follow-through that makes them stick
  • ▹Comfort operating with ambiguity across distributed teams and timezones, and a bias toward writing things down

Nice to have

  • ▹References from both former managers and former direct reports
  • ▹Experience scaling a team through rapid growth (5 to 15+) without losing delivery quality
  • ▹Prior experience in a customer-facing technical function at a cloud, inference or AI infrastructure provider
  • ▹Experience defining process and documentation for a team that had none
  • ▹Hands-on background running LLMs in production and debugging inference workloads at the framework level
  • ▹Work with multimodal AI models (vision-language, speech)
  • ▹Proficiency with DevOps tooling (Docker, Kubernetes) and infrastructure-as-code
  • ▹Preferred technical stack: Python; vLLM, TensorRT-LLM, SGLang, Transformers, OpenAI/Anthropic SDKs; Kubernetes, Docker, Git; AWS (SageMaker, Bedrock), GCP (Vertex AI), Azure (Azure ML)

Soft skills

Excellent communicationHandling difficult conversationsTeam leadership and coachingAmbiguity tolerance

What we offer

  • ▹Competitive compensation
  • ▹Career growth and learning opportunities
  • ▹Flexibility and ownership
  • ▹Collaborative and innovative culture
  • ▹Opportunity to work on impactful AI projects
  • ▹International environment and talented teams

About the company

Nebius is building a full-stack AI cloud platform for the global AI economy, from data and model training through to production deployment. Headquartered in Amsterdam and listed on Nasdaq (NBIS), it has a team of 1,500+ with R&D hubs across Europe, the UK, North America and Israel.

Similar jobs