← Back to list
Job · Senior

Senior Machine Learning Engineer, Agent Oversight

AI / ML Engineer • Senior • On-site • Full-time United States San Francisco, USA

Senior ML engineering role on Scale AI's Agent Oversight team (part of the Scale Generative AI Platform), owning the end-to-end lifecycle that keeps production agents reliable and continuously improving — from observability tooling and evaluation frameworks to improvement loops.

Stack

Responsibilities

  • Build or contribute to observability into agent behavior in production — instrumentation that shows what an agent is actually doing, not just whether it succeeded or failed
  • Design evaluation methodologies and metrics for agentic applications, and work with the platform team to run them automatically at scale across customer use cases
  • Build, ship and own ML systems that detect drift, anomalies or misalignment in production agent behavior — from first prototype through reliable production operation
  • Design and run rigorous experiments to validate model and agent performance improvements before they ship
  • Work alongside software engineers where the work intersects with broader platform infrastructure
  • Collaborate closely with product managers, customers, data annotators and forward-deployed engineers to translate enterprise and government requirements into platform capabilities
  • Contribute to novel methods for agent evaluation and improvement, or build ML systems that hold up reliably at scale in production

Requirements

  • 5+ years of experience as an ML engineer or applied scientist, ideally on a production ML or LLM-powered system — not just consuming a third-party ML API within a feature
  • Strong grounding in at least two of: building/scaling evaluation, monitoring or continuous-learning infrastructure; agent-system design experience (architecture, orchestration, tool use); developing new methods, reward models or model training/fine-tuning approaches
  • Hands-on experience with LLMs and agent architectures — tool use, planning, multi-agent orchestration
  • Comfortable partnering with software engineers to productionize research and experimental work, not just deliver a one-off analysis
  • Rigorous approach to experimentation: clear hypotheses, real statistical grounding, results that hold up under scrutiny
  • Track record of collaborating across functions (product, forward-deployed engineering, etc.) to navigate ambiguous requirements and bring them to production
  • Gives direct, substantive feedback on designs and code, takes it the same way, and mentors others as they grow

Nice to have

  • Experience building or contributing to RLHF, SFT or other fine-tuning/RL workflows, reward modeling, or verifiable-reward systems
  • Experience with model or systems optimization (latency, cost or inference efficiency)
  • Published research, open-source contributions, or patents in agentic systems, LLMs, or applied ML
  • Experience working in regulated or enterprise contexts
  • Track record of taking a novel method from prototype to something running reliably in production
  • Experience reviewing others' technical designs or mentoring engineers at a senior/staff level

Soft skills

Rigorous, hypothesis-driven experimentationCross-functional collaboration with product, customers and forward-deployed engineeringDirect feedback cultureMentoringComfort navigating ambiguity

What we offer

  • Compensation includes base salary, equity and benefits
  • Comprehensive health, dental and vision coverage
  • Retirement benefits
  • Learning and development stipend
  • Generous PTO
  • Commuter stipend for eligible roles

About the company

Scale AI's mission is to develop reliable AI systems for the world's most important decisions. The Applied Intelligence Systems team, part of the Scale Generative AI Platform (SGP), builds the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design and evaluation, serving both commercial and public-sector customers.

Similar jobs