← Back to list
Job · Mid-level

Data Scientist — Agent Evaluations & Quality

Data Scientist • Mid-level • Remote • Full-time European Union EU/EMEA

Clera is building an AI executive assistant that operates across email, calendars, meetings and business software; this role owns the measurement system that determines whether the assistant is genuinely improving in ambiguous, real-world environments.

Responsibilities

  • Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces
  • Translate agent capabilities into explicit pass, partial-pass and failure criteria for complex multi-step tasks
  • Build representative gold datasets and regression suites covering real workflows, edge cases and adversarial scenarios
  • Define meaningful metrics: task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost and reliability
  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, track grader agreement
  • Compare models, prompts and implementations using rigorous offline experiments and production evidence
  • Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy
  • Turn production failures into regression cases and continuously close gaps in evaluation coverage
  • Build dashboards and release-quality signals
  • Recommend improvements to capability engineers and verify that fixes raise quality

Requirements

  • 4+ years in Applied Data Science or Machine Learning roles, with a track record of building and delivering evaluation systems, automated data pipelines or production ML infrastructure
  • Experience designing and implementing automated evaluation frameworks, success criteria and regression suites for complex AI/ML or agentic systems
  • Production-grade proficiency in Python and SQL, with experience building and maintaining automated analytical pipelines on large datasets
  • Applied statistical and experimental skills: significance testing, variance analysis, sampling to evaluate non-deterministic AI/ML systems
  • Experience developing labeled datasets, annotation guidelines and quality-control processes
  • Solid understanding of LLM agent behaviors: tool use, multi-step execution, retrieval, practical failure modes
  • Demonstrated ability to analyze model traces, tool calls and outputs to identify root causes
  • Experience using production telemetry and observability data to build dashboards and analyze real-world user outcomes

Nice to have

  • Hands-on experience with LLM-as-a-judge systems, model-based grading or AI benchmarking platforms
  • Experience shipping or operating production ML products, agentic systems or customer-facing consumer software
  • Experience reviewing and adapting public research benchmarks or academic evaluation methodologies to real-world product problems

Soft skills

Product-oriented mindset

About the company

Clera is building an AI executive assistant that operates across email, calendars, meetings and business software.

Similar jobs