← Back to list

Job
· Lead
Lead Research Engineer, Data Quality
AI Research Scientist
• Lead
• On-site
• Full-time
•
San Francisco, USA
About the Role This is a senior individual contributor and team lead position at an early-stage AI infrastructure startup building the tooling and data pipelines that power frontier RL-based agent training. You'll own the strategy and execution of data quality systems end-to-end — from evaluation frameworks to synthetic data validation — and help shape internal research culture around what makes agent training data genuinely useful. What You'll Do Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data, benchmarks, and domain-specific workflows. Define data quality strategy by building QC systems, enforcing standards, and designing experiments to grade agent outputs. Develop new methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing. Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows. Turn qualitative research insights into production systems — internal tools, dashboards, validation pipelines, and feedback loops. Build internal research taste around what makes agent training data realistic, learnable, diverse, and reliable — not just superficially correct. Mentor research engineers to maintain a high bar for technical rigor, clarity, and execution speed. What We're Looking For 5+ years of experience in research or data quality engineering, specifically building systems for AI/ML data evaluation. Demonstrated track record leading technical teams or projects on ambiguous problems from definition through to iteration. Advanced proficiency in Python, Docker, and Linux environments. Experience building QC systems, evals, benchmarks, synthetic data pipelines, or model evaluation infrastructure. Strong intuition for the characteristics of high-quality training data for AI agents — realistic, learnable, diverse, reliable, and useful. Ability to design metrics, experiments, and QA/QC processes, not just execute them. Experience working with subject-matter experts to capture domain judgment and convert it into scalable review or generation systems. Strong written communication skills with the ability to explain methodology clearly to technical and non-technical audiences. Comfort navigating complex systems involving domain experts, vendors, model outputs, graders, and infrastructure. Prior experience in an early-stage startup environment; able to work independently and move quickly. Compensation & Benefits Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available. Location On-site in San Francisco, CA, United States .
Similar jobs

Job
Research Engineer, Synthetic Data
Clera
AI/MLLlm
$150,000–$250,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
Research Engineer, Benchmarks
Clera
Llm
$150,000–$250,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
Research Engineer
Clera
$150,000–$250,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
Research Engineer, QC Automation
Clera
Llm
$150,000–$250,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
Offensive Cyber Research Engineer
Twenty
AI/ML
+4
$159,000–$263,000/yr
gross
🏢 On-site
Washington DC
🗣️ EN

Job
Medical AI Researcher
Clera
AI/ML
+1
$150,000–$230,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN