← Zurück zur Liste

Stelle
Research Engineer, Benchmarks
AI Research Scientist
• Vor Ort
• Vollzeit
•
San Francisco, USA
About the Role This role sits at the core of a small, technical team building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand true agent capability. Your work directly shapes the credibility and rigor of a product at the frontier of AI evaluation. What You'll Do Design, implement, and own the quality of internal benchmarks for evaluating frontier AI agents on domain-specific tasks. Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks. Build reliable infrastructure to run models and agents against benchmark tasks at scale. Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes. Validate that benchmark performance correlates meaningfully with real-world agent behavior and customer needs. Write clear documentation and benchmark reports that make results legible and credible to technical audiences. What We're Looking For 2–4 years of experience in research engineering or ML engineering, with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments. Strong proficiency in Python, Docker, and Linux environments for research or production infrastructure. Experience designing and running evaluations for AI agents or large language models. Experience building and operating infrastructure to reliably run AI models or agents against benchmark tasks at scale. Experience developing metrics and validation studies to assess benchmark difficulty, reliability, and real-world correlation. Ability to collaborate with subject-matter experts and translate domain workflows into evaluation criteria. Strong attention to detail and a habit of spotting subtle inconsistencies and edge cases. Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces. Excellent written communication skills for cross-functional collaboration across time zones. Bonus: Published papers or technical writing on AI benchmarking, model evaluation, or failure modes; experience with RL training pipelines or widely used public benchmark projects. Compensation & Benefits Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available. Location On-site in San Francisco, CA, United States .
Ähnliche Stellen

Stelle
Research Engineer, Synthetic Data
Clera
AI/MLLlm
128 138–213 564 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
Research Engineer, QC Automation
Clera
Llm
128 138–213 564 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
Research Engineer
Clera
128 138–213 564 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
Research Engineer, Life Sciences
Anthropic
AI/MLLlm
298 990–427 128 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
Research Engineer / Research Scientist — Robotics & Physical AI
Clera
AI/MLCppLlm
85 426–256 277 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
AI Research Resident
Normal Computing
AI/MLLlm
💰 Gehalt: keine Angabe
🏢 Vor Ort
New York City
🗣️ EN