About the Role Join a small technical research and engineering team building benchmarks for evaluating AI agents on realistic, domain-specific workflows. You will own the design and implementation of rigorous evaluations that help technical teams understand how agents perform in real-world tasks. What You'll Do Design, implement, and maintain benchmarks for evaluating AI agents on domain-specific tasks. Work with subject-matter experts to translate real workflows into benchmark tasks and evaluation criteria. Build and operate reliable infrastructure for running models and agents against tasks at scale. Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes. Validate whether benchmark results align with real-world performance and technical user needs. Write clear documentation and reports for research and engineering audiences. What We're Looking For Two to four years of experience in research engineering, machine learning engineering, or a related technical role, including at least two years building AI benchmarks, evaluation infrastructure, or agent environments. Strong Python skills and practical experience with Docker and Linux. Experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models. Experience collaborating with subject-matter experts and analyzing workflows across technical or business domains. Experience developing metrics, statistical analyses, or validation studies for evaluations. Strong technical writing, attention to detail, and ability to work independently in an early-stage environment. A technical educational background; experience with reinforcement learning pipelines, published evaluation work, or widely used benchmarks is a plus. Compensation & Benefits Salary range: USD 150,000 to 250,000 annually. Visa sponsorship is available. Location On-site in Singapore, Singapore.
Ähnliche Stellen

Stelle
Research Engineer, Synthetic Data
Clera
AI/ML
133 695–222 826 €/Jahr
brutto
🏢 Vor Ort
🗣️ EN

Stelle
Software Research Engineer(PAYE)
Huawei R&D UK
AI/MLLlm
💰 Gehalt: keine Angabe
🏢 Vor Ort
Edinburgh
🗣️ EN

Stelle
Research Engineer, Life Sciences
Anthropic
AI/MLLlm
311 956–445 652 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
· Medior
Mid-Level Research Engineer, QC Automation
Clera
Llm
133 695–222 826 €/Jahr
brutto
🏢 Vor Ort
🗣️ EN
Stelle
Research Engineer
Afterquery
AI/MLLlm
187 174–401 086 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN
Stelle
Research Engineer, World Models
Waabi
AI/MLKubeflowLlm
138 152–239 761 €/Jahr
brutto
🏢 Vor Ort
Toronto
🗣️ EN