← Back to list
Job · Mid-level

DevOps Engineer - AI Model Evaluator

DevOps / SRE • Mid-level • Remote • Full-time Germany Germany

Contract role using frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks across cloud, Kubernetes, and CI/CD environments.

Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks
  • Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, and infrastructure automation
  • Identify bugs, edge cases, reliability issues, and failure modes in model outputs
  • Compare outputs from multiple frontier models to assess strengths and weaknesses
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios

Requirements

  • 2+ years of professional DevOps, SRE, or Cloud Engineering experience
  • Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling
  • Regular use of AI coding agents like Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions

Nice to have

  • Experience supporting production-scale systems

Soft skills

Professional engineering judgmentAnalytical and comparative thinking

About the company

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, its investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.

Similar jobs