← Zurück zur Liste
Stelle · Senior

Senior Python Engineer - LLM Code Evaluation

Sonstige • Senior • Remote • Vollzeit • Europäische Union EU/EMEA

A contract role of 10-20 hours a week evaluating the quality of interactions with modern coding agents (OpenAI Codex, Claude Code). You do not write production code: you judge whether the model thinks like a great engineer.

Responsibilities

  • ▹Evaluate AI-generated coding interactions end-to-end
  • ▹Judge whether outputs are useful, correct at a high level and aligned with how a strong engineer would think
  • ▹Assess the quality of explanations and reasoning, not just code
  • ▹Distinguish between levels of response quality (e.g. what makes something a 2 vs 4)
  • ▹Provide clear, opinionated feedback on what worked, what did not and what felt off or misleading
  • ▹Help define what great looks like when interacting with tools like Cursor

Requirements

  • ▹Staff / Principal-level engineer (or equivalent experience)
  • ▹Strong background in TypeScript / JavaScript or Python
  • ▹Hands-on experience using OpenAI Codex, Claude Code and Cursor
  • ▹Deep familiarity with modern AI-assisted dev workflows
  • ▹Able to evaluate code without fully executing or deeply reviewing every line
  • ▹Comfortable giving direct, opinionated feedback with a high bar for good engineering

Nice to have

  • ▹Experience with tools like Cursor or similar AI-first IDEs
  • ▹Prior exposure to prompt design or evaluation workflows
  • ▹Experience mentoring senior engineers or defining engineering standards

Soft skills

Rigorous, reasoned judgmentCommunication

What we offer

  • ▹Rate: $100-$200/hour
  • ▹About 10-20 hours/week
  • ▹Duration through early May, with possible extension
  • ▹Start ASAP

Ähnliche Stellen