AI Agent Evaluation
Benchmark environments for AI coding agents
Tasks that test whether an AI agent can build a real, rule-enforcing web app — plus the tooling to quality-check them in hours instead of days.


The problem
Evaluating AI coding agents needs tasks that are hard but fair: the reference solution must score perfectly, the model must land in a narrow difficulty band, and grading must reward real server-side behaviour rather than a convincing UI.
What I built
As a contractor with Turing, I designed full-stack web-app tasks (brief, seed data, reference app, LLM-judge rubric) and built the tooling around them: a launcher that reproduces the grader's runtime, a static QC linter, deterministic packaging, and a multi-agent review loop.
Core features
- -Owner-voiced briefs that describe business incidents instead of handing over a specification
- -Reference apps with sign-in, roles, and server-enforced business rules
- -LLM-judge rubrics run through a headless browser against the live app
- -Runtime-faithful launcher with real restarts for persistence checks
- -Static linter covering about 20 families of platform rules
Engineering highlights
- -Grading by request replay: the judge records the app's own requests and replays them as another user, so a hidden button never counts as enforcement
- -Blind, parallel graders with a proposer/attacker red team per fix; a task is done after two consecutive clean runs
- -Deterministic judging: binary numbered checks, a fixed clock, and every pinned figure verified by both the reference app and an independent re-implementation
- -Deterministic packaging with a round-trip diff, so what ships is exactly what was tested
Tech stack
Node.jsExpressSQLitePythonPlaywrightBashLLM-as-judgeClaude Code agents
Impact
- -Brought one task's model score from about 98% to about 42% while the reference solution stayed at 100%
- -QC iteration moved from manual review to an automated, repeatable pipeline