[← back to projects]

[in progress]

AI Agent Evaluation

Benchmark environments for AI coding agents

contract (ongoing)

Tasks that test whether an AI agent can build a real, rule-enforcing web app — plus the tooling to quality-check them in hours instead of days.

AI Agent Evaluation screenshot 1
AI Agent Evaluation screenshot 2

The problem

Evaluating AI coding agents needs tasks that are hard but fair: the reference solution must score perfectly, the model must land in a narrow difficulty band, and grading must reward real server-side behaviour rather than a convincing UI.

What I built

As a contractor with Turing, I designed full-stack web-app tasks (brief, seed data, reference app, LLM-judge rubric) and built the tooling around them: a launcher that reproduces the grader's runtime, a static QC linter, deterministic packaging, and a multi-agent review loop.

Core features

  • -Owner-voiced briefs that describe business incidents instead of handing over a specification
  • -Reference apps with sign-in, roles, and server-enforced business rules
  • -LLM-judge rubrics run through a headless browser against the live app
  • -Runtime-faithful launcher with real restarts for persistence checks
  • -Static linter covering about 20 families of platform rules

Engineering highlights

  • -Grading by request replay: the judge records the app's own requests and replays them as another user, so a hidden button never counts as enforcement
  • -Blind, parallel graders with a proposer/attacker red team per fix; a task is done after two consecutive clean runs
  • -Deterministic judging: binary numbered checks, a fixed clock, and every pinned figure verified by both the reference app and an independent re-implementation
  • -Deterministic packaging with a round-trip diff, so what ships is exactly what was tested

Tech stack

Node.jsExpressSQLitePythonPlaywrightBashLLM-as-judgeClaude Code agents

Impact

  • -Brought one task's model score from about 98% to about 42% while the reference solution stayed at 100%
  • -QC iteration moved from manual review to an automated, repeatable pipeline

Want something similar built?

Happy to talk it through — drop me a line.

[get in touch]