evals
Run evals on Jetty

03 — Investigation · Are the evals doing anything?

Do coding agents grade their own drawings accurately?

task pelican-bicycle-svg98 scored drawings14 lineagesjudge anthropic/claude-opus-5.5 · T=0 · 3 samplescomputed 2026-09-28

Decision it informs: Whether the runbook can trust the agent's own score as its judge check, or needs an external judge.

Key takeaway generated

No. Agents score their own drawings a median 9.6 points (of 40) above the external judge, and across 98 drawings the two scores correlate at only 0.28. As a gate, the self-score passes 98 of 98 runs at 28/40 while the judge passes 52, so the v2 runbook's self-score check cannot fail a weak drawing. Score with an external judge instead.

At a glance computed

Runs
98
14 agent × model lineages
Spend
$11.90
10 metered runs + $2.77 judge; 11 runs not metered
Time
5.5 min
median per run (32 timed)
Pass rate
53%
52/98 derived verdicts pass

Key figure computed

Every agent scores itself above the judge, by a median of 9.6 points

One row per agent × model, best round as picked by the agent. Filled dot: the external judge's total out of 40; ring: the agent's own. The line between them is the inflation. Click a row to open that lineage's runs.

external judgeagent self-score
Table view
Agent · modelCohortRoundSelf /40Judge /40InflationJudge spread
opencode · Gemini 3.8 Flashv2v138.532.8+5.70.5
claude-code · Opus 5.5v2v33632.8+3.20.5
hermes · Jev Routerv2v43731.7+5.30.5
gemini-cli · Gemini 3.5 Flashv1v14029.8+10.20.5
opencode · Gemini 3.5 Flashv1v24029.3+10.71
hermes · Fusion routerv1v13428.8+5.22
claude-code · Opus 4.7v1v33628.7+7.31.5
opencode · Fusion routerv1v23628.3+7.71.5
hermes · Gemini 3.5 Flashv1v939.528.2+11.32
opencode · GPT-5.6 Terrav2v13828+103.5
hermes · GLM 5.1v1v236.527.3+9.20.5
claude-code · Sonnet 4.6v1v53726.2+10.81
claude-code · Sonnet 5v2v43625.3+10.71
hermes · Sonnet 4.6v1v53624+121

Evaluations computed

These runs used runbook v1, which wrote no validation report, so the scorecard is replayed here: the v2 runbook's code checks re-run on every stored SVG, plus two judge checks at 28/40 (the external judge, and the agent's own self-score, which is what the v2 runbook gates on). The overall row is the derived verdict shown on Runs.

EvaluationPassedRateShare
Code check · svg-parses98/98100%
Code check · svg-size-under-50kb98/98100%
Code check · no-image-tags98/98100%
Code check · png-rendered98/98100%
Judge · external judge ≥ 28/4052/9853%
Judge · agent self-score ≥ 28/40 (the v2 runbook's judge check)98/98100%
Overall · derived verdict (code checks + external judge)52/9853%

Findings generated

Every agent flatters itself, by a median of 9.6 points

high confidence

All 14 of 14 lineages gave their best round a higher total than the judge did, by 3.2 to 12 points. Across every round the mean gap is 8.2.

hermes-sonnet +12v2-claude-opus55 +3.298 scored drawings

A self-score barely predicts the judge's

high confidence

Across all 98 drawings the correlation between self-score and judge score is 0.28; ranking the 14 lineages by self-score and by judge score gives a rank correlation of 0.20. The self-score cannot rank agents against each other.

r = 0.28 over 98 drawingsrank r = 0.20

As a pass/fail gate, the self-score never fails a run

high confidence

The v2 runbook's judge check (self-score ≥ 28/40) would pass 98 of 98 stored runs; the external judge passes 52. That is 46 runs the self-score would wave through. No self threshold fixes it: the best one (38/40) agrees with the judge on only 66 of 98 runs.

self ≥ 28: 98/98judge ≥ 28: 52/98best threshold 38: 66/98

The code checks guard the format, not the drawing

high confidence

All 98 of 98 stored SVGs pass every replayable code check (parses, under 50 KB, no <image>, renders). They are worth keeping as cheap guards, but they never split runs, so today the only eval that separates a good pelican from a weak one is the external judge.

svg-parses: 98/98svg-size-under-50kb: 98/98no-image-tags: 98/98

Bicycle is where agents are most generous

medium confidence

Averaged over every round, self-scores exceed the judge by 2.2 on Bicycle and 2 on Composition. The v1 Gemini Flash lineages gave 7 of their 30 drawings a perfect 40/40.

pelican: +2bicycle: +2.2composition: +2polish: +2

Proposed improvements generated

eval

Replace the self-score judge check with an external judge

Would have failed 46 of the 98 stored runs that the self-score passes, matching the judge's 52/98.

backed by: F1, F3

runbook

Add a checklist item scored from the PNG by a second model

A cheap vision pass on 'is the pelican riding?' targets the most inflated axis without a full judge call.

backed by: F5

eval

Keep the code checks, but don't read their pass rate as quality

They passed 98/98; report them as guards and gate quality on the judge.

backed by: F4

Caveats

  • The judge is one model (Claude Opus 5.5) at temperature 0 with three samples; its own bias is unmeasured.
  • The judge is also a contestant (v2 Opus 5.5 lineage).
  • Self-scores come from each agent's report.md; a few v1 lineages had no image input and scored from SVG source.
  • The runs used runbook v1, whose self-score was advisory; agents may score differently when the score gates their run.
Runs and provenance

Computed on 2026-09-28 from results.json and judge-raw.json. Task pelican-bicycle-svg. Every run is listed on 02 · Runs.

LineageCohortCollectionRunsTrajectories
claude-code · Opus 5.5v2jettygrowthteam55361e30a ba932f3e 35c47f4e 2e046a89 0e41c0a0
claude-code · Sonnet 5v2jettygrowthteam5ea416761 b77ea720 6e98127a 240351f2 1925c665
hermes · Jev Routerv2jettygrowthteam51a125d51 7a63ffb0 ded99127 3ba61c00 ef67915d
opencode · Gemini 3.8 Flashv2jettygrowthteam1225e57cd
opencode · GPT-5.6 Terrav2jettygrowthteam5ccbdd583 699fefde ce858eaf 9fe62cbb de112455
claude-code · Opus 4.7v1jettyio10c3121188 cb2a7f37 d4a107dc 8ec2e90c 4160dcf1 4848f8a3 df18c780 20f46c4d 4897cbca 38634df1
claude-code · Sonnet 4.6v1jettyio117ddfcf08 414e644d b169371d 063dfed9 7ddc3c19 39febb11 ddf779bd 79fdee62 5f980543 2043d7e3 3a4c3357
gemini-cli · Gemini 3.5 Flashv1jettyio10f5cd6a31 c990e7c4 5d360721 4ed5de35 a652e4e3 eca90d64 f08684e3 ce2de30f 4e122677 24091968
hermes · Gemini 3.5 Flashv1jettyio10c118fb80 2128ffc1 2405db86 822a3bca 60502506 2c2a74e8 c165de92 d391e018 e6d4548b 3e21bf05
hermes · Fusion routerv1jettyio32d0c10c4 b38f9569 cb138790
hermes · GLM 5.1v1jettyio10c1f2da04 4b324d80 f443b217 7ce1468b ab8cb01d 9fd19804 83e8efe2 21b9b8c1 9f9e50be fff60669
hermes · Sonnet 4.6v1jettyio10697b57d4 acb2ba11 99a389e4 28f94566 8969a1a0 ebc95c4b fdf47c73 25323cb8 5b9cf6ae 34f10b3e
opencode · Gemini 3.5 Flashv1jettyio1067836d67 660d2a66 5cd01e9e b130c9c2 d94ff2e1 4a378be2 e3059c87 70898b5c b56c0c03 2a46de05
opencode · Fusion routerv1jettyio340df25cb bfa3612b 9776d557

Ready to stop guessing if your outputs are good?

Get Started Book a 15-minute walkthrough
Connect your agent to Jetty →