03 — Investigation · Are the evals doing anything?
Do coding agents grade their own drawings accurately?
Decision it informs: Whether the runbook can trust the agent's own score as its judge check, or needs an external judge.
Key takeaway generated
No. Agents score their own drawings a median 9.6 points (of 40) above the external judge, and across 98 drawings the two scores correlate at only 0.28. As a gate, the self-score passes 98 of 98 runs at 28/40 while the judge passes 52, so the v2 runbook's self-score check cannot fail a weak drawing. Score with an external judge instead.
At a glance computed
Key figure computed
Every agent scores itself above the judge, by a median of 9.6 points
One row per agent × model, best round as picked by the agent. Filled dot: the external judge's total out of 40; ring: the agent's own. The line between them is the inflation. Click a row to open that lineage's runs.
Table view
| Agent · model | Cohort | Round | Self /40 | Judge /40 | Inflation | Judge spread |
|---|---|---|---|---|---|---|
| opencode · Gemini 3.8 Flash | v2 | v1 | 38.5 | 32.8 | +5.7 | 0.5 |
| claude-code · Opus 5.5 | v2 | v3 | 36 | 32.8 | +3.2 | 0.5 |
| hermes · Jev Router | v2 | v4 | 37 | 31.7 | +5.3 | 0.5 |
| gemini-cli · Gemini 3.5 Flash | v1 | v1 | 40 | 29.8 | +10.2 | 0.5 |
| opencode · Gemini 3.5 Flash | v1 | v2 | 40 | 29.3 | +10.7 | 1 |
| hermes · Fusion router | v1 | v1 | 34 | 28.8 | +5.2 | 2 |
| claude-code · Opus 4.7 | v1 | v3 | 36 | 28.7 | +7.3 | 1.5 |
| opencode · Fusion router | v1 | v2 | 36 | 28.3 | +7.7 | 1.5 |
| hermes · Gemini 3.5 Flash | v1 | v9 | 39.5 | 28.2 | +11.3 | 2 |
| opencode · GPT-5.6 Terra | v2 | v1 | 38 | 28 | +10 | 3.5 |
| hermes · GLM 5.1 | v1 | v2 | 36.5 | 27.3 | +9.2 | 0.5 |
| claude-code · Sonnet 4.6 | v1 | v5 | 37 | 26.2 | +10.8 | 1 |
| claude-code · Sonnet 5 | v2 | v4 | 36 | 25.3 | +10.7 | 1 |
| hermes · Sonnet 4.6 | v1 | v5 | 36 | 24 | +12 | 1 |
Evaluations computed
These runs used runbook v1, which wrote no validation report, so the scorecard is replayed here: the v2 runbook's code checks re-run on every stored SVG, plus two judge checks at 28/40 (the external judge, and the agent's own self-score, which is what the v2 runbook gates on). The overall row is the derived verdict shown on Runs.
| Evaluation | Passed | Rate | Share |
|---|---|---|---|
| Code check · svg-parses | 98/98 | 100% | |
| Code check · svg-size-under-50kb | 98/98 | 100% | |
| Code check · no-image-tags | 98/98 | 100% | |
| Code check · png-rendered | 98/98 | 100% | |
| Judge · external judge ≥ 28/40 | 52/98 | 53% | |
| Judge · agent self-score ≥ 28/40 (the v2 runbook's judge check) | 98/98 | 100% | |
| Overall · derived verdict (code checks + external judge) | 52/98 | 53% |
Findings generated
Every agent flatters itself, by a median of 9.6 points
high confidenceAll 14 of 14 lineages gave their best round a higher total than the judge did, by 3.2 to 12 points. Across every round the mean gap is 8.2.
A self-score barely predicts the judge's
high confidenceAcross all 98 drawings the correlation between self-score and judge score is 0.28; ranking the 14 lineages by self-score and by judge score gives a rank correlation of 0.20. The self-score cannot rank agents against each other.
As a pass/fail gate, the self-score never fails a run
high confidenceThe v2 runbook's judge check (self-score ≥ 28/40) would pass 98 of 98 stored runs; the external judge passes 52. That is 46 runs the self-score would wave through. No self threshold fixes it: the best one (38/40) agrees with the judge on only 66 of 98 runs.
The code checks guard the format, not the drawing
high confidenceAll 98 of 98 stored SVGs pass every replayable code check (parses, under 50 KB, no <image>, renders). They are worth keeping as cheap guards, but they never split runs, so today the only eval that separates a good pelican from a weak one is the external judge.
Bicycle is where agents are most generous
medium confidenceAveraged over every round, self-scores exceed the judge by 2.2 on Bicycle and 2 on Composition. The v1 Gemini Flash lineages gave 7 of their 30 drawings a perfect 40/40.
Proposed improvements generated
Replace the self-score judge check with an external judge
Would have failed 46 of the 98 stored runs that the self-score passes, matching the judge's 52/98.
Add a checklist item scored from the PNG by a second model
A cheap vision pass on 'is the pelican riding?' targets the most inflated axis without a full judge call.
Keep the code checks, but don't read their pass rate as quality
They passed 98/98; report them as guards and gate quality on the judge.
Caveats
- The judge is one model (Claude Opus 5.5) at temperature 0 with three samples; its own bias is unmeasured.
- The judge is also a contestant (v2 Opus 5.5 lineage).
- Self-scores come from each agent's report.md; a few v1 lineages had no image input and scored from SVG source.
- The runs used runbook v1, whose self-score was advisory; agents may score differently when the score gates their run.
Runs and provenance
Computed on 2026-09-28 from results.json and judge-raw.json. Task pelican-bicycle-svg. Every run is listed on 02 · Runs.
| Lineage | Cohort | Collection | Runs | Trajectories |
|---|---|---|---|---|
| claude-code · Opus 5.5 | v2 | jettygrowthteam | 5 | 5361e30a ba932f3e 35c47f4e 2e046a89 0e41c0a0 |
| claude-code · Sonnet 5 | v2 | jettygrowthteam | 5 | ea416761 b77ea720 6e98127a 240351f2 1925c665 |
| hermes · Jev Router | v2 | jettygrowthteam | 5 | 1a125d51 7a63ffb0 ded99127 3ba61c00 ef67915d |
| opencode · Gemini 3.8 Flash | v2 | jettygrowthteam | 1 | 225e57cd |
| opencode · GPT-5.6 Terra | v2 | jettygrowthteam | 5 | ccbdd583 699fefde ce858eaf 9fe62cbb de112455 |
| claude-code · Opus 4.7 | v1 | jettyio | 10 | c3121188 cb2a7f37 d4a107dc 8ec2e90c 4160dcf1 4848f8a3 df18c780 20f46c4d 4897cbca 38634df1 |
| claude-code · Sonnet 4.6 | v1 | jettyio | 11 | 7ddfcf08 414e644d b169371d 063dfed9 7ddc3c19 39febb11 ddf779bd 79fdee62 5f980543 2043d7e3 3a4c3357 |
| gemini-cli · Gemini 3.5 Flash | v1 | jettyio | 10 | f5cd6a31 c990e7c4 5d360721 4ed5de35 a652e4e3 eca90d64 f08684e3 ce2de30f 4e122677 24091968 |
| hermes · Gemini 3.5 Flash | v1 | jettyio | 10 | c118fb80 2128ffc1 2405db86 822a3bca 60502506 2c2a74e8 c165de92 d391e018 e6d4548b 3e21bf05 |
| hermes · Fusion router | v1 | jettyio | 3 | 2d0c10c4 b38f9569 cb138790 |
| hermes · GLM 5.1 | v1 | jettyio | 10 | c1f2da04 4b324d80 f443b217 7ce1468b ab8cb01d 9fd19804 83e8efe2 21b9b8c1 9f9e50be fff60669 |
| hermes · Sonnet 4.6 | v1 | jettyio | 10 | 697b57d4 acb2ba11 99a389e4 28f94566 8969a1a0 ebc95c4b fdf47c73 25323cb8 5b9cf6ae 34f10b3e |
| opencode · Gemini 3.5 Flash | v1 | jettyio | 10 | 67836d67 660d2a66 5cd01e9e b130c9c2 d94ff2e1 4a378be2 e3059c87 70898b5c b56c0c03 2a46de05 |
| opencode · Fusion router | v1 | jettyio | 3 | 40df25cb bfa3612b 9776d557 |