Judge vs self-score
The runbook asks each agent to score its own drawing on four axes. Those numbers are useful for steering one agent's climb, but they can't rank agents against each other: in v1, Gemini Flash gave itself 40/40 while Sonnet topped out at 37. So for v2 we re-scored every round's final drawing, old and new, with one fixed vision judge.
The gap, agent by agent
Best round per lineage, as chosen by the agent. Filled dot: the judge's score. Ring: the agent's own score. The line between them is how much the agent flattered itself.
| Agent | Judge rank | Self rank | Judge | Self | Inflation | Judge spread | Judge's favourite round |
|---|---|---|---|---|---|---|---|
| opencode · Gemini 3.8 Flash v2 · Sep 2026 | #1 | #4 | 32.8 | 38.5 | +5.7 | 0.5 | v1 · 32.8 |
| claude-code · Opus 5.5 v2 · Sep 2026 | #2 | #13 | 32.8 | 36 | +3.2 | 0.5 | v5 · 33.3 |
| hermes · Jev Router v2 · Sep 2026 | #3 | #7 | 31.7 | 37 | +5.3 | 0.5 | v4 · 31.7 |
| gemini-cli · Gemini 3.5 Flash v1 · May 2026 | #4 | #1 | 29.8 | 40 | +10.2 | 0.5 | v5 · 32.3 |
| opencode · Gemini 3.5 Flash v1 · May 2026 | #5 | #2 | 29.3 | 40 | +10.7 | 1 | v9 · 31.5 |
| hermes · Fusion router v1 · May 2026 | #6 | #14 | 28.8 | 34 | +5.2 | 2 | v3 · 29 |
| claude-code · Opus 4.7 v1 · May 2026 | #7 | #9 | 28.7 | 36 | +7.3 | 1.5 | v4 · 29.7 |
| opencode · Fusion router v1 · May 2026 | #8 | #11 | 28.3 | 36 | +7.7 | 1.5 | v1 · 29.3 |
| hermes · Gemini 3.5 Flash v1 · May 2026 | #9 | #3 | 28.2 | 39.5 | +11.3 | 2 | v5 · 30 |
| opencode · GPT-5.6 Terra v2 · Sep 2026 | #10 | #5 | 28 | 38 | +10 | 3.5 | v5 · 29.8 |
| hermes · GLM 5.1 v1 · May 2026 | #11 | #8 | 27.3 | 36.5 | +9.2 | 0.5 | v3 · 28.8 |
| claude-code · Sonnet 4.6 v1 · May 2026 | #12 | #6 | 26.2 | 37 | +10.8 | 1 | v10 · 28 |
| claude-code · Sonnet 5 v2 · Sep 2026 | #13 | #12 | 25.3 | 36 | +10.7 | 1 | v2 · 27.3 |
| hermes · Sonnet 4.6 v1 · May 2026 | #14 | #10 | 24 | 36 | +12 | 1 | v2 · 27 |
Judge's favourite round: the round the judge scored highest across the whole climb, which isn't always the agent's own pick.
What the gap tells us
Self-scores climb. The judge's mostly don't.
Across 13 lineages with three or more rounds, the agents' self-scores rose by 2.3 points on average from round 1 to their best, while the judge's score for the last round was 0.9 points lower than round 1. 8 of 13 ended below where they started by the judge's measure. The steepest slide: hermes · Sonnet 4.6, from 26.5 to 19.8.
A self-score barely predicts the judge's
Over all 98 scored drawings the correlation between self-score and judge score is 0.28. The hill climb steers by the self-score (it targets the weakest self-scored axis), so it is steering by an instrument that only loosely tracks what a viewer sees. That's the likeliest reason long climbs drift.
Some agents are much kinder to themselves
hermes · Sonnet 4.6 gave its best round 36/40; the judge gave it 24 (+12). The most honest was claude-code · Opus 5.5: 36 self vs 32.8 judge (+3.2). Gemini Flash lineages in v1 awarded themselves 39.5 to 40.
Current models draw better pelicans
The v2 cohort averages 30.1/40 from the judge against 27.9 for v1. Best v2: opencode · Gemini 3.8 Flash at 32.8. Best v1: gemini-cli · Gemini 3.5 Flash at 29.8. v2 did this in five rounds instead of ten.
Where the gap lives
Averaged over every lineage's best round. Composition is where agents are most generous with themselves: they know the pelican is meant to be riding, so they score the intent, not the pixels.
Self vs judge, mean by axis
0–10, all 14 lineages
Show as table
| Axis | Mean self | Mean judge | Inflation |
|---|---|---|---|
| Pelican | 9.25 | 7.32 | +1.93 |
| Bicycle | 9.21 | 6.81 | +2.40 |
| Composition | 9.54 | 7.32 | +2.22 |
| Polish | 9.18 | 7.21 | +1.97 |
v1 vs v2 models, judge only
Mean judge score by axis for each cohort
Show as table
| Cohort | Pelican | Bicycle | Composition | Polish |
|---|---|---|---|---|
| v1 | 6.96 | 6.72 | 7.13 | 7.04 |
| v2 | 7.97 | 6.97 | 7.67 | 7.53 |
How the judge works
Every round's final.svg is rasterized with rsvg-convert -w 800 on a white background (librsvg, the same rasterizer the runbook tells agents to use) and sent to the judge with the rubric below. The system prompt, rubric and model are fixed across all 98 drawings. Each drawing gets three independent calls at temperature 0 and we average them. Raw responses, including token usage, are published in judge-raw.json.
A single LLM judge has biases of its own. It's one fixed yardstick rather than ground truth, and it's the same yardstick for every agent, which is the point. The judge model (Opus 5.5) is also one of the v2 contestants. If anything that should make it lenient toward its own lineage, so weigh that row accordingly.
The image below is a rasterized SVG that a coding agent hand-wrote for the prompt:
"Generate an SVG of a pelican riding a bicycle."
Score it on four axes, each an integer or half-integer from 0 to 10:
1. pelican - Pelican recognizability. Would a stranger immediately say "that's a pelican"? Look for the long beak with a throat pouch, pelican body shape, eye, wings, legs/feet. 10 = unmistakably a pelican; 5 = a generic bird; 0 = no bird.
2. bicycle - Bicycle recognizability and correctness. Two wheels (with spokes), a plausible frame (top tube, down tube, seat tube), handlebars, seat, pedals/crank. 10 = a correct, well-formed bicycle; 5 = recognizable but structurally wrong; 0 = no bicycle.
3. composition - Is the pelican clearly RIDING the bicycle? Body on the seat, feet on or near the pedals, wing tips on the handlebars, parts connected rather than floating, sensible scale. 10 = convincingly riding; 5 = next to or loosely attached; 0 = unrelated.
4. polish - Aesthetic polish: line quality, color, balance, absence of rendering glitches, clutter or stray shapes. 10 = professional illustration quality; 5 = passable clip-art; 0 = broken.
Calibration: be strict. Reserve 9-10 for genuinely excellent work with no visible defects. Most competent attempts land between 5 and 8. Judge only what is visible in the image, not what was intended.
Reply with ONLY a JSON object, no prose outside it:
{"pelican": <0-10>, "bicycle": <0-10>, "composition": <0-10>, "polish": <0-10>, "notes": "<one or two sentences naming the main strengths and defects>"}