03 — Investigation · Did the change help?
Does hill-climbing on self-scores actually improve the drawing?
Decision it informs: Whether to keep running multi-round self-scored hill climbs, and how to steer and stop them.
Key takeaway generated
Mostly not. Across 13 lineages with three or more rounds, self-scores rose 2.3 points from round 1 to their best, but the judge scored the last round 0.9 points lower than the first on average, and 8 of 13 ended below where they started. The climb steers by the self-score, which barely tracks the judge; steer it by the judge and stop when the judge stops improving.
At a glance computed
Key figure computed
hermes · Sonnet 4.6's self-score held while the judge's score slid
Score out of 40 for each round's final drawing. The highlighted lineage shows its self-score and the judge's score; grey lines are every other lineage's judge score; the dashed line is the 28/40 threshold. Pick another lineage to highlight it; hover or use ←/→ for values.
Table view
| Lineage | v1 | v2 | v3 | v4 | v5 | v6 | v7 | v8 | v9 | v10 | v11 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| opencode · Gemini 3.8 Flash · judge | 32.8 | – | – | – | – | – | – | – | – | – | – |
| opencode · Gemini 3.8 Flash · self | 38.5 | – | – | – | – | – | – | – | – | – | – |
| claude-code · Opus 5.5 · judge | 32.7 | 33 | 32.8 | 33 | 33.3 | – | – | – | – | – | – |
| claude-code · Opus 5.5 · self | 35 | 35 | 36 | 36 | 35 | – | – | – | – | – | – |
| hermes · Jev Router · judge | 29.8 | 31.5 | 30.7 | 31.7 | 30.8 | – | – | – | – | – | – |
| hermes · Jev Router · self | 36 | 35 | 34.5 | 37 | 36 | – | – | – | – | – | – |
| gemini-cli · Gemini 3.5 Flash · judge | 29.8 | 31 | 29.3 | 31.7 | 32.3 | 30.5 | 31 | 31.7 | 32 | 31.2 | – |
| gemini-cli · Gemini 3.5 Flash · self | 40 | 38.5 | 38.4 | 40 | 39.5 | 39.5 | 38.7 | 40 | 37.5 | 39.2 | – |
| opencode · Gemini 3.5 Flash · judge | 28.2 | 29.3 | 28.7 | 29 | 29 | 30.2 | 30.2 | 31.5 | 31.5 | 26.8 | – |
| opencode · Gemini 3.5 Flash · self | 39.5 | 40 | 39 | 40 | 38.6 | 38.7 | 40 | 38 | 39 | 40 | – |
| hermes · Fusion router · judge | 28.8 | 27.7 | 29 | – | – | – | – | – | – | – | – |
| hermes · Fusion router · self | 34 | 34 | 34 | – | – | – | – | – | – | – | – |
| claude-code · Opus 4.7 · judge | 29 | 29.3 | 28.7 | 29.7 | 28 | 27.3 | 27.5 | 27 | 28 | 27.7 | – |
| claude-code · Opus 4.7 · self | 33 | 35 | 36 | 32 | 34.5 | 36 | 33 | 33 | 33.5 | 34 | – |
| opencode · Fusion router · judge | 29.3 | 28.3 | 28 | – | – | – | – | – | – | – | – |
| opencode · Fusion router · self | 34 | 36 | 34 | – | – | – | – | – | – | – | – |
| hermes · Gemini 3.5 Flash · judge | 28.5 | 28.2 | 26.5 | 28.7 | 30 | 26.8 | 27 | 26.3 | 28.2 | 26.5 | – |
| hermes · Gemini 3.5 Flash · self | 39 | 38 | 37 | 39 | 38.8 | 38 | 38.5 | 38 | 39.5 | 38.5 | – |
| opencode · GPT-5.6 Terra · judge | 28 | 27.5 | 27.2 | 27.3 | 29.8 | – | – | – | – | – | – |
| opencode · GPT-5.6 Terra · self | 38 | 37.2 | 38 | 37.5 | 38 | – | – | – | – | – | – |
| hermes · GLM 5.1 · judge | 28.7 | 27.3 | 28.8 | 28.7 | 27.7 | 25.7 | 25 | 24.7 | 24.2 | 25.7 | – |
| hermes · GLM 5.1 · self | 32 | 36.5 | 36 | 33 | 33 | 34.5 | 36 | 34 | 33 | 36.5 | – |
| claude-code · Sonnet 4.6 · judge | 26.5 | 25.2 | 24.5 | 24.5 | 26.2 | 26.5 | 27.7 | 26.3 | 25.8 | 28 | 25.3 |
| claude-code · Sonnet 4.6 · self | 31 | 32 | 34 | 36 | 37 | 36 | 34 | 34 | 33 | 35 | 35 |
| claude-code · Sonnet 5 · judge | 26.5 | 27.3 | 27 | 25.3 | 26.2 | – | – | – | – | – | – |
| claude-code · Sonnet 5 · self | 28 | 34 | 34 | 36 | 36 | – | – | – | – | – | – |
| hermes · Sonnet 4.6 · judge | 26.5 | 27 | 24 | 25 | 24 | 22.2 | 19.2 | 18 | 18.7 | 19.8 | – |
| hermes · Sonnet 4.6 · self | 32 | 34 | 34 | 34 | 36 | 36 | 36 | 36 | 36 | 36 | – |
Evaluations computed
These runs used runbook v1, which wrote no validation report, so the scorecard is replayed here: the v2 runbook's code checks re-run on every stored SVG, plus two judge checks at 28/40 (the external judge, and the agent's own self-score, which is what the v2 runbook gates on). The overall row is the derived verdict shown on Runs.
| Evaluation | Passed | Rate | Share |
|---|---|---|---|
| Code check · svg-parses | 98/98 | 100% | |
| Code check · svg-size-under-50kb | 98/98 | 100% | |
| Code check · no-image-tags | 98/98 | 100% | |
| Code check · png-rendered | 98/98 | 100% | |
| Judge · external judge ≥ 28/40 | 52/98 | 53% | |
| Judge · agent self-score ≥ 28/40 (the v2 runbook's judge check) | 98/98 | 100% | |
| Overall · derived verdict (code checks + external judge) | 52/98 | 53% |
Findings generated
Self-scores climb; the judge's mostly don't
high confidenceAcross 13 lineages, the best self-score was 2.3 points above round 1 on average, while the judge's last-round score moved -0.9. 8 lineages ended below round 1 by the judge, 5 above.
Long climbs drift down
medium confidenceThe ten-round v1 climbs moved -2 on average by the judge (6 of 7 ended lower); the shorter climbs moved +0.3. The steepest slide: hermes · Sonnet 4.6, from 26.5 to 19.8, while its self-score held at 36.
The judge's best drawing usually comes early
medium confidenceThe judge's favourite round came at a median of round 4, and the last round scored 2.1 points below the judge's favourite on average. The best improvement was opencode · GPT-5.6 Terra (+1.8), a five-round v2 climb.
Even hand-edited runbooks didn't move the judge
high confidenceJon's eleven hand edits took the claude-code + Sonnet 4.6 self-score from 31 to a peak of 37, while the judge went from 26.5 to 25.3. The edits optimised what the agent saw in its own scores.
Drawings get bigger, not better
medium confidenceThe SVG grew from round 1 to the last round in 12 of 13 lineages (median +56%), and across all drawings file size correlates with the judge's score at -0.21. The climb adds detail the rubric doesn't reward.
Proposed improvements generated
Steer the climb by the external judge and keep its best round
Keeping the judge's best round instead of the last would have been worth 2.1 points per lineage on these runs.
Stop after two rounds without a judge improvement
The judge's best came at a median round 4; later rounds mostly spent money moving the score down.
Cap scene additions in the runbook
Tell the agent to fix anatomy and contact points, not add scenery, so size growth stops standing in for progress.
Caveats
- Each round starts from the previous best SVG, so rounds within a lineage are not independent samples.
- One lineage per setup; the direction is consistent but the per-lineage deltas are single observations.
- The judge scores the round's final drawing only, not the inner self-critique passes.
- v1 and v2 climbs differ in length (10 vs 5 rounds) and in model era.
Runs and provenance
Computed on 2026-09-28 from results.json and judge-raw.json. Task pelican-bicycle-svg. Every run is listed on 02 · Runs.
| Lineage | Cohort | Collection | Runs | Trajectories |
|---|---|---|---|---|
| claude-code · Opus 5.5 | v2 | jettygrowthteam | 5 | 5361e30a ba932f3e 35c47f4e 2e046a89 0e41c0a0 |
| claude-code · Sonnet 5 | v2 | jettygrowthteam | 5 | ea416761 b77ea720 6e98127a 240351f2 1925c665 |
| hermes · Jev Router | v2 | jettygrowthteam | 5 | 1a125d51 7a63ffb0 ded99127 3ba61c00 ef67915d |
| opencode · Gemini 3.8 Flash | v2 | jettygrowthteam | 1 | 225e57cd |
| opencode · GPT-5.6 Terra | v2 | jettygrowthteam | 5 | ccbdd583 699fefde ce858eaf 9fe62cbb de112455 |
| claude-code · Opus 4.7 | v1 | jettyio | 10 | c3121188 cb2a7f37 d4a107dc 8ec2e90c 4160dcf1 4848f8a3 df18c780 20f46c4d 4897cbca 38634df1 |
| claude-code · Sonnet 4.6 | v1 | jettyio | 11 | 7ddfcf08 414e644d b169371d 063dfed9 7ddc3c19 39febb11 ddf779bd 79fdee62 5f980543 2043d7e3 3a4c3357 |
| gemini-cli · Gemini 3.5 Flash | v1 | jettyio | 10 | f5cd6a31 c990e7c4 5d360721 4ed5de35 a652e4e3 eca90d64 f08684e3 ce2de30f 4e122677 24091968 |
| hermes · Gemini 3.5 Flash | v1 | jettyio | 10 | c118fb80 2128ffc1 2405db86 822a3bca 60502506 2c2a74e8 c165de92 d391e018 e6d4548b 3e21bf05 |
| hermes · Fusion router | v1 | jettyio | 3 | 2d0c10c4 b38f9569 cb138790 |
| hermes · GLM 5.1 | v1 | jettyio | 10 | c1f2da04 4b324d80 f443b217 7ce1468b ab8cb01d 9fd19804 83e8efe2 21b9b8c1 9f9e50be fff60669 |
| hermes · Sonnet 4.6 | v1 | jettyio | 10 | 697b57d4 acb2ba11 99a389e4 28f94566 8969a1a0 ebc95c4b fdf47c73 25323cb8 5b9cf6ae 34f10b3e |
| opencode · Gemini 3.5 Flash | v1 | jettyio | 10 | 67836d67 660d2a66 5cd01e9e b130c9c2 d94ff2e1 4a378be2 e3059c87 70898b5c b56c0c03 2a46de05 |
| opencode · Fusion router | v1 | jettyio | 3 | 40df25cb bfa3612b 9776d557 |