evals
Run evals on Jetty

03 — Investigation · Did the change help?

Does hill-climbing on self-scores actually improve the drawing?

task pelican-bicycle-svg13 climbs of 3+ rounds98 runsorchestrator hill_climb.pycomputed 2026-09-28

Decision it informs: Whether to keep running multi-round self-scored hill climbs, and how to steer and stop them.

Key takeaway generated

Mostly not. Across 13 lineages with three or more rounds, self-scores rose 2.3 points from round 1 to their best, but the judge scored the last round 0.9 points lower than the first on average, and 8 of 13 ended below where they started. The climb steers by the self-score, which barely tracks the judge; steer it by the judge and stop when the judge stops improving.

At a glance computed

Runs
98
14 agent × model lineages
Spend
$11.90
10 metered runs + $2.77 judge; 11 runs not metered
Time
5.5 min
median per run (32 timed)
Pass rate
53%
52/98 derived verdicts pass

Key figure computed

hermes · Sonnet 4.6's self-score held while the judge's score slid

Score out of 40 for each round's final drawing. The highlighted lineage shows its self-score and the judge's score; grey lines are every other lineage's judge score; the dashed line is the 28/40 threshold. Pick another lineage to highlight it; hover or use ←/→ for values.

selected, self-scoreselected, judgeother lineages, judgepass ≥ 28
Table view
Lineagev1v2v3v4v5v6v7v8v9v10v11
opencode · Gemini 3.8 Flash · judge32.8––––––––––
opencode · Gemini 3.8 Flash · self38.5––––––––––
claude-code · Opus 5.5 · judge32.73332.83333.3––––––
claude-code · Opus 5.5 · self3535363635––––––
hermes · Jev Router · judge29.831.530.731.730.8––––––
hermes · Jev Router · self363534.53736––––––
gemini-cli · Gemini 3.5 Flash · judge29.83129.331.732.330.53131.73231.2–
gemini-cli · Gemini 3.5 Flash · self4038.538.44039.539.538.74037.539.2–
opencode · Gemini 3.5 Flash · judge28.229.328.7292930.230.231.531.526.8–
opencode · Gemini 3.5 Flash · self39.540394038.638.740383940–
hermes · Fusion router · judge28.827.729––––––––
hermes · Fusion router · self343434––––––––
claude-code · Opus 4.7 · judge2929.328.729.72827.327.5272827.7–
claude-code · Opus 4.7 · self3335363234.536333333.534–
opencode · Fusion router · judge29.328.328––––––––
opencode · Fusion router · self343634––––––––
hermes · Gemini 3.5 Flash · judge28.528.226.528.73026.82726.328.226.5–
hermes · Gemini 3.5 Flash · self3938373938.83838.53839.538.5–
opencode · GPT-5.6 Terra · judge2827.527.227.329.8––––––
opencode · GPT-5.6 Terra · self3837.23837.538––––––
hermes · GLM 5.1 · judge28.727.328.828.727.725.72524.724.225.7–
hermes · GLM 5.1 · self3236.536333334.536343336.5–
claude-code · Sonnet 4.6 · judge26.525.224.524.526.226.527.726.325.82825.3
claude-code · Sonnet 4.6 · self3132343637363434333535
claude-code · Sonnet 5 · judge26.527.32725.326.2––––––
claude-code · Sonnet 5 · self2834343636––––––
hermes · Sonnet 4.6 · judge26.52724252422.219.21818.719.8–
hermes · Sonnet 4.6 · self32343434363636363636–

Evaluations computed

These runs used runbook v1, which wrote no validation report, so the scorecard is replayed here: the v2 runbook's code checks re-run on every stored SVG, plus two judge checks at 28/40 (the external judge, and the agent's own self-score, which is what the v2 runbook gates on). The overall row is the derived verdict shown on Runs.

EvaluationPassedRateShare
Code check · svg-parses98/98100%
Code check · svg-size-under-50kb98/98100%
Code check · no-image-tags98/98100%
Code check · png-rendered98/98100%
Judge · external judge ≥ 28/4052/9853%
Judge · agent self-score ≥ 28/40 (the v2 runbook's judge check)98/98100%
Overall · derived verdict (code checks + external judge)52/9853%

Findings generated

Self-scores climb; the judge's mostly don't

high confidence

Across 13 lineages, the best self-score was 2.3 points above round 1 on average, while the judge's last-round score moved -0.9. 8 lineages ended below round 1 by the judge, 5 above.

13 lineagesself +2.3judge -0.9

Long climbs drift down

medium confidence

The ten-round v1 climbs moved -2 on average by the judge (6 of 7 ended lower); the shorter climbs moved +0.3. The steepest slide: hermes · Sonnet 4.6, from 26.5 to 19.8, while its self-score held at 36.

hermes-sonnet v1 697b57d4hermes-sonnet v10 34f10b3e

The judge's best drawing usually comes early

medium confidence

The judge's favourite round came at a median of round 4, and the last round scored 2.1 points below the judge's favourite on average. The best improvement was opencode · GPT-5.6 Terra (+1.8), a five-round v2 climb.

median judge-best round 4v2-opencode-gpt56 +1.8

Even hand-edited runbooks didn't move the judge

high confidence

Jon's eleven hand edits took the claude-code + Sonnet 4.6 self-score from 31 to a peak of 37, while the judge went from 26.5 to 25.3. The edits optimised what the agent saw in its own scores.

claude-sonnet v1 7ddfcf08claude-sonnet v5 7ddc3c19claude-sonnet v11 3a4c3357

Drawings get bigger, not better

medium confidence

The SVG grew from round 1 to the last round in 12 of 13 lineages (median +56%), and across all drawings file size correlates with the judge's score at -0.21. The climb adds detail the rubric doesn't reward.

median growth +56%size vs judge r = -0.21

Proposed improvements generated

runbook

Steer the climb by the external judge and keep its best round

Keeping the judge's best round instead of the last would have been worth 2.1 points per lineage on these runs.

backed by: F1, F3

config

Stop after two rounds without a judge improvement

The judge's best came at a median round 4; later rounds mostly spent money moving the score down.

backed by: F2, F3

runbook

Cap scene additions in the runbook

Tell the agent to fix anatomy and contact points, not add scenery, so size growth stops standing in for progress.

backed by: F5

Caveats

  • Each round starts from the previous best SVG, so rounds within a lineage are not independent samples.
  • One lineage per setup; the direction is consistent but the per-lineage deltas are single observations.
  • The judge scores the round's final drawing only, not the inner self-critique passes.
  • v1 and v2 climbs differ in length (10 vs 5 rounds) and in model era.
Runs and provenance

Computed on 2026-09-28 from results.json and judge-raw.json. Task pelican-bicycle-svg. Every run is listed on 02 · Runs.

LineageCohortCollectionRunsTrajectories
claude-code · Opus 5.5v2jettygrowthteam55361e30a ba932f3e 35c47f4e 2e046a89 0e41c0a0
claude-code · Sonnet 5v2jettygrowthteam5ea416761 b77ea720 6e98127a 240351f2 1925c665
hermes · Jev Routerv2jettygrowthteam51a125d51 7a63ffb0 ded99127 3ba61c00 ef67915d
opencode · Gemini 3.8 Flashv2jettygrowthteam1225e57cd
opencode · GPT-5.6 Terrav2jettygrowthteam5ccbdd583 699fefde ce858eaf 9fe62cbb de112455
claude-code · Opus 4.7v1jettyio10c3121188 cb2a7f37 d4a107dc 8ec2e90c 4160dcf1 4848f8a3 df18c780 20f46c4d 4897cbca 38634df1
claude-code · Sonnet 4.6v1jettyio117ddfcf08 414e644d b169371d 063dfed9 7ddc3c19 39febb11 ddf779bd 79fdee62 5f980543 2043d7e3 3a4c3357
gemini-cli · Gemini 3.5 Flashv1jettyio10f5cd6a31 c990e7c4 5d360721 4ed5de35 a652e4e3 eca90d64 f08684e3 ce2de30f 4e122677 24091968
hermes · Gemini 3.5 Flashv1jettyio10c118fb80 2128ffc1 2405db86 822a3bca 60502506 2c2a74e8 c165de92 d391e018 e6d4548b 3e21bf05
hermes · Fusion routerv1jettyio32d0c10c4 b38f9569 cb138790
hermes · GLM 5.1v1jettyio10c1f2da04 4b324d80 f443b217 7ce1468b ab8cb01d 9fd19804 83e8efe2 21b9b8c1 9f9e50be fff60669
hermes · Sonnet 4.6v1jettyio10697b57d4 acb2ba11 99a389e4 28f94566 8969a1a0 ebc95c4b fdf47c73 25323cb8 5b9cf6ae 34f10b3e
opencode · Gemini 3.5 Flashv1jettyio1067836d67 660d2a66 5cd01e9e b130c9c2 d94ff2e1 4a378be2 e3059c87 70898b5c b56c0c03 2a46de05
opencode · Fusion routerv1jettyio340df25cb bfa3612b 9776d557

Ready to stop guessing if your outputs are good?

Get Started Book a 15-minute walkthrough
Connect your agent to Jetty →