---
title: "Does hill-climbing on self-scores actually improve the drawing? | Pelican benchmark"
url: https://evaljetty.com/pelicans/investigations/hill-climbing.html
description: "Mostly not. Across 13 lineages with three or more rounds, self-scores rose 2.3 points from round 1 to their best, but the judge scored the last round 0.9 points lower than the first on average, and 8 of 13 ended below where they started."
updated: 2026-09-28
publisher: Jetty (https://jetty.io)
---

03 — Investigation · Did the change help?

# Does hill-climbing on self-scores actually improve the drawing?

task pelican-bicycle-svg13 climbs of 3+ rounds98 runsorchestrator hill\_climb.pycomputed 2026-09-28

**Decision it informs:** Whether to keep running multi-round self-scored hill climbs, and how to steer and stop them.

## Key takeaway generated

Mostly not. Across 13 lineages with three or more rounds, self-scores rose 2.3 points from round 1 to their best, but the judge scored the last round 0.9 points lower than the first on average, and 8 of 13 ended below where they started. The climb steers by the self-score, which barely tracks the judge; steer it by the judge and stop when the judge stops improving.

## At a glance computed

Runs

98

14 agent × model lineages

Spend

$11.90

10 metered runs + $2.77 judge; 11 runs not metered

Time

5.5 min

median per run (32 timed)

Pass rate

53%

52/98 derived verdicts pass

## Key figure computed

### hermes · Sonnet 4.6's self-score held while the judge's score slid

Score out of 40 for each round's final drawing. The highlighted lineage shows its self-score and the judge's score; grey lines are every other lineage's judge score; the dashed line is the 28/40 threshold. Pick another lineage to highlight it; hover or use ←/→ for values.

Highlightopencode · Gemini 3.8 Flash (v2)claude-code · Opus 5.5 (v2)hermes · Jev Router (v2)gemini-cli · Gemini 3.5 Flashopencode · Gemini 3.5 Flashhermes · Fusion routerclaude-code · Opus 4.7opencode · Fusion routerhermes · Gemini 3.5 Flashopencode · GPT-5.6 Terra (v2)hermes · GLM 5.1claude-code · Sonnet 4.6claude-code · Sonnet 5 (v2)hermes · Sonnet 4.6

selected, self-scoreselected, judgeother lineages, judgepass ≥ 28

Table view

| Lineage | v1 | v2 | v3 | v4 | v5 | v6 | v7 | v8 | v9 | v10 | v11 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| opencode · Gemini 3.8 Flash · judge | 32.8 | – | – | – | – | – | – | – | – | – | – |
| opencode · Gemini 3.8 Flash · self | 38.5 | – | – | – | – | – | – | – | – | – | – |
| claude-code · Opus 5.5 · judge | 32.7 | 33 | 32.8 | 33 | 33.3 | – | – | – | – | – | – |
| claude-code · Opus 5.5 · self | 35 | 35 | 36 | 36 | 35 | – | – | – | – | – | – |
| hermes · Jev Router · judge | 29.8 | 31.5 | 30.7 | 31.7 | 30.8 | – | – | – | – | – | – |
| hermes · Jev Router · self | 36 | 35 | 34.5 | 37 | 36 | – | – | – | – | – | – |
| gemini-cli · Gemini 3.5 Flash · judge | 29.8 | 31 | 29.3 | 31.7 | 32.3 | 30.5 | 31 | 31.7 | 32 | 31.2 | – |
| gemini-cli · Gemini 3.5 Flash · self | 40 | 38.5 | 38.4 | 40 | 39.5 | 39.5 | 38.7 | 40 | 37.5 | 39.2 | – |
| opencode · Gemini 3.5 Flash · judge | 28.2 | 29.3 | 28.7 | 29 | 29 | 30.2 | 30.2 | 31.5 | 31.5 | 26.8 | – |
| opencode · Gemini 3.5 Flash · self | 39.5 | 40 | 39 | 40 | 38.6 | 38.7 | 40 | 38 | 39 | 40 | – |
| hermes · Fusion router · judge | 28.8 | 27.7 | 29 | – | – | – | – | – | – | – | – |
| hermes · Fusion router · self | 34 | 34 | 34 | – | – | – | – | – | – | – | – |
| claude-code · Opus 4.7 · judge | 29 | 29.3 | 28.7 | 29.7 | 28 | 27.3 | 27.5 | 27 | 28 | 27.7 | – |
| claude-code · Opus 4.7 · self | 33 | 35 | 36 | 32 | 34.5 | 36 | 33 | 33 | 33.5 | 34 | – |
| opencode · Fusion router · judge | 29.3 | 28.3 | 28 | – | – | – | – | – | – | – | – |
| opencode · Fusion router · self | 34 | 36 | 34 | – | – | – | – | – | – | – | – |
| hermes · Gemini 3.5 Flash · judge | 28.5 | 28.2 | 26.5 | 28.7 | 30 | 26.8 | 27 | 26.3 | 28.2 | 26.5 | – |
| hermes · Gemini 3.5 Flash · self | 39 | 38 | 37 | 39 | 38.8 | 38 | 38.5 | 38 | 39.5 | 38.5 | – |
| opencode · GPT-5.6 Terra · judge | 28 | 27.5 | 27.2 | 27.3 | 29.8 | – | – | – | – | – | – |
| opencode · GPT-5.6 Terra · self | 38 | 37.2 | 38 | 37.5 | 38 | – | – | – | – | – | – |
| hermes · GLM 5.1 · judge | 28.7 | 27.3 | 28.8 | 28.7 | 27.7 | 25.7 | 25 | 24.7 | 24.2 | 25.7 | – |
| hermes · GLM 5.1 · self | 32 | 36.5 | 36 | 33 | 33 | 34.5 | 36 | 34 | 33 | 36.5 | – |
| claude-code · Sonnet 4.6 · judge | 26.5 | 25.2 | 24.5 | 24.5 | 26.2 | 26.5 | 27.7 | 26.3 | 25.8 | 28 | 25.3 |
| claude-code · Sonnet 4.6 · self | 31 | 32 | 34 | 36 | 37 | 36 | 34 | 34 | 33 | 35 | 35 |
| claude-code · Sonnet 5 · judge | 26.5 | 27.3 | 27 | 25.3 | 26.2 | – | – | – | – | – | – |
| claude-code · Sonnet 5 · self | 28 | 34 | 34 | 36 | 36 | – | – | – | – | – | – |
| hermes · Sonnet 4.6 · judge | 26.5 | 27 | 24 | 25 | 24 | 22.2 | 19.2 | 18 | 18.7 | 19.8 | – |
| hermes · Sonnet 4.6 · self | 32 | 34 | 34 | 34 | 36 | 36 | 36 | 36 | 36 | 36 | – |

## Evaluations computed

These runs used runbook v1, which wrote no validation report, so the scorecard is replayed here: the v2 runbook's code checks re-run on every stored SVG, plus two judge checks at 28/40 (the external judge, and the agent's own self-score, which is what the v2 runbook gates on). The overall row is the derived verdict shown on Runs.

| Evaluation | Passed | Rate | Share |
| --- | --- | --- | --- |
| Code check · svg-parses | 98/98 | 100% |  |
| Code check · svg-size-under-50kb | 98/98 | 100% |  |
| Code check · no-image-tags | 98/98 | 100% |  |
| Code check · png-rendered | 98/98 | 100% |  |
| Judge · external judge ≥ 28/40 | 52/98 | 53% |  |
| Judge · agent self-score ≥ 28/40 (the v2 runbook's judge check) | 98/98 | 100% |  |
| Overall · derived verdict (code checks + external judge) | 52/98 | 53% |  |

## Findings generated

### Self-scores climb; the judge's mostly don't

high confidence

Across 13 lineages, the best self-score was 2.3 points above round 1 on average, while the judge's last-round score moved -0.9. 8 lineages ended below round 1 by the judge, 5 above.

13 lineagesself +2.3judge -0.9

### Long climbs drift down

medium confidence

The ten-round v1 climbs moved -2 on average by the judge (6 of 7 ended lower); the shorter climbs moved +0.3. The steepest slide: hermes · Sonnet 4.6, from 26.5 to 19.8, while its self-score held at 36.

hermes-sonnet v1 697b57d4hermes-sonnet v10 34f10b3e

### The judge's best drawing usually comes early

medium confidence

The judge's favourite round came at a median of round 4, and the last round scored 2.1 points below the judge's favourite on average. The best improvement was opencode · GPT-5.6 Terra (+1.8), a five-round v2 climb.

median judge-best round 4v2-opencode-gpt56 +1.8

### Even hand-edited runbooks didn't move the judge

high confidence

Jon's eleven hand edits took the claude-code + Sonnet 4.6 self-score from 31 to a peak of 37, while the judge went from 26.5 to 25.3. The edits optimised what the agent saw in its own scores.

claude-sonnet v1 7ddfcf08claude-sonnet v5 7ddc3c19claude-sonnet v11 3a4c3357

### Drawings get bigger, not better

medium confidence

The SVG grew from round 1 to the last round in 12 of 13 lineages (median +56%), and across all drawings file size correlates with the judge's score at -0.21. The climb adds detail the rubric doesn't reward.

median growth +56%size vs judge r = -0.21

## Proposed improvements generated

runbook

### Steer the climb by the external judge and keep its best round

Keeping the judge's best round instead of the last would have been worth 2.1 points per lineage on these runs.

backed by: F1, F3

config

### Stop after two rounds without a judge improvement

The judge's best came at a median round 4; later rounds mostly spent money moving the score down.

backed by: F2, F3

runbook

### Cap scene additions in the runbook

Tell the agent to fix anatomy and contact points, not add scenery, so size growth stops standing in for progress.

backed by: F5

## Caveats

- Each round starts from the previous best SVG, so rounds within a lineage are not independent samples.
- One lineage per setup; the direction is consistent but the per-lineage deltas are single observations.
- The judge scores the round's final drawing only, not the inner self-critique passes.
- v1 and v2 climbs differ in length (10 vs 5 rounds) and in model era.

**Runs and provenance**

Computed on 2026-09-28 from [results.json](https://evaljetty.com/pelicans/data/results.json) and [judge-raw.json](https://evaljetty.com/pelicans/data/judge-raw.json). Task pelican-bicycle-svg. Every run is listed on [02 · Runs](https://evaljetty.com/pelicans/runs/index.html#all-runs).

| Lineage | Cohort | Collection | Runs | Trajectories |
| --- | --- | --- | --- | --- |
| [claude-code · Opus 5.5](https://evaljetty.com/pelicans/runs/agents/v2-claude-opus55.html) | v2 | jettygrowthteam | 5 | 5361e30a ba932f3e 35c47f4e 2e046a89 0e41c0a0 |
| [claude-code · Sonnet 5](https://evaljetty.com/pelicans/runs/agents/v2-claude-sonnet5.html) | v2 | jettygrowthteam | 5 | ea416761 b77ea720 6e98127a 240351f2 1925c665 |
| [hermes · Jev Router](https://evaljetty.com/pelicans/runs/agents/v2-hermes-jev.html) | v2 | jettygrowthteam | 5 | 1a125d51 7a63ffb0 ded99127 3ba61c00 ef67915d |
| [opencode · Gemini 3.8 Flash](https://evaljetty.com/pelicans/runs/agents/v2-opencode-gemini38.html) | v2 | jettygrowthteam | 1 | 225e57cd |
| [opencode · GPT-5.6 Terra](https://evaljetty.com/pelicans/runs/agents/v2-opencode-gpt56.html) | v2 | jettygrowthteam | 5 | ccbdd583 699fefde ce858eaf 9fe62cbb de112455 |
| [claude-code · Opus 4.7](https://evaljetty.com/pelicans/runs/agents/claude-opus.html) | v1 | jettyio | 10 | c3121188 cb2a7f37 d4a107dc 8ec2e90c 4160dcf1 4848f8a3 df18c780 20f46c4d 4897cbca 38634df1 |
| [claude-code · Sonnet 4.6](https://evaljetty.com/pelicans/runs/agents/claude-sonnet.html) | v1 | jettyio | 11 | 7ddfcf08 414e644d b169371d 063dfed9 7ddc3c19 39febb11 ddf779bd 79fdee62 5f980543 2043d7e3 3a4c3357 |
| [gemini-cli · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/gemini-cli.html) | v1 | jettyio | 10 | f5cd6a31 c990e7c4 5d360721 4ed5de35 a652e4e3 eca90d64 f08684e3 ce2de30f 4e122677 24091968 |
| [hermes · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/hermes-flash.html) | v1 | jettyio | 10 | c118fb80 2128ffc1 2405db86 822a3bca 60502506 2c2a74e8 c165de92 d391e018 e6d4548b 3e21bf05 |
| [hermes · Fusion router](https://evaljetty.com/pelicans/runs/agents/hermes-fusion.html) | v1 | jettyio | 3 | 2d0c10c4 b38f9569 cb138790 |
| [hermes · GLM 5.1](https://evaljetty.com/pelicans/runs/agents/hermes-glm51.html) | v1 | jettyio | 10 | c1f2da04 4b324d80 f443b217 7ce1468b ab8cb01d 9fd19804 83e8efe2 21b9b8c1 9f9e50be fff60669 |
| [hermes · Sonnet 4.6](https://evaljetty.com/pelicans/runs/agents/hermes-sonnet.html) | v1 | jettyio | 10 | 697b57d4 acb2ba11 99a389e4 28f94566 8969a1a0 ebc95c4b fdf47c73 25323cb8 5b9cf6ae 34f10b3e |
| [opencode · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/opencode.html) | v1 | jettyio | 10 | 67836d67 660d2a66 5cd01e9e b130c9c2 d94ff2e1 4a378be2 e3059c87 70898b5c b56c0c03 2a46de05 |
| [opencode · Fusion router](https://evaljetty.com/pelicans/runs/agents/opencode-fusion.html) | v1 | jettyio | 3 | 40df25cb bfa3612b 9776d557 |

---
Source: https://evaljetty.com/pelicans/investigations/hill-climbing.html · Evals by Jetty · Run your own evals: https://jetty.io/?utm_source=evaljetty&utm_medium=llms
