A pelican learns to ride
This is a worked example of hill-climbing a runbook on Jetty. We took Simon Willison's pelican-on-a-bicycle prompt, wrote it up as a runbook with a self-critique loop, and ran it across 14 agent and model pairs. Every run is a trajectory you can open.
What you're looking at
Jetty runs a runbook, a markdown file with instructions, constraints and a scoring rubric, inside a pinned sandbox. A coding agent (claude-code, opencode, hermes, gemini-cli) follows it and every step is captured as a trajectory. To improve the result you edit the runbook and run it again.
v1 (May 2026) asked what moves the score more: iterating the runbook or iterating the model. The answer was about the same. Editing the runbook by hand gained six self-score points; swapping the model gained five. v2 (September 2026) re-runs the climb with five current models and adds an external judge, because v1 also showed that self-scores can't be compared across models.
The task
Hand-write a pure-XML SVG of a pelican riding a bicycle. No <image> tags, under 50 KB, viewBox around 800×600. Both subjects must be unmistakable, and the pelican has to be riding the bike rather than floating next to it.
The rubric
Four axes, 0–10 each: Pelican (would a stranger say "pelican"?), Bicycle (two spoked wheels, a real frame), Composition (is it actually riding?) and Polish (line, color, balance). The agent renders its SVG, looks at the PNG, scores it, and redraws, three times per run. It keeps the best round.
The hill climb
Each hill-climb round is one Jetty run. Between rounds, the orchestrator (a small Python script) embeds the previous round's best SVG into the runbook as the new starting point and rewrites the description to target the lowest-scoring axis. v1 ran 10 rounds per agent (3 for the Fusion sweeps); v2 ran 5. The one exception is the reference lineage: Jon read each claude-code + Sonnet trajectory and edited the runbook by hand, eleven times.
The judge (new in v2)
Every round's final drawing, v1 and v2 alike, is re-scored by one fixed vision judge: Claude Opus 5.5 through OpenRouter, temperature 0, three samples averaged, same rubric. See Judge vs self-score for the prompt and the numbers.
Why this prompt
Simon Willison has run the pelican prompt against nearly every major model release since 2024. He picked it on purpose:
They shouldn't be able to draw anything at all. But they can generate code… and SVG is code.
Simon Willison, June 2025
Most importantly: pelicans can't ride bicycles. They're the wrong shape!
Simon Willison
Everyone needs their own benchmark.
Simon Willison
Drawing a pelican on a bicycle in SVG makes a text model reason in code about shape, anatomy and spatial composition all at once. Pelicans can't ride bicycles, so the result can't be copied from training data. And because SVG allows comments, the models often narrate what they're drawing as they go.
The v1 hand-curated climb, in three steps
Three picks from Jon's eleven runbook edits (claude-code + Sonnet 4.6). The judge scores were added in v2, after the fact.
The edit: Use the v2 SVG verbatim as the round-1 input. Don't start from scratch.
What happened: Round 1 jumped from 28 to 33, five points just from skipping the cold start. The single highest-leverage edit in the sequence.
Lesson: When a runbook can carry a working artifact as a seed, it should.
The edit: Close the 13 px right-wing-to-grip gap. Drop the left-wing tips to y ≈ 230.
What happened: Composition hit 10/10 for the first time. The pelican grips the bars. Sonnet's peak self-score, never beaten later.
Lesson: At the top of the curve, the asks become measurements.
The edit: Add motion lines, wind tufts and a sun. Extend the right-wing covert lines to the wrist.
What happened: A richer scene, but composition slipped back to 9. The new elements added clutter the rubric couldn't reward.
Lesson: More isn't better when the rubric doesn't measure scene density.
What we changed for v2
- Current models, one provider path. Everything routes through OpenRouter with a single collection secret, OPENROUTER_API_KEY, so anyone can rerun it with one key.
- Five rounds, not ten. Most v1 lineages had plateaued by round five.
- An external judge over every drawing from both cohorts, so the leaderboard ranks drawings rather than egos.
- Same runbook. The v2 runbook is v1's baseline with the model defaults updated. The orchestrator logic is unchanged apart from the collection and round count.
Want to try it? Run it yourself.