evals
Run evals on Jetty
v2 · Sep 2026 · jettygrowthteam · auto-evolved runbook (weakest-axis targeting)

claude-code · Opus 5.5

claude-code driving anthropic/claude-opus-5.5 through OpenRouter, 5 hill-climb rounds from the v1 runbook.

agent claude-code · model anthropic/claude-opus-5.5

Judge, best round
32.8/40
round v3
Self-score, best round
36/40
inflation +3.2
Rounds
5
~$4.55 est. spend
claude-code · Opus 5.5 best round
best round v3 · 35c47f4e ↗

Self-score vs judge, by axis

Best round v3, 0–10 per axis

agent self-scoreexternal judge

Judge: A clearly readable white pelican with a large orange throat pouch, crest and eye. It sits on the saddle, its feet reach the pedals and its dark wingtips grip the handlebars. The bicycle is well formed, with spoked wheels, a chainring and a chain, but the frame geometry is slightly odd: the seat tube meets the chainring off-center and the fork is thin. There are minor issues too: the legs are stiff and awkwardly jointed, the body overlaps the saddle, and the pouch sits oddly detached from the upper beak.

Score by round

Out of 40. Each round starts from the previous round's best SVG.

self-scorejudge

Every round

Click a drawing to open it in the head-to-head viewer. The outlined card is the agent's best round.

claude-code · Opus 5.5 round v1
v1judge 32.7 · self 35 5361e30a ↗runbook
claude-code · Opus 5.5 round v2
v2judge 33 · self 35 ba932f3e ↗runbook
claude-code · Opus 5.5 round v3
v3 · bestjudge 32.8 · self 36 35c47f4e ↗runbook
claude-code · Opus 5.5 round v4
v4judge 33 · self 36 2e046a89 ↗runbook
claude-code · Opus 5.5 round v5
v5judge 33.3 · self 35 0e41c0a0 ↗runbook

Scores per round

RoundPelicanBicycleCompositionPolishSelf totalJudge totalJudge spreadMinutesTrajectory
v18.5self 97.7self 98.5self 98self 83532.70.54.35361e30a ↗
v28.5self 98self 98.5self 88self 9353304.6ba932f3e ↗
v38.5self 98self 98.3self 108self 83632.80.54.835c47f4e ↗
v48.5self 98self 98.5self 98self 9363304.52e046a89 ↗
v58.5self 98self 98.5self 88.3self 93533.30.56.40e41c0a0 ↗

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.