evals
Run evals on Jetty

03 — Investigation · Compare setups

Which agent and model draw the best pelican on a bicycle?

task pelican-bicycle-svg98 runs14 lineagescollections jettyio + jettygrowthteamjudge 28/40 thresholdcomputed 2026-09-28

Decision it informs: Which agent and model to use for SVG drawing work, and whether a newer model is worth switching to.

Key takeaway generated

opencode · Gemini 3.8 Flash and claude-code · Opus 5.5 tie for the best pelican at 32.8/40 from the external judge, ahead of hermes · Jev Router at 31.7. Current models beat last spring's: the v2 cohort averages 30.1 against 27.9 for v1. The gaps are a few points on one judge with one lineage per setup, so rank the top three as a group and re-run before choosing between them.

At a glance computed

Runs
98
14 agent × model lineages
Spend
$11.90
10 metered runs + $2.77 judge; 11 runs not metered
Time
5.5 min
median per run (32 timed)
Pass rate
53%
52/98 derived verdicts pass

Key figure computed

opencode · Gemini 3.8 Flash and claude-code · Opus 5.5 share the top; v2 models fill most of it

One bar per agent × model: the external judge's score (out of 40) for the round the agent picked as its best, coloured by cohort. The ring marks the judge's favourite round in that lineage's climb; the dashed line is the 28/40 pass threshold. Click a bar to open that lineage's runs.

v1 cohort (May 2026)v2 cohort (Sep 2026)judge's best round in the climbpass threshold (28)
Table view
Agent · modelCohortAgent's pickJudge /40Self /40Judge's best round
opencode · Gemini 3.8 Flashgoogle/gemini-3.8-flashv2v132.838.5v1 · 32.8
claude-code · Opus 5.5anthropic/claude-opus-5.5v2v332.836v5 · 33.3
hermes · Jev Routertypesafe/jev-routerv2v431.737v4 · 31.7
gemini-cli · Gemini 3.5 Flashgemini-3.5-flashv1v129.840v5 · 32.3
opencode · Gemini 3.5 Flashgoogle/gemini-3.5-flashv1v229.340v9 · 31.5
hermes · Fusion routeropenrouter/fusionv1v128.834v3 · 29
claude-code · Opus 4.7claude-opus-4-7v1v328.736v4 · 29.7
opencode · Fusion routeropenrouter/fusionv1v228.336v1 · 29.3
hermes · Gemini 3.5 Flashopenrouter/google/gemini-3.5-flashv1v928.239.5v5 · 30
opencode · GPT-5.6 Terraopenai/gpt-5.6-terrav2v12838v5 · 29.8
hermes · GLM 5.1openrouter/z-ai/glm-5.1v1v227.336.5v3 · 28.8
claude-code · Sonnet 4.6claude-sonnet-4-6v1v526.237v10 · 28
claude-code · Sonnet 5anthropic/claude-sonnet-5v2v425.336v2 · 27.3
hermes · Sonnet 4.6openrouter/anthropic/claude-sonnet-4.6v1v52436v2 · 27

Evaluations computed

These runs used runbook v1, which wrote no validation report, so the scorecard is replayed here: the v2 runbook's code checks re-run on every stored SVG, plus two judge checks at 28/40 (the external judge, and the agent's own self-score, which is what the v2 runbook gates on). The overall row is the derived verdict shown on Runs.

EvaluationPassedRateShare
Code check · svg-parses98/98100%
Code check · svg-size-under-50kb98/98100%
Code check · no-image-tags98/98100%
Code check · png-rendered98/98100%
Judge · external judge ≥ 28/4052/9853%
Judge · agent self-score ≥ 28/40 (the v2 runbook's judge check)98/98100%
Overall · derived verdict (code checks + external judge)52/9853%

Findings generated

The top three are within 1.2 points, and two of them are v2 models

medium confidence

By the judge's score for each agent's own pick, opencode · Gemini 3.8 Flash and claude-code · Opus 5.5 tie at 32.8/40, then hermes · Jev Router (31.7). opencode · Gemini 3.8 Flash got there in a single round before two OpenRouter outages ended its lineage. Differences this small sit inside what one judge on one lineage can resolve.

v2-opencode-gemini38 v1 · 225e57cdv2-claude-opus55 v3 · 35c47f4ev2-hermes-jev v4 · 3ba61c00

Current models draw better pelicans than last spring's

high confidence

The v2 cohort (5 lineages) averages 30.1/40 from the judge against 27.9 for v1 (9 lineages). 3 of 5 v2 setups beat the best v1 setup, gemini-cli · Gemini 3.5 Flash at 29.8, and v2 did it in at most five rounds instead of ten.

v2 mean 30.1v1 mean 27.9gemini-cli v1 · f5cd6a31

The model matters more than the agent harness around it

medium confidence

The same model run through different agents lands within 2.3 points of itself on its best round by the judge: gemini-cli · Gemini 3.5 Flash 32.3; opencode · Gemini 3.5 Flash 31.5; hermes · Gemini 3.5 Flash 30. The spread between models in the same cohort is larger, so pick the model first.

gemini-cli v5opencode v9hermes-flash v5

Agents often keep the wrong drawing

high confidence

In 12 of 14 lineages the round the agent picked as its best is not the round the judge scored highest; the judge's favourite scored 1.6 points more on average. Worst case: hermes · Sonnet 4.6 kept v5 (24) over v2 (27). A ranking by the agents' picks understates what each setup can draw.

hermes-sonnet v5 vs v212/14 lineages

Proposed improvements generated

eval

Pick the kept round with the external judge, not the self-score

Would change the kept drawing in 12 of 14 lineages, +1.6 judge points each on average.

backed by: F4

sweep

Re-run the top three setups three times each

Tells whether the 1.2-point spread at the top is real; today each setup has one lineage.

backed by: F1, F3

eval

Add a second judge from a different vendor

The judge (Opus 5.5) is also a contestant; a second judge shows whether the lead survives a different yardstick.

backed by: F1

Caveats

  • One external judge (Claude Opus 5.5), and it is also a contestant: weigh the Opus 5.5 row accordingly.
  • One lineage per agent × model; no repeat runs, so there are no error bars on any setup.
  • v1 and v2 differ in round count (10 vs 5) and in the models' era, not only in the model.
  • opencode · Gemini 3.8 Flash has a single round: its lineage ended after two OpenRouter outages.
  • Scores are for the agent's own pick of its best round; the ring marker shows the judge's favourite round.
Runs and provenance

Computed on 2026-09-28 from results.json and judge-raw.json. Task pelican-bicycle-svg. Every run is listed on 02 · Runs.

LineageCohortCollectionRunsTrajectories
claude-code · Opus 5.5v2jettygrowthteam55361e30a ba932f3e 35c47f4e 2e046a89 0e41c0a0
claude-code · Sonnet 5v2jettygrowthteam5ea416761 b77ea720 6e98127a 240351f2 1925c665
hermes · Jev Routerv2jettygrowthteam51a125d51 7a63ffb0 ded99127 3ba61c00 ef67915d
opencode · Gemini 3.8 Flashv2jettygrowthteam1225e57cd
opencode · GPT-5.6 Terrav2jettygrowthteam5ccbdd583 699fefde ce858eaf 9fe62cbb de112455
claude-code · Opus 4.7v1jettyio10c3121188 cb2a7f37 d4a107dc 8ec2e90c 4160dcf1 4848f8a3 df18c780 20f46c4d 4897cbca 38634df1
claude-code · Sonnet 4.6v1jettyio117ddfcf08 414e644d b169371d 063dfed9 7ddc3c19 39febb11 ddf779bd 79fdee62 5f980543 2043d7e3 3a4c3357
gemini-cli · Gemini 3.5 Flashv1jettyio10f5cd6a31 c990e7c4 5d360721 4ed5de35 a652e4e3 eca90d64 f08684e3 ce2de30f 4e122677 24091968
hermes · Gemini 3.5 Flashv1jettyio10c118fb80 2128ffc1 2405db86 822a3bca 60502506 2c2a74e8 c165de92 d391e018 e6d4548b 3e21bf05
hermes · Fusion routerv1jettyio32d0c10c4 b38f9569 cb138790
hermes · GLM 5.1v1jettyio10c1f2da04 4b324d80 f443b217 7ce1468b ab8cb01d 9fd19804 83e8efe2 21b9b8c1 9f9e50be fff60669
hermes · Sonnet 4.6v1jettyio10697b57d4 acb2ba11 99a389e4 28f94566 8969a1a0 ebc95c4b fdf47c73 25323cb8 5b9cf6ae 34f10b3e
opencode · Gemini 3.5 Flashv1jettyio1067836d67 660d2a66 5cd01e9e b130c9c2 d94ff2e1 4a378be2 e3059c87 70898b5c b56c0c03 2a46de05
opencode · Fusion routerv1jettyio340df25cb bfa3612b 9776d557

Step through every drawing.

Ready to stop guessing if your outputs are good?

Get Started Book a 15-minute walkthrough
Connect your agent to Jetty →