02 — Runs · v2 · jettygrowthteam · auto-evolved runbook (weakest-axis targeting)
opencode
openai/gpt-5.6-terra
opencode driving openai/gpt-5.6-terra through OpenRouter, 5 hill-climb rounds from the v1 runbook.
Self-score vs judge, by axis
Best round v1, 0–10 per axis
Judge: The pelican reads well thanks to the big orange beak with a throat pouch, eye, and head crest, and the pleasant flat palette gives a clean cartoon look. However, the body is an egg blob with an odd notched outline, and the wing is a flat slab reaching an ambiguous handlebar; the frame geometry is muddled, the chain/crank connection is awkward, and the legs overlap the rear wheel clumsily.
Score by round
Out of 40. Each round starts from the previous round's best SVG. Dashed line: the 28/40 pass threshold.
5 drawings, oldest first.
Click a drawing to open it in the head-to-head viewer. The outlined card is the agent's best round. Verdicts are derived (code checks on the stored SVG and judge ≥ 28/40); see Runs.
Scores per round
| Round | Pelican | Bicycle | Composition | Polish | Self total | Judge total | Verdict | Judge spread | Cost | Minutes | Trajectory |
|---|---|---|---|---|---|---|---|---|---|---|---|
| v1 | 7.2self 10 | 6.5self 9 | 7.2self 10 | 7.2self 9 | 38 | 28 | pass | 3.5 | not metered | 1.4 | ccbdd583 |
| v2 | 7self 9 | 6.5self 9.5 | 7self 9.5 | 7self 9.2 | 37.2 | 27.5 | fail | 0 | not metered | 1.7 | 699fefde |
| v3 | 8self 9.5 | 6.5self 9.5 | 5.3self 9.5 | 7.3self 9.5 | 38 | 27.2 | fail | 2 | not metered | 1.5 | ce858eaf |
| v4 | 8self 9.5 | 6.3self 9.5 | 5.5self 9.5 | 7.5self 9 | 37.5 | 27.3 | fail | 1 | not metered | 1.4 | 9fe62cbb |
| v5 | 8self 10 | 6.8self 10 | 7.3self 9 | 7.7self 9 | 38 | 29.8 | pass | 1.5 | not metered | 1.5 | de112455 |