03 — Investigations
Questions answered
from the runs.
Each investigation is a findings report in the shape Jetty uses: the question, a takeaway, computed tiles and a key figure, the evaluations scorecard, findings with their evidence, proposed improvements and caveats. Every number is computed from the 98 runs; the prose is written from those numbers and tagged generated.
Which agent and model draw the best pelican on a bicycle?
opencode · Gemini 3.8 Flash and claude-code · Opus 5.5 tie for the best pelican at 32.8/40 from the external judge, ahead of hermes · Jev Router at 31.7. Current models beat last spring's: the v2 cohort averages 30.1 against 27.9 for v1. The gaps are a few points on one judge with one lineage per setup, so rank the top three as a group and re-run before choosing between them.
Are the evals doing anything?Do coding agents grade their own drawings accurately?
No. Agents score their own drawings a median 9.6 points (of 40) above the external judge, and across 98 drawings the two scores correlate at only 0.28. As a gate, the self-score passes 98 of 98 runs at 28/40 while the judge passes 52, so the v2 runbook's self-score check cannot fail a weak drawing. Score with an external judge instead.
Did the change help?Does hill-climbing on self-scores actually improve the drawing?
Mostly not. Across 13 lineages with three or more rounds, self-scores rose 2.3 points from round 1 to their best, but the judge scored the last round 0.9 points lower than the first on average, and 8 of 13 ended below where they started. The climb steers by the self-score, which barely tracks the judge; steer it by the judge and stop when the judge stops improving.