---
title: "Investigations | Pelican benchmark"
url: https://evaljetty.com/pelicans/investigations/
description: "Three questions answered from the pelican-on-a-bicycle runs: which setup draws best, whether agents grade themselves honestly, and whether hill-climbing helps."
updated: 2026-09-28
publisher: Jetty (https://jetty.io)
---

03 — Investigations

# Questions answered from the runs.

Each investigation is a findings report in the shape Jetty uses: the question, a takeaway, computed tiles and a key figure, the evaluations scorecard, findings with their evidence, proposed improvements and caveats. Every number is computed from the 98 runs; the prose is written from those numbers and tagged generated.

[Compare setups

### Which agent and model draw the best pelican on a bicycle?

opencode · Gemini 3.8 Flash and claude-code · Opus 5.5 tie for the best pelican at 32.8/40 from the external judge, ahead of hermes · Jev Router at 31.7. Current models beat last spring's: the v2 cohort averages 30.1 against 27.9 for v1. The gaps are a few points on one judge with one lineage per setup, so rank the top three as a group and re-run before choosing between them.](https://evaljetty.com/pelicans/investigations/best-pelican.html)[Are the evals doing anything?

### Do coding agents grade their own drawings accurately?

No. Agents score their own drawings a median 9.6 points (of 40) above the external judge, and across 98 drawings the two scores correlate at only 0.28. As a gate, the self-score passes 98 of 98 runs at 28/40 while the judge passes 52, so the v2 runbook's self-score check cannot fail a weak drawing. Score with an external judge instead.](https://evaljetty.com/pelicans/investigations/self-grading.html)[Did the change help?

### Does hill-climbing on self-scores actually improve the drawing?

Mostly not. Across 13 lineages with three or more rounds, self-scores rose 2.3 points from round 1 to their best, but the judge scored the last round 0.9 points lower than the first on average, and 8 of 13 ended below where they started. The climb steers by the self-score, which barely tracks the judge; steer it by the judge and stop when the judge stops improving.](https://evaljetty.com/pelicans/investigations/hill-climbing.html)

---
Source: https://evaljetty.com/pelicans/investigations/ · Evals by Jetty · Run your own evals: https://jetty.io/?utm_source=evaljetty&utm_medium=llms
