---
title: "Do coding agents grade their own drawings accurately? | Pelican benchmark"
url: https://evaljetty.com/pelicans/investigations/self-grading.html
description: "No. Agents score their own drawings a median 9.6 points (of 40) above the external judge, and across 98 drawings the two scores correlate at only 0.28. As a gate, the self-score passes 98 of 98 runs at 28/40 while the judge passes 52, so the v2 runbook's self-score check cannot fail a weak drawing."
updated: 2026-09-28
publisher: Jetty (https://jetty.io)
---

03 — Investigation · Are the evals doing anything?

# Do coding agents grade their own drawings accurately?

task pelican-bicycle-svg98 scored drawings14 lineagesjudge anthropic/claude-opus-5.5 · T=0 · 3 samplescomputed 2026-09-28

**Decision it informs:** Whether the runbook can trust the agent's own score as its judge check, or needs an external judge.

## Key takeaway generated

No. Agents score their own drawings a median 9.6 points (of 40) above the external judge, and across 98 drawings the two scores correlate at only 0.28. As a gate, the self-score passes 98 of 98 runs at 28/40 while the judge passes 52, so the v2 runbook's self-score check cannot fail a weak drawing. Score with an external judge instead.

## At a glance computed

Runs

98

14 agent × model lineages

Spend

$11.90

10 metered runs + $2.77 judge; 11 runs not metered

Time

5.5 min

median per run (32 timed)

Pass rate

53%

52/98 derived verdicts pass

## Key figure computed

### Every agent scores itself above the judge, by a median of 9.6 points

One row per agent × model, best round as picked by the agent. Filled dot: the external judge's total out of 40; ring: the agent's own. The line between them is the inflation. Click a row to open that lineage's runs.

external judgeagent self-score

Table view

| Agent · model | Cohort | Round | Self /40 | Judge /40 | Inflation | Judge spread |
| --- | --- | --- | --- | --- | --- | --- |
| [opencode · Gemini 3.8 Flash](https://evaljetty.com/pelicans/runs/agents/v2-opencode-gemini38.html) | v2 | v1 | 38.5 | 32.8 | +5.7 | 0.5 |
| [claude-code · Opus 5.5](https://evaljetty.com/pelicans/runs/agents/v2-claude-opus55.html) | v2 | v3 | 36 | 32.8 | +3.2 | 0.5 |
| [hermes · Jev Router](https://evaljetty.com/pelicans/runs/agents/v2-hermes-jev.html) | v2 | v4 | 37 | 31.7 | +5.3 | 0.5 |
| [gemini-cli · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/gemini-cli.html) | v1 | v1 | 40 | 29.8 | +10.2 | 0.5 |
| [opencode · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/opencode.html) | v1 | v2 | 40 | 29.3 | +10.7 | 1 |
| [hermes · Fusion router](https://evaljetty.com/pelicans/runs/agents/hermes-fusion.html) | v1 | v1 | 34 | 28.8 | +5.2 | 2 |
| [claude-code · Opus 4.7](https://evaljetty.com/pelicans/runs/agents/claude-opus.html) | v1 | v3 | 36 | 28.7 | +7.3 | 1.5 |
| [opencode · Fusion router](https://evaljetty.com/pelicans/runs/agents/opencode-fusion.html) | v1 | v2 | 36 | 28.3 | +7.7 | 1.5 |
| [hermes · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/hermes-flash.html) | v1 | v9 | 39.5 | 28.2 | +11.3 | 2 |
| [opencode · GPT-5.6 Terra](https://evaljetty.com/pelicans/runs/agents/v2-opencode-gpt56.html) | v2 | v1 | 38 | 28 | +10 | 3.5 |
| [hermes · GLM 5.1](https://evaljetty.com/pelicans/runs/agents/hermes-glm51.html) | v1 | v2 | 36.5 | 27.3 | +9.2 | 0.5 |
| [claude-code · Sonnet 4.6](https://evaljetty.com/pelicans/runs/agents/claude-sonnet.html) | v1 | v5 | 37 | 26.2 | +10.8 | 1 |
| [claude-code · Sonnet 5](https://evaljetty.com/pelicans/runs/agents/v2-claude-sonnet5.html) | v2 | v4 | 36 | 25.3 | +10.7 | 1 |
| [hermes · Sonnet 4.6](https://evaljetty.com/pelicans/runs/agents/hermes-sonnet.html) | v1 | v5 | 36 | 24 | +12 | 1 |

## Evaluations computed

These runs used runbook v1, which wrote no validation report, so the scorecard is replayed here: the v2 runbook's code checks re-run on every stored SVG, plus two judge checks at 28/40 (the external judge, and the agent's own self-score, which is what the v2 runbook gates on). The overall row is the derived verdict shown on Runs.

| Evaluation | Passed | Rate | Share |
| --- | --- | --- | --- |
| Code check · svg-parses | 98/98 | 100% |  |
| Code check · svg-size-under-50kb | 98/98 | 100% |  |
| Code check · no-image-tags | 98/98 | 100% |  |
| Code check · png-rendered | 98/98 | 100% |  |
| Judge · external judge ≥ 28/40 | 52/98 | 53% |  |
| Judge · agent self-score ≥ 28/40 (the v2 runbook's judge check) | 98/98 | 100% |  |
| Overall · derived verdict (code checks + external judge) | 52/98 | 53% |  |

## Findings generated

### Every agent flatters itself, by a median of 9.6 points

high confidence

All 14 of 14 lineages gave their best round a higher total than the judge did, by 3.2 to 12 points. Across every round the mean gap is 8.2.

hermes-sonnet +12v2-claude-opus55 +3.298 scored drawings

### A self-score barely predicts the judge's

high confidence

Across all 98 drawings the correlation between self-score and judge score is 0.28; ranking the 14 lineages by self-score and by judge score gives a rank correlation of 0.20. The self-score cannot rank agents against each other.

r = 0.28 over 98 drawingsrank r = 0.20

### As a pass/fail gate, the self-score never fails a run

high confidence

The v2 runbook's judge check (self-score ≥ 28/40) would pass 98 of 98 stored runs; the external judge passes 52. That is 46 runs the self-score would wave through. No self threshold fixes it: the best one (38/40) agrees with the judge on only 66 of 98 runs.

self ≥ 28: 98/98judge ≥ 28: 52/98best threshold 38: 66/98

### The code checks guard the format, not the drawing

high confidence

All 98 of 98 stored SVGs pass every replayable code check (parses, under 50 KB, no <image>, renders). They are worth keeping as cheap guards, but they never split runs, so today the only eval that separates a good pelican from a weak one is the external judge.

svg-parses: 98/98svg-size-under-50kb: 98/98no-image-tags: 98/98

### Bicycle is where agents are most generous

medium confidence

Averaged over every round, self-scores exceed the judge by 2.2 on Bicycle and 2 on Composition. The v1 Gemini Flash lineages gave 7 of their 30 drawings a perfect 40/40.

pelican: +2bicycle: +2.2composition: +2polish: +2

## Proposed improvements generated

eval

### Replace the self-score judge check with an external judge

Would have failed 46 of the 98 stored runs that the self-score passes, matching the judge's 52/98.

backed by: F1, F3

runbook

### Add a checklist item scored from the PNG by a second model

A cheap vision pass on 'is the pelican riding?' targets the most inflated axis without a full judge call.

backed by: F5

eval

### Keep the code checks, but don't read their pass rate as quality

They passed 98/98; report them as guards and gate quality on the judge.

backed by: F4

## Caveats

- The judge is one model (Claude Opus 5.5) at temperature 0 with three samples; its own bias is unmeasured.
- The judge is also a contestant (v2 Opus 5.5 lineage).
- Self-scores come from each agent's report.md; a few v1 lineages had no image input and scored from SVG source.
- The runs used runbook v1, whose self-score was advisory; agents may score differently when the score gates their run.

**Runs and provenance**

Computed on 2026-09-28 from [results.json](https://evaljetty.com/pelicans/data/results.json) and [judge-raw.json](https://evaljetty.com/pelicans/data/judge-raw.json). Task pelican-bicycle-svg. Every run is listed on [02 · Runs](https://evaljetty.com/pelicans/runs/index.html#all-runs).

| Lineage | Cohort | Collection | Runs | Trajectories |
| --- | --- | --- | --- | --- |
| [claude-code · Opus 5.5](https://evaljetty.com/pelicans/runs/agents/v2-claude-opus55.html) | v2 | jettygrowthteam | 5 | 5361e30a ba932f3e 35c47f4e 2e046a89 0e41c0a0 |
| [claude-code · Sonnet 5](https://evaljetty.com/pelicans/runs/agents/v2-claude-sonnet5.html) | v2 | jettygrowthteam | 5 | ea416761 b77ea720 6e98127a 240351f2 1925c665 |
| [hermes · Jev Router](https://evaljetty.com/pelicans/runs/agents/v2-hermes-jev.html) | v2 | jettygrowthteam | 5 | 1a125d51 7a63ffb0 ded99127 3ba61c00 ef67915d |
| [opencode · Gemini 3.8 Flash](https://evaljetty.com/pelicans/runs/agents/v2-opencode-gemini38.html) | v2 | jettygrowthteam | 1 | 225e57cd |
| [opencode · GPT-5.6 Terra](https://evaljetty.com/pelicans/runs/agents/v2-opencode-gpt56.html) | v2 | jettygrowthteam | 5 | ccbdd583 699fefde ce858eaf 9fe62cbb de112455 |
| [claude-code · Opus 4.7](https://evaljetty.com/pelicans/runs/agents/claude-opus.html) | v1 | jettyio | 10 | c3121188 cb2a7f37 d4a107dc 8ec2e90c 4160dcf1 4848f8a3 df18c780 20f46c4d 4897cbca 38634df1 |
| [claude-code · Sonnet 4.6](https://evaljetty.com/pelicans/runs/agents/claude-sonnet.html) | v1 | jettyio | 11 | 7ddfcf08 414e644d b169371d 063dfed9 7ddc3c19 39febb11 ddf779bd 79fdee62 5f980543 2043d7e3 3a4c3357 |
| [gemini-cli · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/gemini-cli.html) | v1 | jettyio | 10 | f5cd6a31 c990e7c4 5d360721 4ed5de35 a652e4e3 eca90d64 f08684e3 ce2de30f 4e122677 24091968 |
| [hermes · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/hermes-flash.html) | v1 | jettyio | 10 | c118fb80 2128ffc1 2405db86 822a3bca 60502506 2c2a74e8 c165de92 d391e018 e6d4548b 3e21bf05 |
| [hermes · Fusion router](https://evaljetty.com/pelicans/runs/agents/hermes-fusion.html) | v1 | jettyio | 3 | 2d0c10c4 b38f9569 cb138790 |
| [hermes · GLM 5.1](https://evaljetty.com/pelicans/runs/agents/hermes-glm51.html) | v1 | jettyio | 10 | c1f2da04 4b324d80 f443b217 7ce1468b ab8cb01d 9fd19804 83e8efe2 21b9b8c1 9f9e50be fff60669 |
| [hermes · Sonnet 4.6](https://evaljetty.com/pelicans/runs/agents/hermes-sonnet.html) | v1 | jettyio | 10 | 697b57d4 acb2ba11 99a389e4 28f94566 8969a1a0 ebc95c4b fdf47c73 25323cb8 5b9cf6ae 34f10b3e |
| [opencode · Gemini 3.5 Flash](https://evaljetty.com/pelicans/runs/agents/opencode.html) | v1 | jettyio | 10 | 67836d67 660d2a66 5cd01e9e b130c9c2 d94ff2e1 4a378be2 e3059c87 70898b5c b56c0c03 2a46de05 |
| [opencode · Fusion router](https://evaljetty.com/pelicans/runs/agents/opencode-fusion.html) | v1 | jettyio | 3 | 40df25cb bfa3612b 9776d557 |

---
Source: https://evaljetty.com/pelicans/investigations/self-grading.html · Evals by Jetty · Run your own evals: https://jetty.io/?utm_source=evaljetty&utm_medium=llms
