---
title: "Author: the pelican runbook | Pelican benchmark"
url: https://evaljetty.com/pelicans/author.html
description: "The pelican-on-a-bicycle runbook (evals v2: code checks, checklist, validation report), its four-axis rubric, how the hill climb rewrites it, the external vision judge, and how to run it on your own Jetty."
updated: 2026-09-28
publisher: Jetty (https://jetty.io)
---

01 — Author

# Tell Jetty the job to be done.

Everything the agents did came from one Markdown runbook: the objective, the constraints, a four-axis rubric and a self-critique loop. Here is that runbook, how the hill climb rewrote it each round, the external judge we added on top, and how to run it on your own Jetty.

[Download the runbook (v2.0.0)](https://evaljetty.com/pelicans/run/pelican-bicycle-svg.runbook.md)
[Task workflow JSON](https://evaljetty.com/pelicans/run/task-workflow.json)
[Run it on your Jetty](https://evaljetty.com/pelicans/author.html#run-it)

## A pelican, *riding a bicycle, in SVG.*

Hand-write a pure-XML SVG of a pelican riding a bicycle. Both subjects must be unmistakable, and the pelican has to be *riding* the bike (on the seat, feet at the pedals, wing tips on the bars) rather than floating next to it. The agent draws, renders the SVG to PNG with rsvg-convert, looks at the PNG, scores it and redraws, then keeps its best round.

### SVG constraints

- Pure SVG elements only: no <image> tags, no base64 or external rasters
- Well-formed XML that renders in any modern browser
- viewBox about 0 0 800 600, file size under 50 KB

### The rubric

| Axis | Question | Scale |
| --- | --- | --- |
| **Pelican** | Would a stranger immediately say “that's a pelican”? | 0–10 |
| **Bicycle** | Two spoked wheels, a real frame, handlebars, seat, pedals? | 0–10 |
| **Composition** | Is the pelican clearly riding the bike? | 0–10 |
| **Polish** | Line quality, colour, balance | 0–10 |

Total out of 40. The agent scores itself with this rubric; the external judge uses the same four axes.

## Version 2.0.0, *in the evals-v2 shape.*

### What v2 evaluates on every run

| Kind | Id | What it checks |
| --- | --- | --- |
| code check | outputs-exist | Every required output file exists and is non-empty |
| code check | svg-parses | final.svg is well-formed XML with an <svg> root |
| code check | svg-size-under-50kb | final.svg is under 50 KB |
| code check | no-image-tags | final.svg has no <image> elements or embedded rasters |
| code check | png-rendered | final.png is a real PNG at least 400 px wide |
| code check | rounds-recorded | Every scored round has its SVG and PNG |
| checklist | pelican-recognizable | the bird reads as a pelican at a glance (long beak with a visible throat pouch) |
| checklist | bicycle-complete | two spoked wheels, a connected frame, handlebars, seat and pedals are all present |
| checklist | actually-riding | the pelican sits on the seat with wing tips on the handlebars and feet at the pedals |
| checklist | scores-match-report | report.md, scores.json and final.svg agree on the per-round scores and the best round |
| judge | self-score | The agent's own 4-axis score of its best round, out of 40; passes at ≥ {{pass\_threshold}} (default 28) |

current · v2.0.0 · evals v2
[download .md](https://evaljetty.com/pelicans/run/pelican-bicycle-svg.runbook.md)

```
---
name: pelican-bicycle-svg
version: "2.0.0"
evals_version: 2
description: Hand-write the best possible SVG of a pelican riding a bicycle, refining it over several self-critique rounds, and write an evals-v2 validation report.
agent: claude-code
model: anthropic/claude-sonnet-5
model_provider: openrouter
snapshot: python312-uv
evaluation: programmatic
primary_outputs:
  - final.svg
  - final.png
  - report.md
secrets:
  OPENROUTER_API_KEY:
    env: OPENROUTER_API_KEY
    description: "OpenRouter key (collection environment variable) used by the agent's model calls"
    required: true
---
```

## Pelican Riding a Bicycle: SVG Runbook

### Objective

Produce the highest-quality **hand-written SVG of a pelican riding a bicycle** and deliver it, a PNG render, the per-round scores and an evals-v2 validation report to `{{results_dir}}`.

Both subjects must be unmistakable, and the composition must be coherent: the pelican is **riding** the bicycle (on the seat, feet at the pedals, wing tips on the handlebars), not floating next to it. You draw, render the SVG to PNG, look at the PNG, score it on a four-axis rubric and redraw, for up to `{{max_rounds}}` rounds, then keep the best round.

Hard constraints on every SVG you write:

- Pure SVG XML: `<path>`, `<circle>`, `<ellipse>`, `<rect>`, `<polygon>`, `<line>`, `<g>`, `<defs>`, gradients and so on
- **No** `<image>` elements and no embedded rasters (`data:image/...`, base64, external bitmaps)
- Well-formed XML that renders in any modern browser
- viewBox approximately `0 0 800 600` (landscape)
- File size under 50 KB

### REQUIRED OUTPUT FILES

- `{{results_dir}}/final.svg`: the best round's SVG
- `{{results_dir}}/final.png`: `final.svg` rendered with `rsvg-convert -w 800`
- `{{results_dir}}/rounds/v<N>.svg` and `rounds/v<N>.png`: every round you attempted
- `{{results_dir}}/scores.json`: your rubric scores per round and the best round (format in step 4)
- `{{results_dir}}/report.md`: per-round score table and notes (format in step 4)
- `{{results_dir}}/validation_report.json`: evals-v2 report, written in the last section

### Parameters

| Parameter | Template Variable | Default | Description |
| --- | --- | --- | --- |
| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |
| Rounds | `{{max_rounds}}` | `3` | Maximum draw, render, critique rounds (the loop exits early at 36/40) |
| Pass threshold | `{{pass_threshold}}` | `28` | Minimum self-score total (out of 40) for the `self-score` judge check to pass |

### Dependencies

- `python312-uv` snapshot (Python 3.12 for the checks and the report)
- An SVG rasterizer: `rsvg-convert` from `librsvg2-bin` (installed in step 1 if missing)
- `OPENROUTER_API_KEY` collection environment variable for the agent's own model calls

### Steps

#### 1. Setup

```
mkdir -p {{results_dir}}/rounds && cd {{results_dir}}
which rsvg-convert || (apt-get update && apt-get install -y librsvg2-bin) || true
which rsvg-convert
```

#### 2. Round 1: first draft

Write your best first attempt to `{{results_dir}}/rounds/v1.svg`. Aim to depict:

- **Pelican**: long beak with throat pouch (the iconic silhouette), body, eye, wings, legs and feet
- **Bicycle**: two wheels with spokes, frame (top tube, down tube, seat tube), handlebars, seat, pedals
- **Riding**: body on the seat, feet on or near the pedals, wing tips on the handlebars

Render it:

```
cd {{results_dir}} && rsvg-convert -w 800 rounds/v1.svg -o rounds/v1.png
```

#### 3. Critique and refine

For each round R from 1 to `{{max_rounds}}`:

1. **Look** at `rounds/v${R}.png` (open it as an image; do not score from the SVG source).
2. **Score** it on four axes, 0 to 10 each:  
   **Pelican**: would a stranger immediately say "that's a pelican"?  
   **Bicycle**: would they say "that's a bicycle"?  
   **Composition**: is the pelican clearly riding the bike?  
   **Polish**: line quality, colour, balance
3. If the total is 36/40 or more, or R equals `{{max_rounds}}`, stop the loop.
4. Otherwise name the lowest-scoring axis, write down two or three concrete fixes, write `rounds/v$((R+1)).svg` applying them (keep what worked) and render `rounds/v$((R+1)).png`.

Score honestly: the benchmark re-scores every drawing with an external vision judge, and agents typically rate themselves several points above it. Score what the PNG shows, not what you intended to draw.

#### 4. Finalize

Pick the round with the highest total (break ties in favour of the later round) and copy it:

```
cd {{results_dir}} && cp rounds/vBEST.svg final.svg && cp rounds/vBEST.png final.png
```

Write `{{results_dir}}/scores.json`:

```
{"best_round": 2, "rounds": [
  {"round": 1, "pelican": 7, "bicycle": 8, "composition": 6, "polish": 7, "total": 28},
  {"round": 2, "pelican": 8, "bicycle": 8, "composition": 7, "polish": 8, "total": 31}]}
```

Write `{{results_dir}}/report.md` with a `## Per-round scores` table (Round, Pelican, Bicycle, Composition, Polish, Total), the best round, and short **What worked**, **What didn't work** and **SVG technique notes** sections (viewBox, shape count, file size).

### Code Checks

Each check is a deterministic command; a failing check fails the run. The report script in the last section runs the same checks, so run these for your own verification first.

#### outputs-exist — Every required output file exists and is non-empty

```
cd {{results_dir}} && for f in final.svg final.png scores.json report.md; do test -s "$f" || { echo "missing $f"; exit 1; }; done
```

#### svg-parses — final.svg is well-formed XML with an <svg> root

```
python3 -c "import xml.etree.ElementTree as E; r=E.parse('{{results_dir}}/final.svg').getroot(); assert r.tag.split('}')[-1]=='svg', r.tag"
```

#### svg-size-under-50kb — final.svg is under 50 KB

```
test "$(wc -c < {{results_dir}}/final.svg)" -lt 51200
```

#### no-image-tags — final.svg has no <image> elements or embedded rasters

```
python3 -c "import xml.etree.ElementTree as E; s=open('{{results_dir}}/final.svg').read(); r=E.fromstring(s); assert not [e for e in r.iter() if e.tag.split('}')[-1]=='image'] and 'data:image' not in s and 'base64' not in s"
```

#### png-rendered — final.png is a real PNG at least 400 px wide

```
python3 -c "import struct; b=open('{{results_dir}}/final.png','rb').read(24); assert b[:8]==b'\x89PNG\r\n\x1a\n', 'not a PNG'; w=struct.unpack('>I', b[16:20])[0]; assert w>=400, w"
```

#### rounds-recorded — Every scored round has its SVG and PNG

```
python3 -c "import json,os; d=json.load(open('{{results_dir}}/scores.json')); ns=[r['round'] for r in d['rounds']]; assert ns and d['best_round'] in ns; miss=[n for n in ns for e in ('svg','png') if not os.path.getsize('{{results_dir}}/rounds/v%d.%s'%(n,e))]; assert not miss, miss"
```

### Checklist

Confirm each item by looking at `final.png`; a failed item fails the run.

- `pelican-recognizable`: the bird reads as a pelican at a glance (long beak with a visible throat pouch)
- `bicycle-complete`: two spoked wheels, a connected frame, handlebars, seat and pedals are all present
- `actually-riding`: the pelican sits on the seat with wing tips on the handlebars and feet at the pedals
- `scores-match-report`: report.md, scores.json and final.svg agree on the per-round scores and the best round

### Write Validation Report

Write `{{results_dir}}/validation_report.json` in the evals-v2 shape by running the script below. It runs every code check above, records one `step` entry per step, one `checklist` entry per item you pass in (status, a colon, then a one-line reason) and your best-round self-score as a `judge` check (score out of 40, threshold `{{pass_threshold}}`). Jetty computes the verdict from these checks. Use `fail: <reason>` for anything that did not hold, and set `ITERATIONS` to the number of rounds you drew.

```
cd {{results_dir}} && ITERATIONS=3 \
STEPS='{"setup": "pass", "draft": "pass", "refine": "pass", "finalize": "pass"}' \
CHECKLIST='{"pelican-recognizable": "pass: <why>", "bicycle-complete": "pass: <why>", "actually-riding": "pass: <why>", "scores-match-report": "pass: <why>"}' \
python3 - <<'PY'
import json, os, struct, xml.etree.ElementTree as E
R = "{{results_dir}}"; T = "{{pass_threshold}}"; T = float(T) if T.replace(".", "", 1).isdigit() else 28.0
checks = []
def add(kind, cid, name, ok, msg, **kw):
    checks.append({"kind": kind, "id": cid, "name": name, "status": "pass" if ok else "fail", "message": str(msg)[:500], **kw})
def run(cid, name, fn):
    try:
        add("code_check", cid, name, True, fn() or "ok")
    except Exception as e:
        add("code_check", cid, name, False, repr(e))
p = lambda f: os.path.join(R, f)
def outputs():
    miss = [f for f in ("final.svg", "final.png", "scores.json", "report.md") if not os.path.exists(p(f)) or not os.path.getsize(p(f))]
    assert not miss, "missing: " + ", ".join(miss)
def parses():
    r = E.parse(p("final.svg")).getroot(); assert r.tag.split("}")[-1] == "svg", r.tag
def size():
    n = os.path.getsize(p("final.svg")); assert n < 51200, n; return f"{n} bytes"
def noimg():
    s = open(p("final.svg")).read(); r = E.fromstring(s)
    assert not [e for e in r.iter() if e.tag.split("}")[-1] == "image"] and "data:image" not in s and "base64" not in s
def png():
    b = open(p("final.png"), "rb").read(24); assert b[:8] == b"\x89PNG\r\n\x1a\n", "not a PNG"
    w = struct.unpack(">I", b[16:20])[0]; assert w >= 400, w; return f"{w} px wide"
def rounds():
    d = json.load(open(p("scores.json"))); ns = [r["round"] for r in d["rounds"]]
    assert ns and d["best_round"] in ns, "best_round not among rounds"
    miss = [f"v{n}.{e}" for n in ns for e in ("svg", "png") if not os.path.exists(p(f"rounds/v{n}.{e}"))]
    assert not miss, miss; return f"{len(ns)} rounds"
run("outputs-exist", "Every required output file exists and is non-empty", outputs)
run("svg-parses", "final.svg is well-formed XML with an <svg> root", parses)
run("svg-size-under-50kb", "final.svg is under 50 KB", size)
run("no-image-tags", "final.svg has no <image> elements or embedded rasters", noimg)
run("png-rendered", "final.png is a real PNG at least 400 px wide", png)
run("rounds-recorded", "Every scored round has its SVG and PNG", rounds)
for sid, st in json.loads(os.environ["STEPS"]).items():
    add("step", sid, sid, st.startswith("pass"), st)
for cid, st in json.loads(os.environ["CHECKLIST"]).items():
    ok = st.strip().lower().startswith("pass"); add("checklist", cid, cid, ok, st.split(":", 1)[-1].strip())
try:
    d = json.load(open(p("scores.json"))); best = next(r for r in d["rounds"] if r["round"] == d["best_round"])
    add("judge", "self-score", "Agent self-score of the best round (4 axes x 0-10)", best["total"] >= T,
        f"round {d['best_round']}: " + ", ".join(f"{k} {best[k]}" for k in ("pelican", "bicycle", "composition", "polish")),
        score=float(best["total"]), max_score=40.0, threshold=T)
except Exception as e:
    checks.append({"kind": "judge", "id": "self-score", "name": "Agent self-score of the best round", "status": "error", "message": repr(e)})
gating = [c for c in checks if c["kind"] != "step"]
passed = all(c["status"] == "pass" for c in gating)
rep = {"version": 2, "verdict": "pass" if passed else "fail", "overall_passed": passed,
       "iterations": int(os.environ.get("ITERATIONS", "1")), "checks": checks}
json.dump(rep, open(p("validation_report.json"), "w"), indent=2)
for c in checks: print(f"{c['status']:5} {c['kind']:10} {c['id']}: {c['message']}")
print("verdict:", rep["verdict"])
PY
```

Do not hand-edit the file and do not invent a different format: it must stay `{"version": 2, "verdict", "overall_passed", "iterations", "checks": [...]}`, with each check carrying `kind` (`step`, `code_check`, `checklist` or `judge`), `id`, `name`, `status` (`pass`, `fail`, `skipped` or `error`) and `message`, plus `score`, `max_score` and `threshold` on the judge.

### Tips

- One round costs roughly $0.50 to $2.50 depending on the model; most of it is the agent's own reasoning.
- Seeding round 1 with a previous run's `final.svg` (paste it into the runbook) is the single biggest improvement we measured in hand-edited climbs: it skips the cold start.
- Coordinate-precise fixes ("close the 13 px gap between the right wing tip and the grip") move scores more than scene-level additions (sun, motion lines, scenery), which tend to add clutter.
- The self-score is a weak signal across models: at evaljetty.com/pelicans it correlates only loosely with an external judge. Treat it as a steering aid for your own climb, not as a ranking.

### Changelog

- 2.0.0: Evals v2 shape (Parameters table, Code Checks, Checklist, Write Validation Report with the self-score as a judge check at `{{pass_threshold}}`/40); `scores.json` output; honest-scoring note.
- 1.x: The runbook used for every run published at evaljetty.com/pelicans (v1 cohort May 2026, v2 cohort Sep 2026): draw, render, self-score and redraw for `max_rounds`, then `final.svg`, `final.png` and `report.md`.

Runbook v1: the version every published run used (download: [.md](https://evaljetty.com/pelicans/run/pelican-bicycle-svg.v1.runbook.md))

```
---
name: pelican-bicycle-svg
version: 2
description: Generate the best possible hand-written SVG depicting a pelican riding a bicycle, with iterative self-refinement.
agent: claude-code
model: anthropic/claude-sonnet-5
model_provider: openrouter
parameters:
  max_rounds:
    type: integer
    default: 3
    description: Number of refinement rounds (each round = critique previous + write improved)
  results_dir:
    type: string
    default: /app/results
secrets:
  openrouter:
    env: OPENROUTER_API_KEY
    description: OpenRouter key (collection environment variable) used by the agent's model calls
evaluation:
  pattern: self-judge
---
```

## Pelican Riding a Bicycle — SVG Runbook

### Mission

Produce the highest-quality hand-written SVG depicting **a pelican riding a bicycle**. Both subjects must be unmistakably recognizable, the composition coherent (pelican IS interacting with the bicycle, not floating next to it), and the file should be pure XML SVG (no `<image>` tags, no base64 data, no external rasters).

### Hard constraints

- Pure SVG XML — use `<path>`, `<circle>`, `<ellipse>`, `<rect>`, `<polygon>`, `<line>`, `<g>`, `<defs>`, `<linearGradient>`, etc.
- **No** `<image>` elements with external or base64 data
- Must validate as well-formed XML and render in any modern browser
- viewBox approximately `0 0 800 600` (landscape) — adjust if you have a specific reason
- Total file size under 50KB

### INPUTS

- `max_rounds` = **3** (override only if explicitly told otherwise)
- `results_dir` = `/app/results`

### Steps

#### 1. Setup

```
mkdir -p {{results_dir}}/rounds
cd {{results_dir}}
```

Check that `rsvg-convert` or an SVG rasterizer is available:

```
which rsvg-convert || which inkscape || which convert
```

If none are available, install one:

```
apt-get update && apt-get install -y librsvg2-bin || true
```

#### 2. Round 1 — first draft

Write your best first attempt to `{{results_dir}}/rounds/v1.svg`. Aim to depict:

- **Pelican**: long beak with throat pouch (the iconic pelican silhouette), body, eye, wings, legs/feet
- **Bicycle**: two wheels with spokes, frame (top tube + down tube + seat tube), handlebars, seat, pedals
- **Riding**: pelican's body is on the seat, feet on or near pedals, "hands" (wing tips) on handlebars

Render it:

```
cd {{results_dir}}
rsvg-convert -w 800 rounds/v1.svg -o rounds/v1.png
```

#### 3. Self-critique + refinement loop

For each round R from 1 to `max_rounds - 1`:

1. **Inspect** `rounds/v${R}.png` visually (read it as an image).
2. **Score** on four axes, 0–10 each:  
   **Pelican recognizability** — would a stranger immediately say "that's a pelican"?  
   **Bicycle recognizability** — would they say "that's a bicycle"?  
   **Composition / riding** — is the pelican clearly riding the bike?  
   **Aesthetic polish** — line quality, color, balance
3. **Identify the lowest-scoring axis** and write down 2-3 concrete fixes.
4. **Write** `rounds/v$((R+1)).svg` applying those fixes. Keep what worked; rewrite what didn't.
5. **Render** `rounds/v$((R+1)).png`.
6. If the total (sum of 4 axes) is ≥ 36/40 — early-exit the loop.

#### 4. Finalize

After the loop, pick the round with the highest total score (break ties by preferring later rounds since they had more iteration).

```
cp {{results_dir}}/rounds/vBEST.svg {{results_dir}}/final.svg
cp {{results_dir}}/rounds/vBEST.png {{results_dir}}/final.png
```

Write `{{results_dir}}/report.md` containing:

```
# Pelican-Bicycle SVG Report

## Per-round scores
| Round | Pelican | Bicycle | Composition | Polish | Total |
|-------|---------|---------|-------------|--------|-------|
| 1 | ? | ? | ? | ? | ? |
| 2 | ? | ? | ? | ? | ? |
| 3 | ? | ? | ? | ? | ? |

## Best round: vN (Total: X/40)

## What worked
- ...

## What didn't work
- ...

## SVG technique notes
- viewBox used:
- Total path / shape count:
- File size:
```

### Final checklist

- `{{results_dir}}/final.svg` exists, is valid XML, renders in a browser, has no `<image>` tags
- `{{results_dir}}/final.png` exists (rasterized version)
- `{{results_dir}}/rounds/v*.svg` and `v*.png` exist for every round attempted
- `{{results_dir}}/report.md` exists with per-round score table and a notes section
- File size of `final.svg` is under 50KB

## How the hill climb *rewrites the runbook.*

Each hill-climb round is one Jetty run. Between rounds a small orchestrator ([hill\_climb.py](https://evaljetty.com/pelicans/run/hill_climb.py)) rewrites the runbook: it embeds the previous round's best SVG as the round-1 seed, reads the per-round scores from report.md, and rewrites the description to target the lowest self-scored axis. v1 ran 10 rounds per agent (3 for the Fusion sweeps); v2 ran 5.

One lineage was edited by hand instead: Jon read each claude-code + Sonnet 4.6 trajectory and edited the runbook himself, eleven times. Three of those edits:

Big leap · v2 → v3 · self 32 → 34 · judge 25.2 → 24.5

Embed the baseline

![v2](https://evaljetty.com/pelicans/img/claude-sonnet/v2.svg)

v2 · before

→
![v3](https://evaljetty.com/pelicans/img/claude-sonnet/v3.svg)

v3 · after

**The edit:** Use the v2 SVG verbatim as the round-1 input. Don't start from scratch.

**What happened:** Round 1 jumped from 28 to 33 self-score points just by skipping the cold start. The single highest-leverage edit in the sequence.

**Lesson:** When a runbook can carry a working artifact as a seed, it should.

Peak · v4 → v5 · self 36 → 37 · judge 24.5 → 26.2

Coordinate-precise asks

![v4](https://evaljetty.com/pelicans/img/claude-sonnet/v4.svg)

v4 · before

→
![v5](https://evaljetty.com/pelicans/img/claude-sonnet/v5.svg)

v5 · after

**The edit:** Close the 13 px right-wing-to-grip gap. Drop the left-wing tips to y ≈ 230.

**What happened:** Composition hit 10/10 on the self-score for the first time. Sonnet's peak self-score, never beaten later.

**Lesson:** At the top of the curve, the asks become measurements.

Cautionary · v5 → v6 · self 37 → 36 · judge 26.2 → 26.5

Scope creep cost a point

![v5](https://evaljetty.com/pelicans/img/claude-sonnet/v5.svg)

v5 · before

→
![v6](https://evaljetty.com/pelicans/img/claude-sonnet/v6.svg)

v6 · after

**The edit:** Add motion lines, wind tufts and a sun. Extend the right-wing covert lines to the wrist.

**What happened:** A richer scene, but composition slipped back to 9. The new elements added clutter the rubric couldn't reward.

**Lesson:** More isn't better when the rubric doesn't measure scene density.

[Runbook diffs

### Diff any two runbook versions in any lineage

From the hand edits to the auto-evolved rounds, side by side or unified, with or without the embedded baseline SVG.](https://evaljetty.com/pelicans/author/versions.html)[Investigation

### Did the climb help?

Does hill-climbing on self-scores actually improve the drawing? The judge's answer, round by round.](https://evaljetty.com/pelicans/investigations/hill-climbing.html)

## One fixed yardstick *for every drawing.*

Judge model

Opus 5.5

anthropic/claude-opus-5.5 via OpenRouter

Sampling

3 × T=0

independent calls, averaged

Renderer

librsvg

rsvg-convert -w 800, white background

Judge spend

$2.77

all 98 drawings

The system prompt, rubric and model are fixed across all drawings. Raw responses, including token usage, are published in [judge-raw.json](https://evaljetty.com/pelicans/data/judge-raw.json). A single LLM judge has biases of its own: it is one fixed yardstick, not ground truth, and the judge model is also one of the v2 contestants.

```
The image below is a rasterized SVG that a coding agent hand-wrote for the prompt:
"Generate an SVG of a pelican riding a bicycle."

Score it on four axes, each an integer or half-integer from 0 to 10:

1. pelican - Pelican recognizability. Would a stranger immediately say "that's a pelican"? Look for the long beak with a throat pouch, pelican body shape, eye, wings, legs/feet. 10 = unmistakably a pelican; 5 = a generic bird; 0 = no bird.
2. bicycle - Bicycle recognizability and correctness. Two wheels (with spokes), a plausible frame (top tube, down tube, seat tube), handlebars, seat, pedals/crank. 10 = a correct, well-formed bicycle; 5 = recognizable but structurally wrong; 0 = no bicycle.
3. composition - Is the pelican clearly RIDING the bicycle? Body on the seat, feet on or near the pedals, wing tips on the handlebars, parts connected rather than floating, sensible scale. 10 = convincingly riding; 5 = next to or loosely attached; 0 = unrelated.
4. polish - Aesthetic polish: line quality, color, balance, absence of rendering glitches, clutter or stray shapes. 10 = professional illustration quality; 5 = passable clip-art; 0 = broken.

Calibration: be strict. Reserve 9-10 for genuinely excellent work with no visible defects. Most competent attempts land between 5 and 8. Judge only what is visible in the image, not what was intended.

Reply with ONLY a JSON object, no prose outside it:
{"pelican": <0-10>, "bicycle": <0-10>, "composition": <0-10>, "polish": <0-10>, "notes": "<one or two sentences naming the main strengths and defects>"}
```

## Run it on *your own Jetty.*

[Task workflow JSON](https://evaljetty.com/pelicans/run/task-workflow.json)
[Get a Jetty account](https://jetty.io/?utm_source=evaljetty&utm_medium=referral&utm_campaign=pelicans&utm_content=run_yourself)

1. ### The Jetty task

   A single runbook step. Agent, model, provider, snapshot, instruction and template variables are read from init\_params, so one task runs any agent/model pair. runbook\_evals\_version: 2 is the stamp that makes Jetty compute and show evaluations for each run. The runbook is inlined as instruction (elided here).

   ```
   {
     "steps": [
       "run"
     ],
     "init_params": {
       "vars": {
         "prompt": "Execute the runbook end-to-end.",
         "max_rounds": 3,
         "results_dir": "/app/results",
         "pass_threshold": 28
       },
       "agent": "claude-code",
       "model": "anthropic/claude-sonnet-5",
       "snapshot": "python312-uv",
       "file_paths": [],
       "instruction": "<contents of pelican-bicycle-svg.runbook.md>",
       "model_provider": "openrouter",
       "runbook_evals_version": 2
     },
     "step_configs": {
       "run": {
         "cpus": 4,
         "memory": "8G",
         "activity": "runbook",
         "agent_path": "init_params.agent",
         "files_path": "init_params.file_paths",
         "model_path": "init_params.model",
         "timeout_sec": 1800,
         "snapshot_path": "init_params.snapshot",
         "network_enabled": true,
         "instruction_path": "init_params.instruction",
         "model_provider_path": "init_params.model_provider",
         "template_variables_path": "init_params.vars"
       }
     }
   }
   ```
2. ### Set the key and create the task

   The collection needs one secret, OPENROUTER\_API\_KEY, stored as a collection environment variable. Secrets never go in init\_params.

   ```
   export JETTY_TOKEN=mlc_...            # your Jetty API key
   export COLLECTION=your-collection   # a collection you own

   # 1. Store your OpenRouter key as a collection environment variable (sent via stdin, not argv)
   printf '{"environment_variables": {"OPENROUTER_API_KEY": "%s"}}' "$OPENROUTER_API_KEY" | \
     curl -s -X PATCH "https://flows-api.jetty.io/api/v1/collections/$COLLECTION/environment" \
       -H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" --data-binary @-

   # 2. Create the task from the workflow JSON (the v2 runbook is embedded as init_params.instruction,
   #    and init_params.runbook_evals_version = 2 tells Jetty to compute evaluations for every run)
   curl -sO https://evaljetty.com/pelicans/run/task-workflow.json
   jq '{name: "pelican-bicycle-svg", description: "Pelican riding a bicycle, SVG", workflow: .}' task-workflow.json | \
     curl -s -X POST "https://flows-api.jetty.io/api/v1/tasks/$COLLECTION" \
       -H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" --data-binary @-
   ```
3. ### Launch runs

   See [02 · Runs](https://evaljetty.com/pelicans/runs/index.html#launch) for the launch and download calls.
4. ### Hill-climb it

   The orchestrator we used for v2: each round embeds the previous best SVG into the runbook and targets the weakest self-scored axis. It stops launching runs past a spend budget.

   ```
   # Optional: hill-climb (5 rounds per agent, agents in parallel, $60 spend guard)
   curl -sO https://evaljetty.com/pelicans/run/hill_climb.py
   curl -so runbook.md https://evaljetty.com/pelicans/run/pelican-bicycle-svg.runbook.md
   # edit COLLECTION at the top of hill_climb.py, then:
   JETTY_TOKEN=$JETTY_TOKEN python3 hill_climb.py --agent all --rounds 5 --budget 60
   ```

## Why *this prompt.*

Simon Willison has run the pelican prompt against nearly every major model release since 2024. He picked it on purpose:

> They shouldn't be able to draw anything at all. But they can generate code… and SVG is code.
>
> Simon Willison, June 2025

> Most importantly: pelicans can't ride bicycles. They're the wrong shape!
>
> Simon Willison

Drawing a pelican on a bicycle in SVG makes a text model reason in code about shape, anatomy and composition all at once, and because pelicans can't ride bicycles the result can't be copied from training data. [Simon's pelican posts ↗](https://simonwillison.net/tags/pelican-riding-a-bicycle/)

---
Source: https://evaljetty.com/pelicans/author.html · Evals by Jetty · Run your own evals: https://jetty.io/?utm_source=evaljetty&utm_medium=llms
