---
name: pelican-bicycle-svg
version: "2.0.0"
evals_version: 2
description: Hand-write the best possible SVG of a pelican riding a bicycle, refining it over several self-critique rounds, and write an evals-v2 validation report.
agent: claude-code
model: anthropic/claude-sonnet-5
model_provider: openrouter
snapshot: python312-uv
evaluation: programmatic
primary_outputs:
  - final.svg
  - final.png
  - report.md
secrets:
  OPENROUTER_API_KEY:
    env: OPENROUTER_API_KEY
    description: "OpenRouter key (collection environment variable) used by the agent's model calls"
    required: true
---

# Pelican Riding a Bicycle: SVG Runbook

## Objective

Produce the highest-quality **hand-written SVG of a pelican riding a bicycle** and deliver it, a PNG render, the
per-round scores and an evals-v2 validation report to `{{results_dir}}`.

Both subjects must be unmistakable, and the composition must be coherent: the pelican is **riding** the bicycle (on the
seat, feet at the pedals, wing tips on the handlebars), not floating next to it. You draw, render the SVG to PNG, look
at the PNG, score it on a four-axis rubric and redraw, for up to `{{max_rounds}}` rounds, then keep the best round.

Hard constraints on every SVG you write:

- Pure SVG XML: `<path>`, `<circle>`, `<ellipse>`, `<rect>`, `<polygon>`, `<line>`, `<g>`, `<defs>`, gradients and so on
- **No** `<image>` elements and no embedded rasters (`data:image/...`, base64, external bitmaps)
- Well-formed XML that renders in any modern browser
- viewBox approximately `0 0 800 600` (landscape)
- File size under 50 KB

## REQUIRED OUTPUT FILES

- `{{results_dir}}/final.svg`: the best round's SVG
- `{{results_dir}}/final.png`: `final.svg` rendered with `rsvg-convert -w 800`
- `{{results_dir}}/rounds/v<N>.svg` and `rounds/v<N>.png`: every round you attempted
- `{{results_dir}}/scores.json`: your rubric scores per round and the best round (format in step 4)
- `{{results_dir}}/report.md`: per-round score table and notes (format in step 4)
- `{{results_dir}}/validation_report.json`: evals-v2 report, written in the last section

## Parameters

| Parameter | Template Variable | Default | Description |
|-----------|-------------------|---------|-------------|
| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |
| Rounds | `{{max_rounds}}` | `3` | Maximum draw, render, critique rounds (the loop exits early at 36/40) |
| Pass threshold | `{{pass_threshold}}` | `28` | Minimum self-score total (out of 40) for the `self-score` judge check to pass |

## Dependencies

- `python312-uv` snapshot (Python 3.12 for the checks and the report)
- An SVG rasterizer: `rsvg-convert` from `librsvg2-bin` (installed in step 1 if missing)
- `OPENROUTER_API_KEY` collection environment variable for the agent's own model calls

## Steps

### 1. Setup

```bash
mkdir -p {{results_dir}}/rounds && cd {{results_dir}}
which rsvg-convert || (apt-get update && apt-get install -y librsvg2-bin) || true
which rsvg-convert
```

### 2. Round 1: first draft

Write your best first attempt to `{{results_dir}}/rounds/v1.svg`. Aim to depict:

- **Pelican**: long beak with throat pouch (the iconic silhouette), body, eye, wings, legs and feet
- **Bicycle**: two wheels with spokes, frame (top tube, down tube, seat tube), handlebars, seat, pedals
- **Riding**: body on the seat, feet on or near the pedals, wing tips on the handlebars

Render it:

```bash
cd {{results_dir}} && rsvg-convert -w 800 rounds/v1.svg -o rounds/v1.png
```

### 3. Critique and refine

For each round R from 1 to `{{max_rounds}}`:

1. **Look** at `rounds/v${R}.png` (open it as an image; do not score from the SVG source).
2. **Score** it on four axes, 0 to 10 each:
   - **Pelican**: would a stranger immediately say "that's a pelican"?
   - **Bicycle**: would they say "that's a bicycle"?
   - **Composition**: is the pelican clearly riding the bike?
   - **Polish**: line quality, colour, balance
3. If the total is 36/40 or more, or R equals `{{max_rounds}}`, stop the loop.
4. Otherwise name the lowest-scoring axis, write down two or three concrete fixes, write `rounds/v$((R+1)).svg`
   applying them (keep what worked) and render `rounds/v$((R+1)).png`.

Score honestly: the benchmark re-scores every drawing with an external vision judge, and agents typically rate
themselves several points above it. Score what the PNG shows, not what you intended to draw.

### 4. Finalize

Pick the round with the highest total (break ties in favour of the later round) and copy it:

```bash
cd {{results_dir}} && cp rounds/vBEST.svg final.svg && cp rounds/vBEST.png final.png
```

Write `{{results_dir}}/scores.json`:

```json
{"best_round": 2, "rounds": [
  {"round": 1, "pelican": 7, "bicycle": 8, "composition": 6, "polish": 7, "total": 28},
  {"round": 2, "pelican": 8, "bicycle": 8, "composition": 7, "polish": 8, "total": 31}]}
```

Write `{{results_dir}}/report.md` with a `## Per-round scores` table (Round, Pelican, Bicycle, Composition, Polish,
Total), the best round, and short **What worked**, **What didn't work** and **SVG technique notes** sections (viewBox,
shape count, file size).

## Code Checks

Each check is a deterministic command; a failing check fails the run. The report script in the last section runs
the same checks, so run these for your own verification first.

### outputs-exist — Every required output file exists and is non-empty

```bash
cd {{results_dir}} && for f in final.svg final.png scores.json report.md; do test -s "$f" || { echo "missing $f"; exit 1; }; done
```

### svg-parses — final.svg is well-formed XML with an <svg> root

```bash
python3 -c "import xml.etree.ElementTree as E; r=E.parse('{{results_dir}}/final.svg').getroot(); assert r.tag.split('}')[-1]=='svg', r.tag"
```

### svg-size-under-50kb — final.svg is under 50 KB

```bash
test "$(wc -c < {{results_dir}}/final.svg)" -lt 51200
```

### no-image-tags — final.svg has no <image> elements or embedded rasters

```bash
python3 -c "import xml.etree.ElementTree as E; s=open('{{results_dir}}/final.svg').read(); r=E.fromstring(s); assert not [e for e in r.iter() if e.tag.split('}')[-1]=='image'] and 'data:image' not in s and 'base64' not in s"
```

### png-rendered — final.png is a real PNG at least 400 px wide

```bash
python3 -c "import struct; b=open('{{results_dir}}/final.png','rb').read(24); assert b[:8]==b'\x89PNG\r\n\x1a\n', 'not a PNG'; w=struct.unpack('>I', b[16:20])[0]; assert w>=400, w"
```

### rounds-recorded — Every scored round has its SVG and PNG

```bash
python3 -c "import json,os; d=json.load(open('{{results_dir}}/scores.json')); ns=[r['round'] for r in d['rounds']]; assert ns and d['best_round'] in ns; miss=[n for n in ns for e in ('svg','png') if not os.path.getsize('{{results_dir}}/rounds/v%d.%s'%(n,e))]; assert not miss, miss"
```

## Checklist

Confirm each item by looking at `final.png`; a failed item fails the run.

- [ ] `pelican-recognizable`: the bird reads as a pelican at a glance (long beak with a visible throat pouch)
- [ ] `bicycle-complete`: two spoked wheels, a connected frame, handlebars, seat and pedals are all present
- [ ] `actually-riding`: the pelican sits on the seat with wing tips on the handlebars and feet at the pedals
- [ ] `scores-match-report`: report.md, scores.json and final.svg agree on the per-round scores and the best round

## Write Validation Report

Write `{{results_dir}}/validation_report.json` in the evals-v2 shape by running the script below. It runs every code
check above, records one `step` entry per step, one `checklist` entry per item you pass in (status, a colon, then a
one-line reason) and your best-round self-score as a `judge` check (score out of 40, threshold `{{pass_threshold}}`).
Jetty computes the verdict from these checks. Use `fail: <reason>` for anything that did not hold, and set
`ITERATIONS` to the number of rounds you drew.

```bash
cd {{results_dir}} && ITERATIONS=3 \
STEPS='{"setup": "pass", "draft": "pass", "refine": "pass", "finalize": "pass"}' \
CHECKLIST='{"pelican-recognizable": "pass: <why>", "bicycle-complete": "pass: <why>", "actually-riding": "pass: <why>", "scores-match-report": "pass: <why>"}' \
python3 - <<'PY'
import json, os, struct, xml.etree.ElementTree as E
R = "{{results_dir}}"; T = "{{pass_threshold}}"; T = float(T) if T.replace(".", "", 1).isdigit() else 28.0
checks = []
def add(kind, cid, name, ok, msg, **kw):
    checks.append({"kind": kind, "id": cid, "name": name, "status": "pass" if ok else "fail", "message": str(msg)[:500], **kw})
def run(cid, name, fn):
    try:
        add("code_check", cid, name, True, fn() or "ok")
    except Exception as e:
        add("code_check", cid, name, False, repr(e))
p = lambda f: os.path.join(R, f)
def outputs():
    miss = [f for f in ("final.svg", "final.png", "scores.json", "report.md") if not os.path.exists(p(f)) or not os.path.getsize(p(f))]
    assert not miss, "missing: " + ", ".join(miss)
def parses():
    r = E.parse(p("final.svg")).getroot(); assert r.tag.split("}")[-1] == "svg", r.tag
def size():
    n = os.path.getsize(p("final.svg")); assert n < 51200, n; return f"{n} bytes"
def noimg():
    s = open(p("final.svg")).read(); r = E.fromstring(s)
    assert not [e for e in r.iter() if e.tag.split("}")[-1] == "image"] and "data:image" not in s and "base64" not in s
def png():
    b = open(p("final.png"), "rb").read(24); assert b[:8] == b"\x89PNG\r\n\x1a\n", "not a PNG"
    w = struct.unpack(">I", b[16:20])[0]; assert w >= 400, w; return f"{w} px wide"
def rounds():
    d = json.load(open(p("scores.json"))); ns = [r["round"] for r in d["rounds"]]
    assert ns and d["best_round"] in ns, "best_round not among rounds"
    miss = [f"v{n}.{e}" for n in ns for e in ("svg", "png") if not os.path.exists(p(f"rounds/v{n}.{e}"))]
    assert not miss, miss; return f"{len(ns)} rounds"
run("outputs-exist", "Every required output file exists and is non-empty", outputs)
run("svg-parses", "final.svg is well-formed XML with an <svg> root", parses)
run("svg-size-under-50kb", "final.svg is under 50 KB", size)
run("no-image-tags", "final.svg has no <image> elements or embedded rasters", noimg)
run("png-rendered", "final.png is a real PNG at least 400 px wide", png)
run("rounds-recorded", "Every scored round has its SVG and PNG", rounds)
for sid, st in json.loads(os.environ["STEPS"]).items():
    add("step", sid, sid, st.startswith("pass"), st)
for cid, st in json.loads(os.environ["CHECKLIST"]).items():
    ok = st.strip().lower().startswith("pass"); add("checklist", cid, cid, ok, st.split(":", 1)[-1].strip())
try:
    d = json.load(open(p("scores.json"))); best = next(r for r in d["rounds"] if r["round"] == d["best_round"])
    add("judge", "self-score", "Agent self-score of the best round (4 axes x 0-10)", best["total"] >= T,
        f"round {d['best_round']}: " + ", ".join(f"{k} {best[k]}" for k in ("pelican", "bicycle", "composition", "polish")),
        score=float(best["total"]), max_score=40.0, threshold=T)
except Exception as e:
    checks.append({"kind": "judge", "id": "self-score", "name": "Agent self-score of the best round", "status": "error", "message": repr(e)})
gating = [c for c in checks if c["kind"] != "step"]
passed = all(c["status"] == "pass" for c in gating)
rep = {"version": 2, "verdict": "pass" if passed else "fail", "overall_passed": passed,
       "iterations": int(os.environ.get("ITERATIONS", "1")), "checks": checks}
json.dump(rep, open(p("validation_report.json"), "w"), indent=2)
for c in checks: print(f"{c['status']:5} {c['kind']:10} {c['id']}: {c['message']}")
print("verdict:", rep["verdict"])
PY
```

Do not hand-edit the file and do not invent a different format: it must stay
`{"version": 2, "verdict", "overall_passed", "iterations", "checks": [...]}`, with each check carrying `kind`
(`step`, `code_check`, `checklist` or `judge`), `id`, `name`, `status` (`pass`, `fail`, `skipped` or `error`) and
`message`, plus `score`, `max_score` and `threshold` on the judge.

## Tips

- One round costs roughly $0.50 to $2.50 depending on the model; most of it is the agent's own reasoning.
- Seeding round 1 with a previous run's `final.svg` (paste it into the runbook) is the single biggest improvement we
  measured in hand-edited climbs: it skips the cold start.
- Coordinate-precise fixes ("close the 13 px gap between the right wing tip and the grip") move scores more than
  scene-level additions (sun, motion lines, scenery), which tend to add clutter.
- The self-score is a weak signal across models: at evaljetty.com/pelicans it correlates only loosely with an external
  judge. Treat it as a steering aid for your own climb, not as a ranking.

## Changelog

- 2.0.0: Evals v2 shape (Parameters table, Code Checks, Checklist, Write Validation Report with the self-score as a
  judge check at `{{pass_threshold}}`/40); `scores.json` output; honest-scoring note.
- 1.x: The runbook used for every run published at evaljetty.com/pelicans (v1 cohort May 2026, v2 cohort Sep 2026):
  draw, render, self-score and redraw for `max_rounds`, then `final.svg`, `final.png` and `report.md`.
