evals
Run evals on Jetty

01 — Author

Tell Jetty
the job to be done.

Everything the agents did came from one Markdown runbook: the objective, the constraints, a four-axis rubric and a self-critique loop. Here is that runbook, how the hill climb rewrote it each round, the external judge we added on top, and how to run it on your own Jetty.

A pelican, riding a bicycle, in SVG.

Hand-write a pure-XML SVG of a pelican riding a bicycle. Both subjects must be unmistakable, and the pelican has to be riding the bike (on the seat, feet at the pedals, wing tips on the bars) rather than floating next to it. The agent draws, renders the SVG to PNG with rsvg-convert, looks at the PNG, scores it and redraws, then keeps its best round.

SVG constraints

  • Pure SVG elements only: no <image> tags, no base64 or external rasters
  • Well-formed XML that renders in any modern browser
  • viewBox about 0 0 800 600, file size under 50 KB

The rubric

AxisQuestionScale
PelicanWould a stranger immediately say “that's a pelican”?0–10
BicycleTwo spoked wheels, a real frame, handlebars, seat, pedals?0–10
CompositionIs the pelican clearly riding the bike?0–10
PolishLine quality, colour, balance0–10

Total out of 40. The agent scores itself with this rubric; the external judge uses the same four axes.

Version 2.0.0, in the evals-v2 shape.

What v2 evaluates on every run

KindIdWhat it checks
code checkoutputs-existEvery required output file exists and is non-empty
code checksvg-parsesfinal.svg is well-formed XML with an <svg> root
code checksvg-size-under-50kbfinal.svg is under 50 KB
code checkno-image-tagsfinal.svg has no <image> elements or embedded rasters
code checkpng-renderedfinal.png is a real PNG at least 400 px wide
code checkrounds-recordedEvery scored round has its SVG and PNG
checklistpelican-recognizablethe bird reads as a pelican at a glance (long beak with a visible throat pouch)
checklistbicycle-completetwo spoked wheels, a connected frame, handlebars, seat and pedals are all present
checklistactually-ridingthe pelican sits on the seat with wing tips on the handlebars and feet at the pedals
checklistscores-match-reportreport.md, scores.json and final.svg agree on the per-round scores and the best round
judgeself-scoreThe agent's own 4-axis score of its best round, out of 40; passes at ≥ {{pass_threshold}} (default 28)
current · v2.0.0 · evals v2 download .md
---
name: pelican-bicycle-svg
version: "2.0.0"
evals_version: 2
description: Hand-write the best possible SVG of a pelican riding a bicycle, refining it over several self-critique rounds, and write an evals-v2 validation report.
agent: claude-code
model: anthropic/claude-sonnet-5
model_provider: openrouter
snapshot: python312-uv
evaluation: programmatic
primary_outputs:
  - final.svg
  - final.png
  - report.md
secrets:
  OPENROUTER_API_KEY:
    env: OPENROUTER_API_KEY
    description: "OpenRouter key (collection environment variable) used by the agent's model calls"
    required: true
---

Pelican Riding a Bicycle: SVG Runbook

Objective

Produce the highest-quality hand-written SVG of a pelican riding a bicycle and deliver it, a PNG render, the per-round scores and an evals-v2 validation report to {{results_dir}}.

Both subjects must be unmistakable, and the composition must be coherent: the pelican is riding the bicycle (on the seat, feet at the pedals, wing tips on the handlebars), not floating next to it. You draw, render the SVG to PNG, look at the PNG, score it on a four-axis rubric and redraw, for up to {{max_rounds}} rounds, then keep the best round.

Hard constraints on every SVG you write:

  • Pure SVG XML: <path>, <circle>, <ellipse>, <rect>, <polygon>, <line>, <g>, <defs>, gradients and so on
  • No <image> elements and no embedded rasters (data:image/..., base64, external bitmaps)
  • Well-formed XML that renders in any modern browser
  • viewBox approximately 0 0 800 600 (landscape)
  • File size under 50 KB

REQUIRED OUTPUT FILES

  • {{results_dir}}/final.svg: the best round's SVG
  • {{results_dir}}/final.png: final.svg rendered with rsvg-convert -w 800
  • {{results_dir}}/rounds/v<N>.svg and rounds/v<N>.png: every round you attempted
  • {{results_dir}}/scores.json: your rubric scores per round and the best round (format in step 4)
  • {{results_dir}}/report.md: per-round score table and notes (format in step 4)
  • {{results_dir}}/validation_report.json: evals-v2 report, written in the last section

Parameters

ParameterTemplate VariableDefaultDescription
Results directory{{results_dir}}/app/resultsOutput directory (persisted by Jetty)
Rounds{{max_rounds}}3Maximum draw, render, critique rounds (the loop exits early at 36/40)
Pass threshold{{pass_threshold}}28Minimum self-score total (out of 40) for the self-score judge check to pass

Dependencies

  • python312-uv snapshot (Python 3.12 for the checks and the report)
  • An SVG rasterizer: rsvg-convert from librsvg2-bin (installed in step 1 if missing)
  • OPENROUTER_API_KEY collection environment variable for the agent's own model calls

Steps

1. Setup

mkdir -p {{results_dir}}/rounds && cd {{results_dir}}
which rsvg-convert || (apt-get update && apt-get install -y librsvg2-bin) || true
which rsvg-convert

2. Round 1: first draft

Write your best first attempt to {{results_dir}}/rounds/v1.svg. Aim to depict:

  • Pelican: long beak with throat pouch (the iconic silhouette), body, eye, wings, legs and feet
  • Bicycle: two wheels with spokes, frame (top tube, down tube, seat tube), handlebars, seat, pedals
  • Riding: body on the seat, feet on or near the pedals, wing tips on the handlebars

Render it:

cd {{results_dir}} && rsvg-convert -w 800 rounds/v1.svg -o rounds/v1.png

3. Critique and refine

For each round R from 1 to {{max_rounds}}:

  1. Look at rounds/v${R}.png (open it as an image; do not score from the SVG source).
  2. Score it on four axes, 0 to 10 each:
    Pelican: would a stranger immediately say "that's a pelican"?
    Bicycle: would they say "that's a bicycle"?
    Composition: is the pelican clearly riding the bike?
    Polish: line quality, colour, balance
  3. If the total is 36/40 or more, or R equals {{max_rounds}}, stop the loop.
  4. Otherwise name the lowest-scoring axis, write down two or three concrete fixes, write rounds/v$((R+1)).svg applying them (keep what worked) and render rounds/v$((R+1)).png.

Score honestly: the benchmark re-scores every drawing with an external vision judge, and agents typically rate themselves several points above it. Score what the PNG shows, not what you intended to draw.

4. Finalize

Pick the round with the highest total (break ties in favour of the later round) and copy it:

cd {{results_dir}} && cp rounds/vBEST.svg final.svg && cp rounds/vBEST.png final.png

Write {{results_dir}}/scores.json:

{"best_round": 2, "rounds": [
  {"round": 1, "pelican": 7, "bicycle": 8, "composition": 6, "polish": 7, "total": 28},
  {"round": 2, "pelican": 8, "bicycle": 8, "composition": 7, "polish": 8, "total": 31}]}

Write {{results_dir}}/report.md with a ## Per-round scores table (Round, Pelican, Bicycle, Composition, Polish, Total), the best round, and short What worked, What didn't work and SVG technique notes sections (viewBox, shape count, file size).

Code Checks

Each check is a deterministic command; a failing check fails the run. The report script in the last section runs the same checks, so run these for your own verification first.

outputs-exist — Every required output file exists and is non-empty

cd {{results_dir}} && for f in final.svg final.png scores.json report.md; do test -s "$f" || { echo "missing $f"; exit 1; }; done

svg-parses — final.svg is well-formed XML with an <svg> root

python3 -c "import xml.etree.ElementTree as E; r=E.parse('{{results_dir}}/final.svg').getroot(); assert r.tag.split('}')[-1]=='svg', r.tag"

svg-size-under-50kb — final.svg is under 50 KB

test "$(wc -c < {{results_dir}}/final.svg)" -lt 51200

no-image-tags — final.svg has no <image> elements or embedded rasters

python3 -c "import xml.etree.ElementTree as E; s=open('{{results_dir}}/final.svg').read(); r=E.fromstring(s); assert not [e for e in r.iter() if e.tag.split('}')[-1]=='image'] and 'data:image' not in s and 'base64' not in s"

png-rendered — final.png is a real PNG at least 400 px wide

python3 -c "import struct; b=open('{{results_dir}}/final.png','rb').read(24); assert b[:8]==b'\x89PNG\r\n\x1a\n', 'not a PNG'; w=struct.unpack('>I', b[16:20])[0]; assert w>=400, w"

rounds-recorded — Every scored round has its SVG and PNG

python3 -c "import json,os; d=json.load(open('{{results_dir}}/scores.json')); ns=[r['round'] for r in d['rounds']]; assert ns and d['best_round'] in ns; miss=[n for n in ns for e in ('svg','png') if not os.path.getsize('{{results_dir}}/rounds/v%d.%s'%(n,e))]; assert not miss, miss"

Checklist

Confirm each item by looking at final.png; a failed item fails the run.

  • pelican-recognizable: the bird reads as a pelican at a glance (long beak with a visible throat pouch)
  • bicycle-complete: two spoked wheels, a connected frame, handlebars, seat and pedals are all present
  • actually-riding: the pelican sits on the seat with wing tips on the handlebars and feet at the pedals
  • scores-match-report: report.md, scores.json and final.svg agree on the per-round scores and the best round

Write Validation Report

Write {{results_dir}}/validation_report.json in the evals-v2 shape by running the script below. It runs every code check above, records one step entry per step, one checklist entry per item you pass in (status, a colon, then a one-line reason) and your best-round self-score as a judge check (score out of 40, threshold {{pass_threshold}}). Jetty computes the verdict from these checks. Use fail: <reason> for anything that did not hold, and set ITERATIONS to the number of rounds you drew.

cd {{results_dir}} && ITERATIONS=3 \
STEPS='{"setup": "pass", "draft": "pass", "refine": "pass", "finalize": "pass"}' \
CHECKLIST='{"pelican-recognizable": "pass: <why>", "bicycle-complete": "pass: <why>", "actually-riding": "pass: <why>", "scores-match-report": "pass: <why>"}' \
python3 - <<'PY'
import json, os, struct, xml.etree.ElementTree as E
R = "{{results_dir}}"; T = "{{pass_threshold}}"; T = float(T) if T.replace(".", "", 1).isdigit() else 28.0
checks = []
def add(kind, cid, name, ok, msg, **kw):
    checks.append({"kind": kind, "id": cid, "name": name, "status": "pass" if ok else "fail", "message": str(msg)[:500], **kw})
def run(cid, name, fn):
    try:
        add("code_check", cid, name, True, fn() or "ok")
    except Exception as e:
        add("code_check", cid, name, False, repr(e))
p = lambda f: os.path.join(R, f)
def outputs():
    miss = [f for f in ("final.svg", "final.png", "scores.json", "report.md") if not os.path.exists(p(f)) or not os.path.getsize(p(f))]
    assert not miss, "missing: " + ", ".join(miss)
def parses():
    r = E.parse(p("final.svg")).getroot(); assert r.tag.split("}")[-1] == "svg", r.tag
def size():
    n = os.path.getsize(p("final.svg")); assert n < 51200, n; return f"{n} bytes"
def noimg():
    s = open(p("final.svg")).read(); r = E.fromstring(s)
    assert not [e for e in r.iter() if e.tag.split("}")[-1] == "image"] and "data:image" not in s and "base64" not in s
def png():
    b = open(p("final.png"), "rb").read(24); assert b[:8] == b"\x89PNG\r\n\x1a\n", "not a PNG"
    w = struct.unpack(">I", b[16:20])[0]; assert w >= 400, w; return f"{w} px wide"
def rounds():
    d = json.load(open(p("scores.json"))); ns = [r["round"] for r in d["rounds"]]
    assert ns and d["best_round"] in ns, "best_round not among rounds"
    miss = [f"v{n}.{e}" for n in ns for e in ("svg", "png") if not os.path.exists(p(f"rounds/v{n}.{e}"))]
    assert not miss, miss; return f"{len(ns)} rounds"
run("outputs-exist", "Every required output file exists and is non-empty", outputs)
run("svg-parses", "final.svg is well-formed XML with an <svg> root", parses)
run("svg-size-under-50kb", "final.svg is under 50 KB", size)
run("no-image-tags", "final.svg has no <image> elements or embedded rasters", noimg)
run("png-rendered", "final.png is a real PNG at least 400 px wide", png)
run("rounds-recorded", "Every scored round has its SVG and PNG", rounds)
for sid, st in json.loads(os.environ["STEPS"]).items():
    add("step", sid, sid, st.startswith("pass"), st)
for cid, st in json.loads(os.environ["CHECKLIST"]).items():
    ok = st.strip().lower().startswith("pass"); add("checklist", cid, cid, ok, st.split(":", 1)[-1].strip())
try:
    d = json.load(open(p("scores.json"))); best = next(r for r in d["rounds"] if r["round"] == d["best_round"])
    add("judge", "self-score", "Agent self-score of the best round (4 axes x 0-10)", best["total"] >= T,
        f"round {d['best_round']}: " + ", ".join(f"{k} {best[k]}" for k in ("pelican", "bicycle", "composition", "polish")),
        score=float(best["total"]), max_score=40.0, threshold=T)
except Exception as e:
    checks.append({"kind": "judge", "id": "self-score", "name": "Agent self-score of the best round", "status": "error", "message": repr(e)})
gating = [c for c in checks if c["kind"] != "step"]
passed = all(c["status"] == "pass" for c in gating)
rep = {"version": 2, "verdict": "pass" if passed else "fail", "overall_passed": passed,
       "iterations": int(os.environ.get("ITERATIONS", "1")), "checks": checks}
json.dump(rep, open(p("validation_report.json"), "w"), indent=2)
for c in checks: print(f"{c['status']:5} {c['kind']:10} {c['id']}: {c['message']}")
print("verdict:", rep["verdict"])
PY

Do not hand-edit the file and do not invent a different format: it must stay {"version": 2, "verdict", "overall_passed", "iterations", "checks": [...]}, with each check carrying kind (step, code_check, checklist or judge), id, name, status (pass, fail, skipped or error) and message, plus score, max_score and threshold on the judge.

Tips

  • One round costs roughly $0.50 to $2.50 depending on the model; most of it is the agent's own reasoning.
  • Seeding round 1 with a previous run's final.svg (paste it into the runbook) is the single biggest improvement we measured in hand-edited climbs: it skips the cold start.
  • Coordinate-precise fixes ("close the 13 px gap between the right wing tip and the grip") move scores more than scene-level additions (sun, motion lines, scenery), which tend to add clutter.
  • The self-score is a weak signal across models: at evaljetty.com/pelicans it correlates only loosely with an external judge. Treat it as a steering aid for your own climb, not as a ranking.

Changelog

  • 2.0.0: Evals v2 shape (Parameters table, Code Checks, Checklist, Write Validation Report with the self-score as a judge check at {{pass_threshold}}/40); scores.json output; honest-scoring note.
  • 1.x: The runbook used for every run published at evaljetty.com/pelicans (v1 cohort May 2026, v2 cohort Sep 2026): draw, render, self-score and redraw for max_rounds, then final.svg, final.png and report.md.
Runbook v1: the version every published run used (download: .md)
---
name: pelican-bicycle-svg
version: 2
description: Generate the best possible hand-written SVG depicting a pelican riding a bicycle, with iterative self-refinement.
agent: claude-code
model: anthropic/claude-sonnet-5
model_provider: openrouter
parameters:
  max_rounds:
    type: integer
    default: 3
    description: Number of refinement rounds (each round = critique previous + write improved)
  results_dir:
    type: string
    default: /app/results
secrets:
  openrouter:
    env: OPENROUTER_API_KEY
    description: OpenRouter key (collection environment variable) used by the agent's model calls
evaluation:
  pattern: self-judge
---

Pelican Riding a Bicycle — SVG Runbook

Mission

Produce the highest-quality hand-written SVG depicting a pelican riding a bicycle. Both subjects must be unmistakably recognizable, the composition coherent (pelican IS interacting with the bicycle, not floating next to it), and the file should be pure XML SVG (no <image> tags, no base64 data, no external rasters).

Hard constraints

  • Pure SVG XML — use <path>, <circle>, <ellipse>, <rect>, <polygon>, <line>, <g>, <defs>, <linearGradient>, etc.
  • No <image> elements with external or base64 data
  • Must validate as well-formed XML and render in any modern browser
  • viewBox approximately 0 0 800 600 (landscape) — adjust if you have a specific reason
  • Total file size under 50KB

INPUTS

  • max_rounds = 3 (override only if explicitly told otherwise)
  • results_dir = /app/results

Steps

1. Setup

mkdir -p {{results_dir}}/rounds
cd {{results_dir}}

Check that rsvg-convert or an SVG rasterizer is available:

which rsvg-convert || which inkscape || which convert

If none are available, install one:

apt-get update && apt-get install -y librsvg2-bin || true

2. Round 1 — first draft

Write your best first attempt to {{results_dir}}/rounds/v1.svg. Aim to depict:

  • Pelican: long beak with throat pouch (the iconic pelican silhouette), body, eye, wings, legs/feet
  • Bicycle: two wheels with spokes, frame (top tube + down tube + seat tube), handlebars, seat, pedals
  • Riding: pelican's body is on the seat, feet on or near pedals, "hands" (wing tips) on handlebars

Render it:

cd {{results_dir}}
rsvg-convert -w 800 rounds/v1.svg -o rounds/v1.png

3. Self-critique + refinement loop

For each round R from 1 to max_rounds - 1:

  1. Inspect rounds/v${R}.png visually (read it as an image).
  2. Score on four axes, 0–10 each:
    Pelican recognizability — would a stranger immediately say "that's a pelican"?
    Bicycle recognizability — would they say "that's a bicycle"?
    Composition / riding — is the pelican clearly riding the bike?
    Aesthetic polish — line quality, color, balance
  3. Identify the lowest-scoring axis and write down 2-3 concrete fixes.
  4. Write rounds/v$((R+1)).svg applying those fixes. Keep what worked; rewrite what didn't.
  5. Render rounds/v$((R+1)).png.
  6. If the total (sum of 4 axes) is ≥ 36/40 — early-exit the loop.

4. Finalize

After the loop, pick the round with the highest total score (break ties by preferring later rounds since they had more iteration).

cp {{results_dir}}/rounds/vBEST.svg {{results_dir}}/final.svg
cp {{results_dir}}/rounds/vBEST.png {{results_dir}}/final.png

Write {{results_dir}}/report.md containing:

# Pelican-Bicycle SVG Report

## Per-round scores
| Round | Pelican | Bicycle | Composition | Polish | Total |
|-------|---------|---------|-------------|--------|-------|
| 1 | ? | ? | ? | ? | ? |
| 2 | ? | ? | ? | ? | ? |
| 3 | ? | ? | ? | ? | ? |

## Best round: vN (Total: X/40)

## What worked
- ...

## What didn't work
- ...

## SVG technique notes
- viewBox used:
- Total path / shape count:
- File size:

Final checklist

  • {{results_dir}}/final.svg exists, is valid XML, renders in a browser, has no <image> tags
  • {{results_dir}}/final.png exists (rasterized version)
  • {{results_dir}}/rounds/v*.svg and v*.png exist for every round attempted
  • {{results_dir}}/report.md exists with per-round score table and a notes section
  • File size of final.svg is under 50KB

How the hill climb rewrites the runbook.

Each hill-climb round is one Jetty run. Between rounds a small orchestrator (hill_climb.py) rewrites the runbook: it embeds the previous round's best SVG as the round-1 seed, reads the per-round scores from report.md, and rewrites the description to target the lowest self-scored axis. v1 ran 10 rounds per agent (3 for the Fusion sweeps); v2 ran 5.

One lineage was edited by hand instead: Jon read each claude-code + Sonnet 4.6 trajectory and edited the runbook himself, eleven times. Three of those edits:

Big leap · v2 → v3 · self 32 → 34 · judge 25.2 → 24.5
Embed the baseline
v2
v2 · before
v3
v3 · after

The edit: Use the v2 SVG verbatim as the round-1 input. Don't start from scratch.

What happened: Round 1 jumped from 28 to 33 self-score points just by skipping the cold start. The single highest-leverage edit in the sequence.

Lesson: When a runbook can carry a working artifact as a seed, it should.

Peak · v4 → v5 · self 36 → 37 · judge 24.5 → 26.2
Coordinate-precise asks
v4
v4 · before
v5
v5 · after

The edit: Close the 13 px right-wing-to-grip gap. Drop the left-wing tips to y ≈ 230.

What happened: Composition hit 10/10 on the self-score for the first time. Sonnet's peak self-score, never beaten later.

Lesson: At the top of the curve, the asks become measurements.

Cautionary · v5 → v6 · self 37 → 36 · judge 26.2 → 26.5
Scope creep cost a point
v5
v5 · before
v6
v6 · after

The edit: Add motion lines, wind tufts and a sun. Extend the right-wing covert lines to the wrist.

What happened: A richer scene, but composition slipped back to 9. The new elements added clutter the rubric couldn't reward.

Lesson: More isn't better when the rubric doesn't measure scene density.

One fixed yardstick for every drawing.

Judge model
Opus 5.5
anthropic/claude-opus-5.5 via OpenRouter
Sampling
3 × T=0
independent calls, averaged
Renderer
librsvg
rsvg-convert -w 800, white background
Judge spend
$2.77
all 98 drawings

The system prompt, rubric and model are fixed across all drawings. Raw responses, including token usage, are published in judge-raw.json. A single LLM judge has biases of its own: it is one fixed yardstick, not ground truth, and the judge model is also one of the v2 contestants.

The image below is a rasterized SVG that a coding agent hand-wrote for the prompt:
"Generate an SVG of a pelican riding a bicycle."

Score it on four axes, each an integer or half-integer from 0 to 10:

1. pelican - Pelican recognizability. Would a stranger immediately say "that's a pelican"? Look for the long beak with a throat pouch, pelican body shape, eye, wings, legs/feet. 10 = unmistakably a pelican; 5 = a generic bird; 0 = no bird.
2. bicycle - Bicycle recognizability and correctness. Two wheels (with spokes), a plausible frame (top tube, down tube, seat tube), handlebars, seat, pedals/crank. 10 = a correct, well-formed bicycle; 5 = recognizable but structurally wrong; 0 = no bicycle.
3. composition - Is the pelican clearly RIDING the bicycle? Body on the seat, feet on or near the pedals, wing tips on the handlebars, parts connected rather than floating, sensible scale. 10 = convincingly riding; 5 = next to or loosely attached; 0 = unrelated.
4. polish - Aesthetic polish: line quality, color, balance, absence of rendering glitches, clutter or stray shapes. 10 = professional illustration quality; 5 = passable clip-art; 0 = broken.

Calibration: be strict. Reserve 9-10 for genuinely excellent work with no visible defects. Most competent attempts land between 5 and 8. Judge only what is visible in the image, not what was intended.

Reply with ONLY a JSON object, no prose outside it:
{"pelican": <0-10>, "bicycle": <0-10>, "composition": <0-10>, "polish": <0-10>, "notes": "<one or two sentences naming the main strengths and defects>"}

Run it on your own Jetty.

  1. The Jetty task

    A single runbook step. Agent, model, provider, snapshot, instruction and template variables are read from init_params, so one task runs any agent/model pair. runbook_evals_version: 2 is the stamp that makes Jetty compute and show evaluations for each run. The runbook is inlined as instruction (elided here).

    {
      "steps": [
        "run"
      ],
      "init_params": {
        "vars": {
          "prompt": "Execute the runbook end-to-end.",
          "max_rounds": 3,
          "results_dir": "/app/results",
          "pass_threshold": 28
        },
        "agent": "claude-code",
        "model": "anthropic/claude-sonnet-5",
        "snapshot": "python312-uv",
        "file_paths": [],
        "instruction": "<contents of pelican-bicycle-svg.runbook.md>",
        "model_provider": "openrouter",
        "runbook_evals_version": 2
      },
      "step_configs": {
        "run": {
          "cpus": 4,
          "memory": "8G",
          "activity": "runbook",
          "agent_path": "init_params.agent",
          "files_path": "init_params.file_paths",
          "model_path": "init_params.model",
          "timeout_sec": 1800,
          "snapshot_path": "init_params.snapshot",
          "network_enabled": true,
          "instruction_path": "init_params.instruction",
          "model_provider_path": "init_params.model_provider",
          "template_variables_path": "init_params.vars"
        }
      }
    }
  2. Set the key and create the task

    The collection needs one secret, OPENROUTER_API_KEY, stored as a collection environment variable. Secrets never go in init_params.

    export JETTY_TOKEN=mlc_...            # your Jetty API key
    export COLLECTION=your-collection   # a collection you own
    
    # 1. Store your OpenRouter key as a collection environment variable (sent via stdin, not argv)
    printf '{"environment_variables": {"OPENROUTER_API_KEY": "%s"}}' "$OPENROUTER_API_KEY" | \
      curl -s -X PATCH "https://flows-api.jetty.io/api/v1/collections/$COLLECTION/environment" \
        -H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" --data-binary @-
    
    # 2. Create the task from the workflow JSON (the v2 runbook is embedded as init_params.instruction,
    #    and init_params.runbook_evals_version = 2 tells Jetty to compute evaluations for every run)
    curl -sO https://evaljetty.com/pelicans/run/task-workflow.json
    jq '{name: "pelican-bicycle-svg", description: "Pelican riding a bicycle, SVG", workflow: .}' task-workflow.json | \
      curl -s -X POST "https://flows-api.jetty.io/api/v1/tasks/$COLLECTION" \
        -H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" --data-binary @-
  3. Launch runs

    See 02 · Runs for the launch and download calls.

  4. Hill-climb it

    The orchestrator we used for v2: each round embeds the previous best SVG into the runbook and targets the weakest self-scored axis. It stops launching runs past a spend budget.

    # Optional: hill-climb (5 rounds per agent, agents in parallel, $60 spend guard)
    curl -sO https://evaljetty.com/pelicans/run/hill_climb.py
    curl -so runbook.md https://evaljetty.com/pelicans/run/pelican-bicycle-svg.runbook.md
    # edit COLLECTION at the top of hill_climb.py, then:
    JETTY_TOKEN=$JETTY_TOKEN python3 hill_climb.py --agent all --rounds 5 --budget 60

Why this prompt.

Simon Willison has run the pelican prompt against nearly every major model release since 2024. He picked it on purpose:

They shouldn't be able to draw anything at all. But they can generate code… and SVG is code.

Simon Willison, June 2025

Most importantly: pelicans can't ride bicycles. They're the wrong shape!

Simon Willison

Drawing a pelican on a bicycle in SVG makes a text model reason in code about shape, anatomy and composition all at once, and because pelicans can't ride bicycles the result can't be copied from training data. Simon's pelican posts ↗

Ready to stop guessing if your outputs are good?

Get Started Book a 15-minute walkthrough
Connect your agent to Jetty →