{
  "steps": [
    "run"
  ],
  "init_params": {
    "vars": {
      "prompt": "Execute the runbook end-to-end.",
      "max_rounds": 3,
      "results_dir": "/app/results",
      "pass_threshold": 28
    },
    "agent": "claude-code",
    "model": "anthropic/claude-sonnet-5",
    "snapshot": "python312-uv",
    "file_paths": [],
    "instruction": "---\nname: pelican-bicycle-svg\nversion: \"2.0.0\"\nevals_version: 2\ndescription: Hand-write the best possible SVG of a pelican riding a bicycle, refining it over several self-critique rounds, and write an evals-v2 validation report.\nagent: claude-code\nmodel: anthropic/claude-sonnet-5\nmodel_provider: openrouter\nsnapshot: python312-uv\nevaluation: programmatic\nprimary_outputs:\n  - final.svg\n  - final.png\n  - report.md\nsecrets:\n  OPENROUTER_API_KEY:\n    env: OPENROUTER_API_KEY\n    description: \"OpenRouter key (collection environment variable) used by the agent's model calls\"\n    required: true\n---\n\n# Pelican Riding a Bicycle: SVG Runbook\n\n## Objective\n\nProduce the highest-quality **hand-written SVG of a pelican riding a bicycle** and deliver it, a PNG render, the\nper-round scores and an evals-v2 validation report to `{{results_dir}}`.\n\nBoth subjects must be unmistakable, and the composition must be coherent: the pelican is **riding** the bicycle (on the\nseat, feet at the pedals, wing tips on the handlebars), not floating next to it. You draw, render the SVG to PNG, look\nat the PNG, score it on a four-axis rubric and redraw, for up to `{{max_rounds}}` rounds, then keep the best round.\n\nHard constraints on every SVG you write:\n\n- Pure SVG XML: `<path>`, `<circle>`, `<ellipse>`, `<rect>`, `<polygon>`, `<line>`, `<g>`, `<defs>`, gradients and so on\n- **No** `<image>` elements and no embedded rasters (`data:image/...`, base64, external bitmaps)\n- Well-formed XML that renders in any modern browser\n- viewBox approximately `0 0 800 600` (landscape)\n- File size under 50 KB\n\n## REQUIRED OUTPUT FILES\n\n- `{{results_dir}}/final.svg`: the best round's SVG\n- `{{results_dir}}/final.png`: `final.svg` rendered with `rsvg-convert -w 800`\n- `{{results_dir}}/rounds/v<N>.svg` and `rounds/v<N>.png`: every round you attempted\n- `{{results_dir}}/scores.json`: your rubric scores per round and the best round (format in step 4)\n- `{{results_dir}}/report.md`: per-round score table and notes (format in step 4)\n- `{{results_dir}}/validation_report.json`: evals-v2 report, written in the last section\n\n## Parameters\n\n| Parameter | Template Variable | Default | Description |\n|-----------|-------------------|---------|-------------|\n| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |\n| Rounds | `{{max_rounds}}` | `3` | Maximum draw, render, critique rounds (the loop exits early at 36/40) |\n| Pass threshold | `{{pass_threshold}}` | `28` | Minimum self-score total (out of 40) for the `self-score` judge check to pass |\n\n## Dependencies\n\n- `python312-uv` snapshot (Python 3.12 for the checks and the report)\n- An SVG rasterizer: `rsvg-convert` from `librsvg2-bin` (installed in step 1 if missing)\n- `OPENROUTER_API_KEY` collection environment variable for the agent's own model calls\n\n## Steps\n\n### 1. Setup\n\n```bash\nmkdir -p {{results_dir}}/rounds && cd {{results_dir}}\nwhich rsvg-convert || (apt-get update && apt-get install -y librsvg2-bin) || true\nwhich rsvg-convert\n```\n\n### 2. Round 1: first draft\n\nWrite your best first attempt to `{{results_dir}}/rounds/v1.svg`. Aim to depict:\n\n- **Pelican**: long beak with throat pouch (the iconic silhouette), body, eye, wings, legs and feet\n- **Bicycle**: two wheels with spokes, frame (top tube, down tube, seat tube), handlebars, seat, pedals\n- **Riding**: body on the seat, feet on or near the pedals, wing tips on the handlebars\n\nRender it:\n\n```bash\ncd {{results_dir}} && rsvg-convert -w 800 rounds/v1.svg -o rounds/v1.png\n```\n\n### 3. Critique and refine\n\nFor each round R from 1 to `{{max_rounds}}`:\n\n1. **Look** at `rounds/v${R}.png` (open it as an image; do not score from the SVG source).\n2. **Score** it on four axes, 0 to 10 each:\n   - **Pelican**: would a stranger immediately say \"that's a pelican\"?\n   - **Bicycle**: would they say \"that's a bicycle\"?\n   - **Composition**: is the pelican clearly riding the bike?\n   - **Polish**: line quality, colour, balance\n3. If the total is 36/40 or more, or R equals `{{max_rounds}}`, stop the loop.\n4. Otherwise name the lowest-scoring axis, write down two or three concrete fixes, write `rounds/v$((R+1)).svg`\n   applying them (keep what worked) and render `rounds/v$((R+1)).png`.\n\nScore honestly: the benchmark re-scores every drawing with an external vision judge, and agents typically rate\nthemselves several points above it. Score what the PNG shows, not what you intended to draw.\n\n### 4. Finalize\n\nPick the round with the highest total (break ties in favour of the later round) and copy it:\n\n```bash\ncd {{results_dir}} && cp rounds/vBEST.svg final.svg && cp rounds/vBEST.png final.png\n```\n\nWrite `{{results_dir}}/scores.json`:\n\n```json\n{\"best_round\": 2, \"rounds\": [\n  {\"round\": 1, \"pelican\": 7, \"bicycle\": 8, \"composition\": 6, \"polish\": 7, \"total\": 28},\n  {\"round\": 2, \"pelican\": 8, \"bicycle\": 8, \"composition\": 7, \"polish\": 8, \"total\": 31}]}\n```\n\nWrite `{{results_dir}}/report.md` with a `## Per-round scores` table (Round, Pelican, Bicycle, Composition, Polish,\nTotal), the best round, and short **What worked**, **What didn't work** and **SVG technique notes** sections (viewBox,\nshape count, file size).\n\n## Code Checks\n\nEach check is a deterministic command; a failing check fails the run. The report script in the last section runs\nthe same checks, so run these for your own verification first.\n\n### outputs-exist \u2014 Every required output file exists and is non-empty\n\n```bash\ncd {{results_dir}} && for f in final.svg final.png scores.json report.md; do test -s \"$f\" || { echo \"missing $f\"; exit 1; }; done\n```\n\n### svg-parses \u2014 final.svg is well-formed XML with an <svg> root\n\n```bash\npython3 -c \"import xml.etree.ElementTree as E; r=E.parse('{{results_dir}}/final.svg').getroot(); assert r.tag.split('}')[-1]=='svg', r.tag\"\n```\n\n### svg-size-under-50kb \u2014 final.svg is under 50 KB\n\n```bash\ntest \"$(wc -c < {{results_dir}}/final.svg)\" -lt 51200\n```\n\n### no-image-tags \u2014 final.svg has no <image> elements or embedded rasters\n\n```bash\npython3 -c \"import xml.etree.ElementTree as E; s=open('{{results_dir}}/final.svg').read(); r=E.fromstring(s); assert not [e for e in r.iter() if e.tag.split('}')[-1]=='image'] and 'data:image' not in s and 'base64' not in s\"\n```\n\n### png-rendered \u2014 final.png is a real PNG at least 400 px wide\n\n```bash\npython3 -c \"import struct; b=open('{{results_dir}}/final.png','rb').read(24); assert b[:8]==b'\\x89PNG\\r\\n\\x1a\\n', 'not a PNG'; w=struct.unpack('>I', b[16:20])[0]; assert w>=400, w\"\n```\n\n### rounds-recorded \u2014 Every scored round has its SVG and PNG\n\n```bash\npython3 -c \"import json,os; d=json.load(open('{{results_dir}}/scores.json')); ns=[r['round'] for r in d['rounds']]; assert ns and d['best_round'] in ns; miss=[n for n in ns for e in ('svg','png') if not os.path.getsize('{{results_dir}}/rounds/v%d.%s'%(n,e))]; assert not miss, miss\"\n```\n\n## Checklist\n\nConfirm each item by looking at `final.png`; a failed item fails the run.\n\n- [ ] `pelican-recognizable`: the bird reads as a pelican at a glance (long beak with a visible throat pouch)\n- [ ] `bicycle-complete`: two spoked wheels, a connected frame, handlebars, seat and pedals are all present\n- [ ] `actually-riding`: the pelican sits on the seat with wing tips on the handlebars and feet at the pedals\n- [ ] `scores-match-report`: report.md, scores.json and final.svg agree on the per-round scores and the best round\n\n## Write Validation Report\n\nWrite `{{results_dir}}/validation_report.json` in the evals-v2 shape by running the script below. It runs every code\ncheck above, records one `step` entry per step, one `checklist` entry per item you pass in (status, a colon, then a\none-line reason) and your best-round self-score as a `judge` check (score out of 40, threshold `{{pass_threshold}}`).\nJetty computes the verdict from these checks. Use `fail: <reason>` for anything that did not hold, and set\n`ITERATIONS` to the number of rounds you drew.\n\n```bash\ncd {{results_dir}} && ITERATIONS=3 \\\nSTEPS='{\"setup\": \"pass\", \"draft\": \"pass\", \"refine\": \"pass\", \"finalize\": \"pass\"}' \\\nCHECKLIST='{\"pelican-recognizable\": \"pass: <why>\", \"bicycle-complete\": \"pass: <why>\", \"actually-riding\": \"pass: <why>\", \"scores-match-report\": \"pass: <why>\"}' \\\npython3 - <<'PY'\nimport json, os, struct, xml.etree.ElementTree as E\nR = \"{{results_dir}}\"; T = \"{{pass_threshold}}\"; T = float(T) if T.replace(\".\", \"\", 1).isdigit() else 28.0\nchecks = []\ndef add(kind, cid, name, ok, msg, **kw):\n    checks.append({\"kind\": kind, \"id\": cid, \"name\": name, \"status\": \"pass\" if ok else \"fail\", \"message\": str(msg)[:500], **kw})\ndef run(cid, name, fn):\n    try:\n        add(\"code_check\", cid, name, True, fn() or \"ok\")\n    except Exception as e:\n        add(\"code_check\", cid, name, False, repr(e))\np = lambda f: os.path.join(R, f)\ndef outputs():\n    miss = [f for f in (\"final.svg\", \"final.png\", \"scores.json\", \"report.md\") if not os.path.exists(p(f)) or not os.path.getsize(p(f))]\n    assert not miss, \"missing: \" + \", \".join(miss)\ndef parses():\n    r = E.parse(p(\"final.svg\")).getroot(); assert r.tag.split(\"}\")[-1] == \"svg\", r.tag\ndef size():\n    n = os.path.getsize(p(\"final.svg\")); assert n < 51200, n; return f\"{n} bytes\"\ndef noimg():\n    s = open(p(\"final.svg\")).read(); r = E.fromstring(s)\n    assert not [e for e in r.iter() if e.tag.split(\"}\")[-1] == \"image\"] and \"data:image\" not in s and \"base64\" not in s\ndef png():\n    b = open(p(\"final.png\"), \"rb\").read(24); assert b[:8] == b\"\\x89PNG\\r\\n\\x1a\\n\", \"not a PNG\"\n    w = struct.unpack(\">I\", b[16:20])[0]; assert w >= 400, w; return f\"{w} px wide\"\ndef rounds():\n    d = json.load(open(p(\"scores.json\"))); ns = [r[\"round\"] for r in d[\"rounds\"]]\n    assert ns and d[\"best_round\"] in ns, \"best_round not among rounds\"\n    miss = [f\"v{n}.{e}\" for n in ns for e in (\"svg\", \"png\") if not os.path.exists(p(f\"rounds/v{n}.{e}\"))]\n    assert not miss, miss; return f\"{len(ns)} rounds\"\nrun(\"outputs-exist\", \"Every required output file exists and is non-empty\", outputs)\nrun(\"svg-parses\", \"final.svg is well-formed XML with an <svg> root\", parses)\nrun(\"svg-size-under-50kb\", \"final.svg is under 50 KB\", size)\nrun(\"no-image-tags\", \"final.svg has no <image> elements or embedded rasters\", noimg)\nrun(\"png-rendered\", \"final.png is a real PNG at least 400 px wide\", png)\nrun(\"rounds-recorded\", \"Every scored round has its SVG and PNG\", rounds)\nfor sid, st in json.loads(os.environ[\"STEPS\"]).items():\n    add(\"step\", sid, sid, st.startswith(\"pass\"), st)\nfor cid, st in json.loads(os.environ[\"CHECKLIST\"]).items():\n    ok = st.strip().lower().startswith(\"pass\"); add(\"checklist\", cid, cid, ok, st.split(\":\", 1)[-1].strip())\ntry:\n    d = json.load(open(p(\"scores.json\"))); best = next(r for r in d[\"rounds\"] if r[\"round\"] == d[\"best_round\"])\n    add(\"judge\", \"self-score\", \"Agent self-score of the best round (4 axes x 0-10)\", best[\"total\"] >= T,\n        f\"round {d['best_round']}: \" + \", \".join(f\"{k} {best[k]}\" for k in (\"pelican\", \"bicycle\", \"composition\", \"polish\")),\n        score=float(best[\"total\"]), max_score=40.0, threshold=T)\nexcept Exception as e:\n    checks.append({\"kind\": \"judge\", \"id\": \"self-score\", \"name\": \"Agent self-score of the best round\", \"status\": \"error\", \"message\": repr(e)})\ngating = [c for c in checks if c[\"kind\"] != \"step\"]\npassed = all(c[\"status\"] == \"pass\" for c in gating)\nrep = {\"version\": 2, \"verdict\": \"pass\" if passed else \"fail\", \"overall_passed\": passed,\n       \"iterations\": int(os.environ.get(\"ITERATIONS\", \"1\")), \"checks\": checks}\njson.dump(rep, open(p(\"validation_report.json\"), \"w\"), indent=2)\nfor c in checks: print(f\"{c['status']:5} {c['kind']:10} {c['id']}: {c['message']}\")\nprint(\"verdict:\", rep[\"verdict\"])\nPY\n```\n\nDo not hand-edit the file and do not invent a different format: it must stay\n`{\"version\": 2, \"verdict\", \"overall_passed\", \"iterations\", \"checks\": [...]}`, with each check carrying `kind`\n(`step`, `code_check`, `checklist` or `judge`), `id`, `name`, `status` (`pass`, `fail`, `skipped` or `error`) and\n`message`, plus `score`, `max_score` and `threshold` on the judge.\n\n## Tips\n\n- One round costs roughly $0.50 to $2.50 depending on the model; most of it is the agent's own reasoning.\n- Seeding round 1 with a previous run's `final.svg` (paste it into the runbook) is the single biggest improvement we\n  measured in hand-edited climbs: it skips the cold start.\n- Coordinate-precise fixes (\"close the 13 px gap between the right wing tip and the grip\") move scores more than\n  scene-level additions (sun, motion lines, scenery), which tend to add clutter.\n- The self-score is a weak signal across models: at evaljetty.com/pelicans it correlates only loosely with an external\n  judge. Treat it as a steering aid for your own climb, not as a ranking.\n\n## Changelog\n\n- 2.0.0: Evals v2 shape (Parameters table, Code Checks, Checklist, Write Validation Report with the self-score as a\n  judge check at `{{pass_threshold}}`/40); `scores.json` output; honest-scoring note.\n- 1.x: The runbook used for every run published at evaljetty.com/pelicans (v1 cohort May 2026, v2 cohort Sep 2026):\n  draw, render, self-score and redraw for `max_rounds`, then `final.svg`, `final.png` and `report.md`.\n",
    "model_provider": "openrouter",
    "runbook_evals_version": 2
  },
  "step_configs": {
    "run": {
      "cpus": 4,
      "memory": "8G",
      "activity": "runbook",
      "agent_path": "init_params.agent",
      "files_path": "init_params.file_paths",
      "model_path": "init_params.model",
      "timeout_sec": 1800,
      "snapshot_path": "init_params.snapshot",
      "network_enabled": true,
      "instruction_path": "init_params.instruction",
      "model_provider_path": "init_params.model_provider",
      "template_variables_path": "init_params.vars"
    }
  }
}