evals
Create your eval

01 — Author

Tell Jetty
the job to be done.

One runbook plus a sandbox snapshot. Jetty scores every run against the runbook's checks and the task goal.

The job

Micropolis is the city simulator Electronic Arts released as open source (GPL v3) in 2008: the original SimCity engine, ported to C++. Each run gives one model the mayor's chair on one task and measures the city it leaves behind. The simulation is turn-based: every turn the model reads a text briefing and replies with actions, then the engine simulates 3 game months. Model latency never costs game time; it only costs wall-clock time, and each run gives the model a 45-minute budget, after which the city runs on untouched to the deadline so every outcome is measured at the same game date.

Tasks and goals

TaskNameLengthGoal (the judge)
growGrow a city20 yrs · 80 turnsPopulation at the deadline (bankrupt = 0)
tokyoTokyo 1957: monster attack5 yrs · 20 turnsScenario won (city score > 500 in 1962)
sanfranciscoSan Francisco 1906: earthquake5 yrs · 20 turnsScenario won (Metropolis, 100,000 people, in 1911)
detroitDetroit 1972: crime10 yrs · 40 turnsScenario won (crime < 60 in 1982, population >= 25,000)

Grow starts on an empty generated 120×100 map with $20,000 (seed = map seed); bankruptcy (treasury under $500 for 8 turns in a row) zeroes the score. The scenarios are Micropolis's own, with the engine's own win rules; Detroit adds a population floor of 25,000 because an idle mayor "wins" it when the city empties out and crime falls with it. Same rules everywhere: no random disasters (scripted scenario disasters still happen), manual budget, trees and rubble auto-cleared when building, easy difficulty.

What the model sees and can do

Each turn's briefing has the date and deadline, funds and last January's budget, population by zone type, residential / commercial / industrial demand, the city score and the citizens' top complaints, crime, pollution, traffic and land value, every power plant and service building with coordinates, unpowered zones, free 3×3 sites next to a road, large open areas, hot spots, a 60×50 text map (one character per 2×2 tiles), last turn's events and action results, and the model's own note from last turn. It replies with one JSON object: its reasoning, a note for next turn, and up to 40 actions.

ACTIONS (reply with up to 40 per turn, applied in order; coordinates are tiles, x = column 0-119 west->east, y = row 0-99 north->south):
- {"type": "zone", "zone": "residential"|"commercial"|"industrial", "x": X, "y": Y}  3x3 zone, (X,Y) = top-left tile, $100.
    Optional "count": 2-8 and "dir": "east"|"south" places a row of adjacent zones (next one at X+3 or Y+3).
- {"type": "build", "what": B, "x": X, "y": Y}  (X,Y) = top-left tile. B is one of:
    "coal_power" 4x4 $3000 | "nuclear_power" 4x4 $5000 | "police" 3x3 $500 | "fire_station" 3x3 $500 |
    "park" 1x1 $10 | "stadium" 4x4 $5000 | "seaport" 4x4 $3000 (must touch water) | "airport" 6x6 $10000
- {"type": "road"|"rail"|"power_line", "x1": X1, "y1": Y1, "x2": X2, "y2": Y2}  a line of tiles from (X1,Y1) to
    (X2,Y2); if not straight it goes horizontally first, then vertically (an L). Road $10/tile, rail $20, power line $5;
    crossing water costs more (bridges). Max 60 tiles per line.
- {"type": "bulldoze", "x": X, "y": Y, "w": 1-12, "h": 1-12}  clear a rectangle (rubble, trees, roads, buildings). $1/tile+.
- {"type": "tax", "rate": 0-20}  city tax rate in percent (applies from now on; collected every January).
- {"type": "funding", "road": 0-100, "fire": 0-100, "police": 0-100}  percent of each department's requested budget.
- {"type": "zoom", "x": X, "y": Y}  next turn, also show a full-resolution 40x25 view with top-left (X,Y).

Every action goes through the engine's own tools (the same code paths as clicking in the game), so costs and placement rules are Micropolis's. Rejected actions come back next turn with the reason (for example "blocked by road at (51,83)"). The full system prompt ships in the source download.

How a run is checked (evals v2)

Code checks decide whether a run is a valid measurement; a judge check decides whether the model met the task goal; a reviewer checklist covers what code cannot. city-report writes them into validation_report.json and Jetty computes the verdict: a run passes only if it is valid and the goal was met.

Code checkWhat it verifies
outputs-existEvery required output file exists and is non-empty ```bash cd {{results_dir}} && for f in result.json metrics.json decisions.jsonl series.json timelapse.mp4 thumb.jpg final_map.png city.cty runtime_config.json summary.md; do test -s "$f" || { echo "missing $f"; exit 1; }; done ``` ### sim-reached-deadline — The simulation ran to the task's deadline (or the scenario was judged) ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); assert r['reached_deadline'], r['final']['date']; print(r['final']['date'], r['turns'])" ``` ### decisions-logged — decisions.jsonl has one row per turn ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); n=sum(1 for _ in open('{{results_dir}}/decisions.jsonl')); assert n and n==r['turns'], (n, r['turns'])" ``` ### llm-turns-valid — At most 20% of the model's turns failed ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); c,f=r['llm_calls'],r['llm_failures']; assert c and f<=max(2,0.2*c), (f,c)" ``` ### video-decodes — timelapse.mp4 decodes ```bash test "$(ffprobe -v error -show_entries format=duration -of csv=p=0 {{results_dir}}/timelapse.mp4 | cut -d. -f1)" -gt 0 ``` ### limits-respected — The run finished within the wall-clock budget ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); assert r['wall_seconds']/60 <= {{wall_minutes}}+10, r['wall_seconds']" ``` ## Checklist Confirm each item by inspection; a failed item fails the run. - [ ] `notes-explain-outcome`: summary.md has a Notes section explaining the outcome, citing the model's own thoughts - [ ] `player-matches-request`: the player, model, task and seed in result.json are the ones requested in `{{player}}`, `{{task}}` and `{{seed}}` - [ ] `no-silent-outage`: run.log shows no long stretch of failed or empty model turns that the code checks did not already flag ## Write Validation Report Write `{{results_dir}}/validation_report.json` in the evals-v2 shape with `city-report`. It runs every code check above, adds one `judge` check for the task goal (from `result.json`: for `grow` the score is the final population, threshold 10,000, and a bankrupt city scores 0; for a scenario the score is 1 if the scenario was won and 0 if lost, threshold 1), records one `step` entry per step you pass in, and one `checklist` entry per item you pass in (status, a colon, then a one-line reason). Jetty computes the verdict from these checks: the run passes only if it is valid (every code check and checklist item passes) AND the model met the goal (judge score >= threshold).
Checklist itemReviewer confirms
notes-explain-outcomesummary.md has a Notes section explaining the outcome, citing the model's own thoughts
player-matches-requestthe player, model, task and seed in result.json are the ones requested in `{{player}}`, `{{task}}` and `{{seed}}`
no-silent-outagerun.log shows no long stretch of failed or empty model turns that the code checks did not already flag
JudgeScoreThreshold
city-goalGrow: population in January 1920, 0 if bankrupt10,000
scenario-goalScenario: 1 if won (engine verdict, plus Detroit's population floor), else 01

The environment: sandbox snapshot micropolis-city

4 vCPU / 8 GB. Debian bookworm with the Micropolis engine (SimHacker/micropolis@c98f6b0) compiled as a Python extension through a small pybind11 binding, the game's scenario files and tiles, Python 3.11, ffmpeg, Node and Claude Code for the runbook agent. A 20-year city simulates in well under a second; almost all of a run's wall clock is the model thinking. The timelapse is rendered from the engine's own 16-pixel tile set, one frame per game week.

Players

A preset ID, openrouter:<model>|<Display name> for any OpenRouter chat model, or idle for a baseline mayor that never acts.

SpecNameModelReasoning
jevJev Routertypesafe/jev-routernone (router)
sonnet5Claude Sonnet 5anthropic/claude-sonnet-5effort: low
gpt56terraGPT-5.6 Terraopenai/gpt-5.6-terraeffort: low
gpt56lunaGPT-5.6 Lunaopenai/gpt-5.6-lunaeffort: low

The runbook

A runbook is one Markdown file that gives an agent its job, its bar for done and its checks. Jetty runs it in a sandbox and computes a verdict from the checks. This one has these parts:

  1. Objective: play one city task with the requested model as mayor and deliver a timelapse, the final map, a decision log and a results file.
  2. Parameters: the player, the task and seed, a run ID and the model's wall-clock budget.
  3. Steps: check the sandbox, launch the city, supervise it in the foreground until the deadline, render the timelapse and write short notes.
  4. Code checks: deterministic commands that verify the outputs, listed above.
  5. Judge: the task goal (population, city score or crime) read from the engine at the deadline.
  6. Checklist and validation report: items the agent confirms, then validation_report.json in Jetty's evals-v2 shape, from which Jetty computes the verdict.

Learn more about runbooks in the Jetty docs

Improve: what changed

Lessons from the other evals are built in from the start: supervision is foreground-only (a runbook agent that moves polling into a background task ends its session and loses the sandbox), results are written before the video is finalized, and a crashed run gets one retry. Pilot runs before round 1 led to two harness changes: rejected building placements now name the tile that blocks them, and the briefing says that the city score and census are only re-evaluated each January.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →