---
title: "Author: the Micropolis runbook, tasks and checks | Evals by Jetty"
url: https://evaljetty.com/micropolis/author.html
description: "The Jetty runbook behind the Micropolis LLM city-builder eval: tasks and goals, what the model sees, the action set, evals-v2 code checks, judge and checklist, the sandbox snapshot."
updated: 2026-09-28
publisher: Jetty (https://jetty.io)
---

01 — Author

# Tell Jetty the job to be done.

One runbook plus a sandbox snapshot. Jetty scores every run against the runbook's checks and the task goal.

[Snapshot Dockerfile](https://evaljetty.com/micropolis/downloads/Dockerfile)
[Source](https://evaljetty.com/micropolis/downloads/micropolis-eval-src.tar.gz)

## The job

[Micropolis](https://github.com/SimHacker/micropolis) is the city simulator Electronic Arts released as open
source (GPL v3) in 2008: the original SimCity engine, ported to C++. Each run gives one model the mayor's chair on one task
and measures the city it leaves behind. The simulation is **turn-based**: every turn the model reads a text
briefing and replies with actions, then the engine simulates 3 game months. Model latency never
costs game time; it only costs wall-clock time, and each run gives the model a 45-minute budget, after which the city runs on
untouched to the deadline so every outcome is measured at the same game date.

## Tasks and goals

| Task | Name | Length | Goal (the judge) |
| --- | --- | --- | --- |
| grow | Grow a city | 20 yrs · 80 turns | Population at the deadline (bankrupt = 0) |
| tokyo | Tokyo 1957: monster attack | 5 yrs · 20 turns | Scenario won (city score > 500 in 1962) |
| sanfrancisco | San Francisco 1906: earthquake | 5 yrs · 20 turns | Scenario won (Metropolis, 100,000 people, in 1911) |
| detroit | Detroit 1972: crime | 10 yrs · 40 turns | Scenario won (crime < 60 in 1982, population >= 25,000) |

Grow starts on an empty generated 120×100 map with $20,000 (seed = map seed); bankruptcy (treasury under
$500 for 8 turns in a row) zeroes the score. The scenarios are Micropolis's own, with the engine's
own win rules; Detroit adds a population floor of 25,000 because an idle mayor "wins" it when the
city empties out and crime falls with it. Same rules everywhere: no random disasters (scripted scenario disasters still
happen), manual budget, trees and rubble auto-cleared when building, easy difficulty.

## What the model sees and can do

Each turn's briefing has the date and deadline, funds and last January's budget, population by zone type, residential /
commercial / industrial demand, the city score and the citizens' top complaints, crime, pollution, traffic and land value,
every power plant and service building with coordinates, unpowered zones, free 3×3 sites next to a road, large open areas,
hot spots, a 60×50 text map (one character per 2×2 tiles), last turn's events and action results, and the model's own note
from last turn. It replies with one JSON object: its reasoning, a note for next turn, and up to 40 actions.

```
ACTIONS (reply with up to 40 per turn, applied in order; coordinates are tiles, x = column 0-119 west->east, y = row 0-99 north->south):
- {"type": "zone", "zone": "residential"|"commercial"|"industrial", "x": X, "y": Y}  3x3 zone, (X,Y) = top-left tile, $100.
    Optional "count": 2-8 and "dir": "east"|"south" places a row of adjacent zones (next one at X+3 or Y+3).
- {"type": "build", "what": B, "x": X, "y": Y}  (X,Y) = top-left tile. B is one of:
    "coal_power" 4x4 $3000 | "nuclear_power" 4x4 $5000 | "police" 3x3 $500 | "fire_station" 3x3 $500 |
    "park" 1x1 $10 | "stadium" 4x4 $5000 | "seaport" 4x4 $3000 (must touch water) | "airport" 6x6 $10000
- {"type": "road"|"rail"|"power_line", "x1": X1, "y1": Y1, "x2": X2, "y2": Y2}  a line of tiles from (X1,Y1) to
    (X2,Y2); if not straight it goes horizontally first, then vertically (an L). Road $10/tile, rail $20, power line $5;
    crossing water costs more (bridges). Max 60 tiles per line.
- {"type": "bulldoze", "x": X, "y": Y, "w": 1-12, "h": 1-12}  clear a rectangle (rubble, trees, roads, buildings). $1/tile+.
- {"type": "tax", "rate": 0-20}  city tax rate in percent (applies from now on; collected every January).
- {"type": "funding", "road": 0-100, "fire": 0-100, "police": 0-100}  percent of each department's requested budget.
- {"type": "zoom", "x": X, "y": Y}  next turn, also show a full-resolution 40x25 view with top-left (X,Y).
```

Every action goes through the engine's own tools (the same code paths as clicking in the game), so
costs and placement rules are Micropolis's. Rejected actions come back next turn with the reason (for example "blocked by
road at (51,83)"). The full system prompt ships in the source download.

## How a run is checked (evals v2)

Code checks decide whether a run is a valid measurement; a judge check decides whether the model met the
task goal; a reviewer checklist covers what code cannot. `city-report` writes them into
`validation_report.json` and Jetty computes the verdict: a run passes only if it is valid *and* the goal was met.

| Code check | What it verifies |
| --- | --- |
| outputs-exist | Every required output file exists and is non-empty ```bash cd {{results\_dir}} && for f in result.json metrics.json decisions.jsonl series.json timelapse.mp4 thumb.jpg final\_map.png city.cty runtime\_config.json summary.md; do test -s "$f" || { echo "missing $f"; exit 1; }; done ``` ### sim-reached-deadline — The simulation ran to the task's deadline (or the scenario was judged) ```bash python3 -c "import json; r=json.load(open('{{results\_dir}}/result.json')); assert r['reached\_deadline'], r['final']['date']; print(r['final']['date'], r['turns'])" ``` ### decisions-logged — decisions.jsonl has one row per turn ```bash python3 -c "import json; r=json.load(open('{{results\_dir}}/result.json')); n=sum(1 for \_ in open('{{results\_dir}}/decisions.jsonl')); assert n and n==r['turns'], (n, r['turns'])" ``` ### llm-turns-valid — At most 20% of the model's turns failed ```bash python3 -c "import json; r=json.load(open('{{results\_dir}}/result.json')); c,f=r['llm\_calls'],r['llm\_failures']; assert c and f<=max(2,0.2\*c), (f,c)" ``` ### video-decodes — timelapse.mp4 decodes ```bash test "$(ffprobe -v error -show\_entries format=duration -of csv=p=0 {{results\_dir}}/timelapse.mp4 | cut -d. -f1)" -gt 0 ``` ### limits-respected — The run finished within the wall-clock budget ```bash python3 -c "import json; r=json.load(open('{{results\_dir}}/result.json')); assert r['wall\_seconds']/60 <= {{wall\_minutes}}+10, r['wall\_seconds']" ``` ## Checklist Confirm each item by inspection; a failed item fails the run. - [ ] `notes-explain-outcome`: summary.md has a Notes section explaining the outcome, citing the model's own thoughts - [ ] `player-matches-request`: the player, model, task and seed in result.json are the ones requested in `{{player}}`, `{{task}}` and `{{seed}}` - [ ] `no-silent-outage`: run.log shows no long stretch of failed or empty model turns that the code checks did not already flag ## Write Validation Report Write `{{results\_dir}}/validation\_report.json` in the evals-v2 shape with `city-report`. It runs every code check above, adds one `judge` check for the task goal (from `result.json`: for `grow` the score is the final population, threshold 10,000, and a bankrupt city scores 0; for a scenario the score is 1 if the scenario was won and 0 if lost, threshold 1), records one `step` entry per step you pass in, and one `checklist` entry per item you pass in (status, a colon, then a one-line reason). Jetty computes the verdict from these checks: the run passes only if it is valid (every code check and checklist item passes) AND the model met the goal (judge score >= threshold). |

| Checklist item | Reviewer confirms |
| --- | --- |
| notes-explain-outcome | summary.md has a Notes section explaining the outcome, citing the model's own thoughts |
| player-matches-request | the player, model, task and seed in result.json are the ones requested in `{{player}}`, `{{task}}` and `{{seed}}` |
| no-silent-outage | run.log shows no long stretch of failed or empty model turns that the code checks did not already flag |

| Judge | Score | Threshold |
| --- | --- | --- |
| city-goal | Grow: population in January 1920, 0 if bankrupt | 10,000 |
| scenario-goal | Scenario: 1 if won (engine verdict, plus Detroit's population floor), else 0 | 1 |

## The environment: sandbox snapshot `micropolis-city`

4 vCPU / 8 GB. Debian bookworm with the Micropolis engine (`SimHacker/micropolis@c98f6b0`) compiled as a Python
extension through a small pybind11 binding, the game's scenario files and tiles, Python 3.11, ffmpeg, Node and Claude Code for
the runbook agent. A 20-year city simulates in well under a second; almost all of a run's wall clock is the model thinking.
The timelapse is rendered from the engine's own 16-pixel tile set, one frame per game week.

## Players

A preset ID, `openrouter:<model>|<Display name>` for any OpenRouter chat model, or
`idle` for a baseline mayor that never acts.

| Spec | Name | Model | Reasoning |
| --- | --- | --- | --- |
| jev | Jev Router | typesafe/jev-router | none (router) |
| sonnet5 | Claude Sonnet 5 | anthropic/claude-sonnet-5 | effort: low |
| gpt56terra | GPT-5.6 Terra | openai/gpt-5.6-terra | effort: low |
| gpt56luna | GPT-5.6 Luna | openai/gpt-5.6-luna | effort: low |

## The runbook

A runbook is one Markdown file that gives an agent its job, its bar for done and its checks. Jetty runs it in a sandbox and computes a verdict from the checks. This one has these parts:

1. **Objective:** play one city task with the requested model as mayor and deliver a timelapse, the final map, a decision log and a results file.
2. **Parameters:** the player, the task and seed, a run ID and the model's wall-clock budget.
3. **Steps:** check the sandbox, launch the city, supervise it in the foreground until the deadline, render the timelapse and write short notes.
4. **Code checks:** deterministic commands that verify the outputs, listed above.
5. **Judge:** the task goal (population, city score or crime) read from the engine at the deadline.
6. **Checklist and validation report:** items the agent confirms, then validation\_report.json in Jetty's evals-v2 shape, from which Jetty computes the verdict.

[Learn more about runbooks in the Jetty docs](https://jetty.io/docs/ai-workloads/runbooks?utm_source=evaljetty&utm_medium=referral&utm_campaign=micropolis)

## Improve: what changed

Lessons from the other evals are built in from the start: supervision is foreground-only (a runbook agent that moves
polling into a background task ends its session and loses the sandbox), results are written before the video is finalized,
and a crashed run gets one retry. Pilot runs before round 1 led to two harness changes: rejected building placements now name
the tile that blocks them, and the briefing says that the city score and census are only re-evaluated each January.

---
Source: https://evaljetty.com/micropolis/author.html · Evals by Jetty · Run your own evals: https://jetty.io/?utm_source=evaljetty&utm_medium=llms
