evals
Create your eval

03 — Investigation · Compare setups

Which model plans a city best?

8 runsRound 13 game months per turnrunbook v1.0.0collection jettygrowthteamtask micropolis-cityupdated 2026-09-28

Decision it informs: Which model to trust with a long-horizon planning task where early mistakes compound for years.

Key takeaway generated

Jev Router built the largest cities (mean 12,120 people after 20 years), ahead of GPT-5.6 Luna (7,210), GPT-5.6 Terra (3,290) and Claude Sonnet 5 (1,070). With two maps per model, the order is a signal, not a verdict.

At a glance computed

Runs
8
one city per Jetty run
Model spend
$10.00
OpenRouter, the mayor models
Goals met
0/8
judge score ≥ threshold
Eval pass rate
0%
0/8 runs valid and goal met

Key figure computed

Population in January 1920, every grow run

One bar per run, colored by model; the dashed line is the 10,000 goal. Hover for the city score.

Table view
RunTaskSeedMayorOutcomePopulationCity scoreEvaluationModel costWall minJetty run ID
jev-grow-s2Grow a city2Jev Router19,660 people (bankrupt)19,660596fail$2.1446.35de24fe4
gpt56luna-grow-s2Grow a city2GPT-5.6 Luna8,060 people8,060417fail$0.148.4831ede26
gpt56luna-grow-s6Grow a city6GPT-5.6 Luna6,360 people6,360311fail$0.159.7349eb345
jev-grow-s6Grow a city6Jev Router4,580 people4,580367fail$1.9146.6981094de
gpt56terra-grow-s2Grow a city2GPT-5.6 Terra3,660 people3,660322fail$1.5413.0476d94f4
gpt56terra-grow-s6Grow a city6GPT-5.6 Terra2,920 people (bankrupt)2,920756fail$1.4112.64c859227
sonnet5-grow-s2Grow a city2Claude Sonnet 51,200 people1,200431fail$1.3914.240c8cf3e
sonnet5-grow-s6Grow a city6Claude Sonnet 5940 people940514fail$1.3213.10e6f2bec

Evaluations computed

Every run's evals-v2 validation report, scored by Jetty: code checks (is the run valid), the judge (was the goal met) and the reviewer checklist.

EvaluationPassedRateShare
Code checks (run is valid)48/48100%
Checklist items24/24100%
Judge: task goal met0/80%
Valid runs (all code checks and checklist items)8/8100%
Runs with a passing verdict (valid and goal met)0/80%

Findings generated

Jev Router builds the biggest cities

medium confidence

Mean population in January 1920: Jev Router 12,120 (best 19,660); GPT-5.6 Luna 7,210 (best 8,060); GPT-5.6 Terra 3,290 (best 3,660); Claude Sonnet 5 1,070 (best 1,200). 0 of 8 grow runs met the goal (10,000 people in 1920 without going bankrupt).

jev-grow-s2jev-grow-s68 grow runs

The best single city: 19,660 people (Jev Router, seed 2)

high confidence

It peaked at 25,140 with city score 596 and $2,545 left, but it still missed the goal: the treasury sat under $500 for two years first, which counts as bankruptcy. The weakest model, Claude Sonnet 5, averaged 1,070.

jev-grow-s2score 596

7 of 8 cities peaked and then shrank by 40% or more

medium confidence

The usual pattern: early overbuilding drains the treasury, the model cannot afford the next power plant, raises taxes to rebuild cash, and demand collapses. Claude Sonnet 5 s2: peak 14,880 → 1,200; GPT-5.6 Terra s2: peak 20,900 → 3,660; GPT-5.6 Luna s2: peak 25,820 → 8,060; Jev Router s6: peak 24,600 → 4,580.

sonnet5-grow-s2gpt56terra-grow-s2gpt56luna-grow-s2jev-grow-s6

Placement accuracy differs a lot

high confidence

Share of actions the engine rejected (blocked site, no money, off-map), all tasks: Jev Router 4% of 3.5 actions/turn; GPT-5.6 Terra 7% of 4.7 actions/turn; GPT-5.6 Luna 17% of 6.5 actions/turn; Claude Sonnet 5 20% of 4.6 actions/turn.

actions_ok / actions_rejected in runs.csv

Proposed improvements generated

Run

Add seeds

Two maps per model separate large gaps but not close ones; four to six seeds would.

backed by: 8 grow runs

Runbook

Give the mayor a budget forecast

Most collapses start with an unaffordable power plant; a projected January budget in the briefing would test planning, not arithmetic.

backed by: 7 collapsed cities

Eval

Judge the plan, not just the result

An LLM judge over decisions.jsonl could score zoning ratios, power planning and budget discipline per turn.

backed by: decision logs for every run

Caveats

  • Two generated maps (seeds 2 and 6), one run each per model.
  • Routers such as Jev Router pick a model per request; the model used each turn is in the decision log.
  • Each run gives the model 45 wall-clock minutes; slower models may hand the last years to autopilot (flagged per run).
  • Population is Micropolis's own census; the city score is a tie-breaker, not part of the goal.
Runs and provenance
RunTaskSeedMayorOutcomePopulationCity scoreEvaluationModel costWall minJetty run ID
jev-grow-s2Grow a city2Jev Router19,660 people (bankrupt)19,660596fail$2.1446.35de24fe4
sonnet5-grow-s2Grow a city2Claude Sonnet 51,200 people1,200431fail$1.3914.240c8cf3e
gpt56terra-grow-s2Grow a city2GPT-5.6 Terra3,660 people3,660322fail$1.5413.0476d94f4
gpt56luna-grow-s2Grow a city2GPT-5.6 Luna8,060 people8,060417fail$0.148.4831ede26
jev-grow-s6Grow a city6Jev Router4,580 people4,580367fail$1.9146.6981094de
sonnet5-grow-s6Grow a city6Claude Sonnet 5940 people940514fail$1.3213.10e6f2bec
gpt56terra-grow-s6Grow a city6GPT-5.6 Terra2,920 people (bankrupt)2,920756fail$1.4112.64c859227
gpt56luna-grow-s6Grow a city6GPT-5.6 Luna6,360 people6,360311fail$0.159.7349eb345

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →