evals
Create your eval

03 — Investigation · Break down by case

Which scenarios break which models?

24 runsRound 13 game months per turnrunbook v1.0.0collection jettygrowthteamtask micropolis-cityupdated 2026-09-28

Decision it informs: Whether a model can recover a system in crisis (a disaster, a failing city), not just grow one from scratch.

Key takeaway generated

Scenario wins: Tokyo 1957 1/8; San Francisco 1906 0/8; Detroit 1972 0/8. Best at recovery: GPT-5.6 Terra (1/6).

At a glance computed

Runs
24
one city per Jetty run
Model spend
$18.08
OpenRouter, the mayor models
Goals met
1/24
judge score ≥ threshold
Eval pass rate
4%
1/24 runs valid and goal met

Key figure computed

Outcome by model and task

✓ goal met, ✗ missed, one mark per seed; hover a mark for the judge's detail. The last row is an idle mayor that never acts.

MayorGrow a cityTokyo 1957San Francisco 1906Detroit 1972Goals met
GPT-5.6 Terra✗ 3.7k · ✗ 2.9k✓ won · ✗ lost✗ lost · ✗ lost✗ lost · ✗ lost1/8
Jev Router✗ 19.7k · ✗ 4.6k✗ lost · ✗ lost✗ lost · ✗ lost✗ lost · ✗ lost0/8
GPT-5.6 Luna✗ 8.1k · ✗ 6.4k✗ lost · ✗ lost✗ lost · ✗ lost✗ lost · ✗ lost0/8
Claude Sonnet 5✗ 1.2k · ✗ 0.9k✗ lost · ✗ lost✗ lost · ✗ lost✗ lost · ✗ lost0/8
Idle mayor (baseline: no actions)✗ 0.0k · ✗ 0.0k✗ lost · ✗ lost✗ lost · ✗ lost✗ lost · ✗ lost0/8
Table view
RunMayorResultJudge detail
jev-tokyo-s1Jev RouterlostLOST: City score 372 (goal > 500); engine verdict: lost
sonnet5-tokyo-s1Claude Sonnet 5lostLOST: City score 333 (goal > 500); engine verdict: lost
gpt56terra-tokyo-s1GPT-5.6 TerrawonWON: City score 501 (goal > 500); engine verdict: won
gpt56luna-tokyo-s1GPT-5.6 LunalostLOST: City score 330 (goal > 500); engine verdict: lost
jev-tokyo-s2Jev RouterlostLOST: City score 368 (goal > 500); engine verdict: lost
sonnet5-tokyo-s2Claude Sonnet 5lostLOST: City score 346 (goal > 500); engine verdict: lost
gpt56terra-tokyo-s2GPT-5.6 TerralostLOST: City score 237 (goal > 500); engine verdict: lost
gpt56luna-tokyo-s2GPT-5.6 LunalostLOST: City score 458 (goal > 500); engine verdict: lost
jev-sanfrancisco-s1Jev RouterlostLOST: Population 36,300 (goal >= 100,000); engine verdict: lost
sonnet5-sanfrancisco-s1Claude Sonnet 5lostLOST: Population 36,840 (goal >= 100,000); engine verdict: lost
gpt56terra-sanfrancisco-s1GPT-5.6 TerralostLOST: Population 47,680 (goal >= 100,000); engine verdict: lost
gpt56luna-sanfrancisco-s1GPT-5.6 LunalostLOST: Population 23,680 (goal >= 100,000); engine verdict: lost
jev-sanfrancisco-s2Jev RouterlostLOST: Population 92,360 (goal >= 100,000); engine verdict: lost
sonnet5-sanfrancisco-s2Claude Sonnet 5lostLOST: Population 72,300 (goal >= 100,000); engine verdict: lost
gpt56terra-sanfrancisco-s2GPT-5.6 TerralostLOST: Population 38,880 (goal >= 100,000); engine verdict: lost
gpt56luna-sanfrancisco-s2GPT-5.6 LunalostLOST: Population 48,220 (goal >= 100,000); engine verdict: lost
jev-detroit-s1Jev RouterlostLOST: Average crime 22 (goal < 60), population 7,380 (floor 25,000); engine verdict: won
sonnet5-detroit-s1Claude Sonnet 5lostLOST: Average crime 30 (goal < 60), population 3,520 (floor 25,000); engine verdict: won
gpt56terra-detroit-s1GPT-5.6 TerralostLOST: Average crime 33 (goal < 60), population 6,120 (floor 25,000); engine verdict: won
gpt56luna-detroit-s1GPT-5.6 LunalostLOST: Average crime 24 (goal < 60), population 7,400 (floor 25,000); engine verdict: won
jev-detroit-s2Jev RouterlostLOST: Average crime 40 (goal < 60), population 10,040 (floor 25,000); engine verdict: won
sonnet5-detroit-s2Claude Sonnet 5lostLOST: Average crime 28 (goal < 60), population 3,280 (floor 25,000); engine verdict: won
gpt56terra-detroit-s2GPT-5.6 TerralostLOST: Average crime 56 (goal < 60), population 6,880 (floor 25,000); engine verdict: won
gpt56luna-detroit-s2GPT-5.6 LunalostLOST: Average crime 30 (goal < 60), population 15,280 (floor 25,000); engine verdict: won

Evaluations computed

Every run's evals-v2 validation report, scored by Jetty: code checks (is the run valid), the judge (was the goal met) and the reviewer checklist.

EvaluationPassedRateShare
Code checks (run is valid)144/144100%
Checklist items72/72100%
Judge: task goal met1/244%
Valid runs (all code checks and checklist items)24/24100%
Runs with a passing verdict (valid and goal met)1/244%

Findings generated

Tokyo 1957: monster attack: 1 of 8 runs won

medium confidence

Goal: Scenario won (city score > 500 in 1962). Jev Router s1 lost (city score 372); Claude Sonnet 5 s1 lost (city score 333); GPT-5.6 Terra s1 won (city score 501); GPT-5.6 Luna s1 lost (city score 330); Jev Router s2 lost (city score 368); Claude Sonnet 5 s2 lost (city score 346); GPT-5.6 Terra s2 lost (city score 237); GPT-5.6 Luna s2 lost (city score 458).

gpt56terra-tokyo-s1

San Francisco 1906: earthquake: 0 of 8 runs won

medium confidence

Goal: Scenario won (Metropolis, 100,000 people, in 1911). Jev Router s1 lost (population 36,300); Claude Sonnet 5 s1 lost (population 36,840); GPT-5.6 Terra s1 lost (population 47,680); GPT-5.6 Luna s1 lost (population 23,680); Jev Router s2 lost (population 92,360); Claude Sonnet 5 s2 lost (population 72,300); GPT-5.6 Terra s2 lost (population 38,880); GPT-5.6 Luna s2 lost (population 48,220).

jev-sanfrancisco-s1sonnet5-sanfrancisco-s1gpt56terra-sanfrancisco-s1gpt56luna-sanfrancisco-s1

Detroit 1972: crime: 0 of 8 runs won

medium confidence

Goal: Scenario won (crime < 60 in 1982, population >= 25,000). Jev Router s1 lost (average crime 22, pop 7,380); Claude Sonnet 5 s1 lost (average crime 30, pop 3,520); GPT-5.6 Terra s1 lost (average crime 33, pop 6,120); GPT-5.6 Luna s1 lost (average crime 24, pop 7,400); Jev Router s2 lost (average crime 40, pop 10,040); Claude Sonnet 5 s2 lost (average crime 28, pop 3,280); GPT-5.6 Terra s2 lost (average crime 56, pop 6,880); GPT-5.6 Luna s2 lost (average crime 30, pop 15,280).

jev-detroit-s1sonnet5-detroit-s1gpt56terra-detroit-s1gpt56luna-detroit-s1

Doing nothing loses every scenario

high confidence

An idle mayor (same harness and seeds, no actions) as a reference: Detroit 1972 s1 LOST: Average crime 56 (goal < 60), population 9,320 (floor 25,000); Detroit 1972 s2 LOST: Average crime 46 (goal < 60), population 2,020 (floor 25,000); San Francisco 1906 s1 LOST: Population 32,240 (goal >= 100,000); San Francisco 1906 s2 LOST: Population 40,900 (goal >= 100,000); Tokyo 1957 s1 LOST: City score 360 (goal > 500); Tokyo 1957 s2 LOST: City score 315 (goal > 500). These cities decline on their own, so a model has to act to win.

runs/micropolis-baseline (local, no model)

GPT-5.6 Terra handles disasters best

medium confidence

Scenario wins by model: GPT-5.6 Terra 1/6; Jev Router 0/6; GPT-5.6 Luna 0/6; Claude Sonnet 5 0/6.

24 scenario runs

Proposed improvements generated

Runbook

Show the scenario rule's trend

Scenarios are judged once, at the deadline; a per-year projection of the judged metric would reward steering over reacting.

backed by: judge messages per run

Run

Run more seeds per scenario

One or two repeats per model and scenario cannot separate a lucky win from a skill; four to six would.

backed by: runs per scenario in the matrix

Eval

Score partial progress

A 0/1 judge hides near misses; recording the judged metric's distance to the goal would separate close losses from collapses.

backed by: scenario-goal judge

Caveats

  • Repeats per model: Tokyo 1957 2 (seeds 1, 2); San Francisco 1906 2 (seeds 1, 2); Detroit 1972 2 (seeds 1, 2). The seed moves where disasters land.
  • Win rules are the engine's own; Detroit adds this eval's population floor (25,000).
  • Scenario cities decline without intervention in this engine, so holding the line is part of the job.
Runs and provenance
RunTaskSeedMayorOutcomePopulationCity scoreEvaluationModel costWall minJetty run ID
jev-tokyo-s1Tokyo 19571Jev Routerlost23,120372fail$1.2114.496f24e6b
sonnet5-tokyo-s1Tokyo 19571Claude Sonnet 5lost23,800333fail$0.393.458cccc15
gpt56terra-tokyo-s1Tokyo 19571GPT-5.6 Terrawon29,640501pass$0.616.000ff7869
gpt56luna-tokyo-s1Tokyo 19571GPT-5.6 Lunalost23,700330fail$0.042.54675adce
jev-tokyo-s2Tokyo 19572Jev Routerlost21,900368fail$0.6510.1d78fe456
sonnet5-tokyo-s2Tokyo 19572Claude Sonnet 5lost22,400346fail$0.383.172ed89e1
gpt56terra-tokyo-s2Tokyo 19572GPT-5.6 Terralost31,020237fail$0.555.42cf5d9c4
gpt56luna-tokyo-s2Tokyo 19572GPT-5.6 Lunalost29,120458fail$0.042.5a57b0db6
jev-sanfrancisco-s1San Francisco 19061Jev Routerlost36,300286fail$0.8610.7b6be516e
sonnet5-sanfrancisco-s1San Francisco 19061Claude Sonnet 5lost36,84023fail$0.423.7df629564
gpt56terra-sanfrancisco-s1San Francisco 19061GPT-5.6 Terralost47,68039fail$0.656.4d5eb8764
gpt56luna-sanfrancisco-s1San Francisco 19061GPT-5.6 Lunalost23,68021fail$0.052.67a99b8d9
jev-sanfrancisco-s2San Francisco 19062Jev Routerlost92,360500fail$2.0018.4d76eb040
sonnet5-sanfrancisco-s2San Francisco 19062Claude Sonnet 5lost72,30027fail$0.454.3fd728217
gpt56terra-sanfrancisco-s2San Francisco 19062GPT-5.6 Terralost38,88037fail$0.686.4b8e5bc93
gpt56luna-sanfrancisco-s2San Francisco 19062GPT-5.6 Lunalost48,22022fail$0.042.5cd12e212
jev-detroit-s1Detroit 19721Jev Routerlost7,380595fail$2.2126.2aea16d25
sonnet5-detroit-s1Detroit 19721Claude Sonnet 5lost3,520586fail$0.736.9650b3bc0
gpt56terra-detroit-s1Detroit 19721GPT-5.6 Terralost6,120597fail$1.069.8175f2d53
gpt56luna-detroit-s1Detroit 19721GPT-5.6 Lunalost7,400547fail$0.084.8fde1c078
jev-detroit-s2Detroit 19722Jev Routerlost10,040456fail$3.1145.985fe296a
sonnet5-detroit-s2Detroit 19722Claude Sonnet 5lost3,280341fail$0.726.6c55e7a2b
gpt56terra-detroit-s2Detroit 19722GPT-5.6 Terralost6,880693fail$1.0510.8ef93e671
gpt56luna-detroit-s2Detroit 19722GPT-5.6 Lunalost15,280704fail$0.095.4aba6000b

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →