02 — Run · Round 2 · Round robin · r2-rr-jev-gpt56terra-g3
Jev Router vs GPT-5.6 Terra
Jev Router destroyed the enemy base after 14.6 game-minutes · 21.8 min wall clock
WIN russia · west
Jev Router
typesafe/jev-router
Value destroyed$36,800Value lost$10,300
Units killed / lost114 / 45Buildings killed / lost10 / 3
Peak army$12,600Orders issued218
Decision turns (failed)110 (4)Mean latency5.7sModel cost$0.211
LOSS ukraine · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$10,300Value lost$36,800
Units killed / lost45 / 114Buildings killed / lost3 / 10
Peak army$4,050Orders issued144
Decision turns (failed)110 (0)Mean latency3.9sModel cost$0.897
Evaluation
passchecks 10/10 verdict computed from the gating checks
| Status | Kind | Check | Detail |
|---|---|---|---|
| pass | step | environment | Environment check (snapshot, engine, content, key) |
| pass | step | launch | Launch run-match as a detached process |
| pass | step | supervise | Supervise in the foreground until run-match exits |
| pass | step | render | Render the 1080p recording |
| pass | step | notes | Append Notes to summary.md |
| pass | code_check | outputs-exist | all present |
| pass | code_check | game-has-winner | jev won: base destroyed |
| pass | code_check | decisions-logged | turns per slot: {'Multi0': 110, 'Multi1': 110} |
| pass | code_check | llm-turns-valid | failed/total per slot: {'Multi0': '1/110', 'Multi1': '0/110'} |
| pass | code_check | video-decodes | duration 148.1s |
| pass | code_check | video-complete | video ends at tick 21909 of 21909 |
| pass | code_check | limits-respected | ended after 21.8 wall-minutes (budget 30.0 min): base destroyed |
| pass | checklist | notes-explain-win | summary.md Notes section cites final thoughts of both players and explains how Jev Router overcame an early deficit to destroy GPT Terra base |
| pass | checklist | players-match-request | result.json records Multi0=jev (typesafe/jev-router) and Multi1=gpt56terra (openai/gpt-5.6-terra) matching the requested players |
| pass | checklist | no-silent-outage | run.log shows continuous tick-by-tick progress with 1 LLM failure out of 220 total calls and no long stretches of failed turns |
Decision timeline
What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.