02 — Run · Round 2 · Round robin · r2-rr-jev-gpt56luna-g2
GPT-5.6 Luna vs Jev Router
Jev Router destroyed the enemy base after 4.9 game-minutes · 6.6 min wall clock
LOSS russia · west
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$1,300Value lost$8,300
Units killed / lost13 / 14Buildings killed / lost0 / 5
Peak army$550Orders issued48
Decision turns (failed)37 (0)Mean latency2.9sModel cost$0.025
WIN russia · east
Jev Router
typesafe/jev-router
Value destroyed$8,300Value lost$1,300
Units killed / lost14 / 13Buildings killed / lost5 / 0
Peak army$3,500Orders issued74
Decision turns (failed)37 (0)Mean latency3.8sModel cost$0.051
Evaluation
passchecks 10/10 verdict computed from the gating checks
| Status | Kind | Check | Detail |
|---|---|---|---|
| pass | step | environment | Environment check (snapshot, engine, content, key) |
| pass | step | launch | Launch run-match as a detached process |
| pass | step | supervise | Supervise in the foreground until run-match exits |
| pass | step | render | Render the 1080p recording |
| pass | step | notes | Append Notes to summary.md |
| pass | code_check | outputs-exist | all present |
| pass | code_check | game-has-winner | jev won: base destroyed |
| pass | code_check | decisions-logged | turns per slot: {'Multi0': 37, 'Multi1': 37} |
| pass | code_check | llm-turns-valid | failed/total per slot: {'Multi0': '0/37', 'Multi1': '0/37'} |
| pass | code_check | video-decodes | duration 51.3s |
| pass | code_check | video-complete | video ends at tick 7386 of 7386 |
| pass | code_check | limits-respected | ended after 6.6 wall-minutes (budget 30.0 min): base destroyed |
| pass | checklist | notes-explain-win | summary.md Notes section cites both players final reasoning and explains Jev Router won by base destruction via superior economy and army |
| pass | checklist | players-match-request | result.json records Multi0=gpt56luna (openai/gpt-5.6-luna) and Multi1=jev (typesafe/jev-router) matching the requested west=gpt56luna east=jev players |
| pass | checklist | no-silent-outage | run.log shows continuous 3-8s decision turns for all 37 calls per slot with zero llm_failures recorded |
Decision timeline
What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.