02 — Run · Round 2 · Round robin · r2-rr-jev-gpt6luna-g3
Jev Router vs GPT-6 Luna
Jev Router destroyed the enemy base after 9.3 game-minutes · 16.6 min wall clock
WIN england · west
Jev Router
typesafe/jev-router
Value destroyed$19,000Value lost$4,700
Units killed / lost24 / 17Buildings killed / lost9 / 0
Peak army$7,550Orders issued94
Decision turns (failed)71 (20)Mean latency6.0sModel cost$0.118
LOSS germany · east
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$4,600Value lost$18,900
Units killed / lost16 / 23Buildings killed / lost0 / 9
Peak army$2,300Orders issued48
Decision turns (failed)71 (0)Mean latency3.0sModel cost$0.024
Evaluation
passchecks 10/10 verdict computed from the gating checks
| Status | Kind | Check | Detail |
|---|---|---|---|
| pass | step | environment | Environment check (snapshot, engine, content, key) |
| pass | step | launch | Launch run-match as a detached process |
| pass | step | supervise | Supervise in the foreground until run-match exits |
| pass | step | render | Render the 1080p recording |
| pass | step | notes | Append Notes to summary.md |
| pass | code_check | outputs-exist | all present |
| pass | code_check | game-has-winner | jev won: base destroyed |
| pass | code_check | decisions-logged | turns per slot: {'Multi0': 71, 'Multi1': 71} |
| pass | code_check | llm-turns-valid | failed/total per slot: {'Multi0': '13/71', 'Multi1': '0/71'} |
| pass | code_check | video-decodes | duration 95.4s |
| pass | code_check | video-complete | video ends at tick 14001 of 14001 |
| pass | code_check | limits-respected | ended after 16.6 wall-minutes (budget 30.0 min): base destroyed |
| pass | checklist | notes-explain-win | Notes section cites final thoughts from both Jev Router and GPT-6 Luna and explains how Jev Router won by economy and aggression |
| pass | checklist | players-match-request | result.json confirms Multi0=jev (Jev Router/typesafe/jev-router) and Multi1=gpt6luna (GPT-6 Luna/openai/gpt-6-luna) matching the requested players |
| pass | checklist | no-silent-outage | 13 isolated LLM failures for Multi0 (18.3% of 71 calls, under 20% threshold) visible in run.log as delayed turns; no consecutive empty stretch; Multi1 had 0 failures |
Decision timeline
What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.