02 — Run · Round 2 · Round robin · r2-rr-jev-gpt6luna-g4
GPT-6 Luna vs Jev Router
Jev Router destroyed the enemy base after 9.3 game-minutes · 16.9 min wall clock
LOSS england · west
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$6,300Value lost$23,600
Units killed / lost19 / 35Buildings killed / lost1 / 10
Peak army$2,700Orders issued54
Decision turns (failed)70 (0)Mean latency3.3sModel cost$0.025
WIN germany · east
Jev Router
typesafe/jev-router
Value destroyed$22,200Value lost$4,900
Units killed / lost35 / 19Buildings killed / lost9 / 0
Peak army$9,550Orders issued80
Decision turns (failed)70 (13)Mean latency6.0sModel cost$0.128
Evaluation
passchecks 10/10 verdict computed from the gating checks
| Status | Kind | Check | Detail |
|---|---|---|---|
| pass | step | environment | Environment check (snapshot, engine, content, key) |
| pass | step | launch | Launch run-match as a detached process |
| pass | step | supervise | Supervise in the foreground until run-match exits |
| pass | step | render | Render the 1080p recording |
| pass | step | notes | Append Notes to summary.md |
| pass | code_check | outputs-exist | all present |
| pass | code_check | game-has-winner | jev won: base destroyed |
| pass | code_check | decisions-logged | turns per slot: {'Multi0': 70, 'Multi1': 70} |
| pass | code_check | llm-turns-valid | failed/total per slot: {'Multi0': '0/70', 'Multi1': '10/70'} |
| pass | code_check | video-decodes | duration 95.2s |
| pass | code_check | video-complete | video ends at tick 13974 of 13974 |
| pass | code_check | limits-respected | ended after 16.9 wall-minutes (budget 30.0 min): base destroyed |
| pass | checklist | notes-explain-win | summary.md Notes section cites final thoughts from both models explaining the decisive Jev Router push |
| pass | checklist | players-match-request | result.json shows Multi0=gpt6luna(openai/gpt-6-luna) and Multi1=jev(typesafe/jev-router) exactly as requested |
| pass | checklist | no-silent-outage | every decision tick in run.log has a latency entry; Jev Router had 10/70 failed turns (14%) flagged by llm-turns-valid check, no long silent gaps |
Decision timeline
What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.