02 — Run · Round 2 · Round robin · r2-rr-gpt56luna-gpt6luna-g1
GPT-5.6 Luna vs GPT-6 Luna
GPT-5.6 Luna destroyed the enemy base after 5.8 game-minutes · 7.9 min wall clock
WIN england · west
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$8,700Value lost$1,000
Units killed / lost10 / 6Buildings killed / lost6 / 0
Peak army$3,100Orders issued55
Decision turns (failed)44 (0)Mean latency3.5sModel cost$0.034
LOSS germany · east
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$1,000Value lost$8,700
Units killed / lost6 / 10Buildings killed / lost0 / 6
Peak army$1,500Orders issued28
Decision turns (failed)44 (0)Mean latency3.0sModel cost$0.015
Evaluation
passchecks 10/10 verdict computed from the gating checks
| Status | Kind | Check | Detail |
|---|---|---|---|
| pass | step | environment | Environment check (snapshot, engine, content, key) |
| pass | step | launch | Launch run-match as a detached process |
| pass | step | supervise | Supervise in the foreground until run-match exits |
| pass | step | render | Render the 1080p recording |
| pass | step | notes | Append Notes to summary.md |
| pass | code_check | outputs-exist | all present |
| pass | code_check | game-has-winner | gpt56luna won: base destroyed |
| pass | code_check | decisions-logged | turns per slot: {'Multi0': 44, 'Multi1': 44} |
| pass | code_check | llm-turns-valid | failed/total per slot: {'Multi0': '0/44', 'Multi1': '0/44'} |
| pass | code_check | video-decodes | duration 60.0s |
| pass | code_check | video-complete | video ends at tick 8695 of 8695 |
| pass | code_check | limits-respected | ended after 7.9 wall-minutes (budget 30.0 min): base destroyed |
| pass | checklist | notes-explain-win | summary.md Notes section explains GPT-5.6 Luna win by base destruction, citing final thoughts of both models |
| pass | checklist | players-match-request | result.json shows Multi0=gpt56luna (GPT-5.6 Luna) and Multi1=gpt6luna (GPT-6 Luna) matching the requested players |
| pass | checklist | no-silent-outage | run.log shows 44 successful turns per player with no gaps or failed turns; llm_failures are 0 for both slots |
Decision timeline
What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.