02 — Run · Round 2 · Round robin · r2-rr-gpt56terra-gpt56luna-g4
GPT-5.6 Luna vs GPT-5.6 Terra
GPT-5.6 Terra destroyed the enemy base after 6.1 game-minutes · 7.8 min wall clock
LOSS germany · west
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$3,100Value lost$10,800
Units killed / lost18 / 17Buildings killed / lost1 / 6
Peak army$1,900Orders issued92
Decision turns (failed)46 (0)Mean latency3.1sModel cost$0.036
WIN france · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$10,800Value lost$3,100
Units killed / lost17 / 18Buildings killed / lost6 / 1
Peak army$4,000Orders issued55
Decision turns (failed)46 (0)Mean latency3.3sModel cost$0.320
Evaluation
passchecks 10/10 verdict computed from the gating checks
| Status | Kind | Check | Detail |
|---|---|---|---|
| pass | step | environment | Environment check (snapshot, engine, content, key) |
| pass | step | launch | Launch run-match as a detached process |
| pass | step | supervise | Supervise in the foreground until run-match exits |
| pass | step | render | Render the 1080p recording |
| pass | step | notes | Append Notes to summary.md |
| pass | code_check | outputs-exist | all present |
| pass | code_check | game-has-winner | gpt56terra won: base destroyed |
| pass | code_check | decisions-logged | turns per slot: {'Multi0': 46, 'Multi1': 46} |
| pass | code_check | llm-turns-valid | failed/total per slot: {'Multi0': '0/46', 'Multi1': '0/46'} |
| pass | code_check | video-decodes | duration 63.2s |
| pass | code_check | video-complete | video ends at tick 9171 of 9171 |
| pass | code_check | limits-respected | ended after 7.8 wall-minutes (budget 30.0 min): base destroyed |
| pass | checklist | notes-explain-win | Notes section explains GPT-5.6 Terra win by base destruction, citing both models final thoughts and economic strategy |
| pass | checklist | players-match-request | result.json shows gpt56luna as Multi0 (west) and gpt56terra as Multi1 (east), matching the requested parameters |
| pass | checklist | no-silent-outage | run.log shows continuous decision turns from tick 201 to 9001 with zero LLM failures for either model across 46 calls each |
Decision timeline
What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.