02 — Run · Round 1 · Round robin · rr-sonnet5-gpt56terra-g1
Claude Sonnet 5 vs GPT-5.6 Terra
GPT-5.6 Terra destroyed the enemy base after 7.8 game-minutes · 12.2 min wall clock
LOSS ukraine · west
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$7,700Value lost$13,200
Units killed / lost43 / 29Buildings killed / lost0 / 9
Peak army$2,700Orders issued92
Decision turns (failed)59 (0)Mean latency4.7sModel cost$0.399
WIN england · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$13,200Value lost$7,700
Units killed / lost29 / 43Buildings killed / lost9 / 0
Peak army$5,800Orders issued91
Decision turns (failed)59 (0)Mean latency3.7sModel cost$0.455
Evaluation
passchecks 7/7 verdict computed from the gating checks
Round-1 run: written before the runbook moved to evals v2, so these are the older self-reported checks.
| Status | Kind | Check | Detail |
|---|---|---|---|
| pass | code_check | result-json | result json |
| pass | code_check | has-winner | has winner |
| pass | code_check | decisions-logged | decisions logged |
| pass | code_check | replay-saved | replay saved |
| pass | code_check | video-rendered | video rendered |
| pass | code_check | no-llm-outage | no llm outage |
| pass | code_check | video-complete | video complete |
Decision timeline
What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.