evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-gpt56terra-gpt6luna-g1

GPT-5.6 Terra vs GPT-6 Luna

GPT-5.6 Terra destroyed the enemy base after 5.3 game-minutes · 6.6 min wall clock

WIN ukraine · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$8,400Value lost$2,600 Units killed / lost15 / 22Buildings killed / lost4 / 0 Peak army$4,300Orders issued79 Decision turns (failed)40 (0)Mean latency3.6sModel cost$0.283
LOSS russia · east
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$2,600Value lost$8,400 Units killed / lost22 / 15Buildings killed / lost0 / 4 Peak army$1,800Orders issued29 Decision turns (failed)40 (0)Mean latency3.1sModel cost$0.014

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnergpt56terra won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 40, 'Multi1': 40}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/40', 'Multi1': '0/40'}
passcode_checkvideo-decodesduration 54.6s
passcode_checkvideo-completevideo ends at tick 7891 of 7891
passcode_checklimits-respectedended after 6.6 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winNotes section cites final thoughts from both models, explaining Terras infantry-swarm push and Lunas economic collapse
passchecklistplayers-match-requestresult.json records Multi0=gpt56terra (openai/gpt-5.6-terra) and Multi1=gpt6luna (openai/gpt-6-luna) matching the requested players
passchecklistno-silent-outagerun.log shows 41 consecutive decision-turn lines with no gaps; 0/40 LLM failures for both players

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →