evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-gpt56terra-gpt6luna-g3

GPT-5.6 Terra vs GPT-6 Luna

GPT-5.6 Terra destroyed the enemy base after 5.0 game-minutes · 6.3 min wall clock

WIN germany · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$8,100Value lost$2,400 Units killed / lost23 / 18Buildings killed / lost4 / 0 Peak army$4,400Orders issued77 Decision turns (failed)38 (0)Mean latency3.6sModel cost$0.265
LOSS ukraine · east
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$2,400Value lost$8,100 Units killed / lost18 / 23Buildings killed / lost0 / 4 Peak army$1,600Orders issued27 Decision turns (failed)38 (0)Mean latency3.0sModel cost$0.012

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnergpt56terra won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 38, 'Multi1': 38}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/38', 'Multi1': '0/38'}
passcode_checkvideo-decodesduration 51.9s
passcode_checkvideo-completevideo ends at tick 7479 of 7479
passcode_checklimits-respectedended after 6.3 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winsummary.md Notes section cites both models final thoughts and explains Terra win via infantry mass overrunning base at tick 7480
passchecklistplayers-match-requestresult.json Multi0=gpt56terra (GPT-5.6 Terra) west and Multi1=gpt6luna (GPT-6 Luna) east match the requested players
passchecklistno-silent-outagerun.log shows 38 consecutive decision turns per player with no gaps all LLM calls 0 failures

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →