evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-gpt56luna-gpt6luna-g4

GPT-6 Luna vs GPT-5.6 Luna

GPT-6 Luna destroyed the enemy base after 6.4 game-minutes · 10.8 min wall clock

WIN russia · west
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$16,200Value lost$4,100 Units killed / lost11 / 17Buildings killed / lost6 / 0 Peak army$4,000Orders issued44 Decision turns (failed)49 (0)Mean latency2.9sModel cost$0.016
LOSS england · east
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$4,100Value lost$16,200 Units killed / lost17 / 11Buildings killed / lost0 / 6 Peak army$1,400Orders issued29 Decision turns (failed)49 (0)Mean latency3.8sModel cost$0.036

Evaluation

passchecks 7/7 verdict computed from the gating checks

StatusKindCheckDetail
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passcode_checkoutputs-existall present
passcode_checkgame-has-winnergpt6luna won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 49, 'Multi1': 49}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/49', 'Multi1': '0/49'}
passcode_checkvideo-decodesduration 66.4s
passcode_checkvideo-completevideo ends at tick 9651 of 9651
passcode_checklimits-respectedended after 10.8 wall-minutes (budget 30.0 min): base destroyed

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →