evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-gpt56luna-gpt6luna-g2

GPT-6 Luna vs GPT-5.6 Luna

GPT-6 Luna destroyed the enemy base after 11.9 game-minutes · 18.6 min wall clock

WIN russia · west
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$30,300Value lost$11,550 Units killed / lost54 / 42Buildings killed / lost12 / 3 Peak army$4,500Orders issued93 Decision turns (failed)90 (0)Mean latency3.2sModel cost$0.032
LOSS russia · east
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$11,550Value lost$30,300 Units killed / lost42 / 54Buildings killed / lost3 / 12 Peak army$4,200Orders issued103 Decision turns (failed)90 (0)Mean latency4.2sModel cost$0.079

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnergpt6luna won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 90, 'Multi1': 90}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/90', 'Multi1': '0/90'}
passcode_checkvideo-decodesduration 121.1s
passcode_checkvideo-completevideo ends at tick 17860 of 17860
passcode_checklimits-respectedended after 18.6 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winsummary.md Notes section cites final thoughts from both models and explains GPT-6 Luna win by base destruction
passchecklistplayers-match-requestresult.json records gpt6luna (openai/gpt-6-luna) as west and gpt56luna (openai/gpt-5.6-luna) as east, matching the requested players
passchecklistno-silent-outagerun.log shows 90 turns per player with 0 failures and no gaps; every 200-tick interval is present

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →