evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-gpt56luna-gpt6luna-g1

GPT-5.6 Luna vs GPT-6 Luna

GPT-5.6 Luna destroyed the enemy base after 5.8 game-minutes · 7.9 min wall clock

WIN england · west
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$8,700Value lost$1,000 Units killed / lost10 / 6Buildings killed / lost6 / 0 Peak army$3,100Orders issued55 Decision turns (failed)44 (0)Mean latency3.5sModel cost$0.034
LOSS germany · east
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$1,000Value lost$8,700 Units killed / lost6 / 10Buildings killed / lost0 / 6 Peak army$1,500Orders issued28 Decision turns (failed)44 (0)Mean latency3.0sModel cost$0.015

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnergpt56luna won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 44, 'Multi1': 44}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/44', 'Multi1': '0/44'}
passcode_checkvideo-decodesduration 60.0s
passcode_checkvideo-completevideo ends at tick 8695 of 8695
passcode_checklimits-respectedended after 7.9 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winsummary.md Notes section explains GPT-5.6 Luna win by base destruction, citing final thoughts of both models
passchecklistplayers-match-requestresult.json shows Multi0=gpt56luna (GPT-5.6 Luna) and Multi1=gpt6luna (GPT-6 Luna) matching the requested players
passchecklistno-silent-outagerun.log shows 44 successful turns per player with no gaps or failed turns; llm_failures are 0 for both slots

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →