evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-gpt56luna-gpt6luna-g3

GPT-5.6 Luna vs GPT-6 Luna

GPT-6 Luna destroyed the enemy base after 4.0 game-minutes · 5.0 min wall clock

LOSS russia · west
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$0Value lost$10,100 Units killed / lost0 / 10Buildings killed / lost0 / 6 Peak army$1,600Orders issued74 Decision turns (failed)30 (0)Mean latency3.4sModel cost$0.021
WIN germany · east
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$10,100Value lost$0 Units killed / lost10 / 0Buildings killed / lost6 / 0 Peak army$3,800Orders issued28 Decision turns (failed)30 (0)Mean latency3.0sModel cost$0.009

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnergpt6luna won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 30, 'Multi1': 30}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/30', 'Multi1': '0/30'}
passcode_checkvideo-decodesduration 41.6s
passcode_checkvideo-completevideo ends at tick 5939 of 5939
passcode_checklimits-respectedended after 5.0 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winNotes section cites final thoughts from both models explaining GPT-6 Luna won by base destruction at tick 5940
passchecklistplayers-match-requestresult.json confirms Multi0=gpt56luna (openai/gpt-5.6-luna) and Multi1=gpt6luna (openai/gpt-6-luna) as requested
passchecklistno-silent-outagerun.log shows uninterrupted per-tick progress lines with no gaps or model failures across all 60 decision turns

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →