evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-sonnet5-gpt6luna-g4

GPT-6 Luna vs Claude Sonnet 5

GPT-6 Luna destroyed the enemy base after 4.6 game-minutes · 7.5 min wall clock

WIN ukraine · west
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$12,300Value lost$1,200 Units killed / lost4 / 6Buildings killed / lost10 / 0 Peak army$3,800Orders issued45 Decision turns (failed)35 (0)Mean latency3.1sModel cost$0.011
LOSS germany · east
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$1,200Value lost$12,300 Units killed / lost6 / 4Buildings killed / lost0 / 10 Peak army$1,400Orders issued40 Decision turns (failed)35 (0)Mean latency4.6sModel cost$0.201

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnergpt6luna won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 35, 'Multi1': 35}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/35', 'Multi1': '0/35'}
passcode_checkvideo-decodesduration 47.7s
passcode_checkvideo-completevideo ends at tick 6851 of 6851
passcode_checklimits-respectedended after 7.5 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winNotes section added to summary.md citing both models final thoughts on win/loss
passchecklistplayers-match-requestresult.json confirms Multi0=gpt6luna (openai/gpt-6-luna) and Multi1=sonnet5 (anthropic/claude-sonnet-5) matching request
passchecklistno-silent-outage0 LLM failures across 35 turns each for both models, no gaps in run.log decision output

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →