evals
Run evals on Jetty

02 — Run · Round 2 · Round robin · r2-rr-sonnet5-gpt56luna-g2

GPT-5.6 Luna vs Claude Sonnet 5

Claude Sonnet 5 destroyed the enemy base after 5.9 game-minutes · 7.8 min wall clock

LOSS ukraine · west
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$0Value lost$8,900 Units killed / lost0 / 14Buildings killed / lost0 / 5 Peak army$2,200Orders issued92 Decision turns (failed)45 (0)Mean latency3.3sModel cost$0.036
WIN ukraine · east
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$8,900Value lost$0 Units killed / lost14 / 0Buildings killed / lost5 / 0 Peak army$5,950Orders issued103 Decision turns (failed)45 (0)Mean latency3.9sModel cost$0.286

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnersonnet5 won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 45, 'Multi1': 45}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/45', 'Multi1': '0/45'}
passcode_checkvideo-decodesduration 61.1s
passcode_checkvideo-completevideo ends at tick 8870 of 8870
passcode_checklimits-respectedended after 7.8 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winsummary.md Notes section cites both models final thoughts and explains Sonnet 5 coordinated assault vs Luna base collapse
passchecklistplayers-match-requestresult.json shows Multi0=gpt56luna and Multi1=sonnet5 as requested
passchecklistno-silent-outagerun.log shows continuous progress every 200 ticks with no gaps; both models had zero LLM failures across 45 calls each

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Get Started Book a 15-minute walkthrough
Connect your agent to Jetty →