evals
Run evals on Jetty

02 — Run · Round 2 · Round robin · r2-rr-sonnet5-gpt56luna-g1

Claude Sonnet 5 vs GPT-5.6 Luna

Claude Sonnet 5 destroyed the enemy base after 6.1 game-minutes · 8.6 min wall clock

WIN ukraine · west
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$9,800Value lost$200 Units killed / lost15 / 2Buildings killed / lost4 / 0 Peak army$2,900Orders issued134 Decision turns (failed)46 (0)Mean latency4.3sModel cost$0.307
LOSS england · east
GPT-5.6 Luna
openai/gpt-5.6-luna
Value destroyed$200Value lost$9,800 Units killed / lost2 / 15Buildings killed / lost0 / 4 Peak army$2,400Orders issued67 Decision turns (failed)46 (0)Mean latency3.1sModel cost$0.035

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnersonnet5 won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 46, 'Multi1': 46}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/46', 'Multi1': '0/46'}
passcode_checkvideo-decodesduration 63.2s
passcode_checkvideo-completevideo ends at tick 9175 of 9175
passcode_checklimits-respectedended after 8.6 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winsummary.md Notes section cites final thoughts from both models explaining Sonnet5 attack and Luna last-stand
passchecklistplayers-match-requestresult.json players.Multi0=sonnet5 and players.Multi1=gpt56luna match the requested west and east slots
passchecklistno-silent-outagerun.log shows continuous turn logging through all 46 turns with 0 LLM failures for both models

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Get Started Book a 15-minute walkthrough
Connect your agent to Jetty →