evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-sonnet5-gpt6luna-g3

Claude Sonnet 5 vs GPT-6 Luna

Claude Sonnet 5 destroyed the enemy base after 7.5 game-minutes · 9.7 min wall clock

WIN france · west
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$14,350Value lost$2,100 Units killed / lost24 / 1Buildings killed / lost8 / 1 Peak army$10,500Orders issued127 Decision turns (failed)57 (1)Mean latency4.2sModel cost$0.379
LOSS russia · east
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$2,100Value lost$14,350 Units killed / lost1 / 24Buildings killed / lost1 / 8 Peak army$1,800Orders issued54 Decision turns (failed)57 (0)Mean latency3.1sModel cost$0.021

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnersonnet5 won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 57, 'Multi1': 57}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/57', 'Multi1': '0/57'}
passcode_checkvideo-decodesduration 77.1s
passcode_checkvideo-completevideo ends at tick 11259 of 11259
passcode_checklimits-respectedended after 9.7 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winNotes section cites last thoughts of both models and explains tank push victory
passchecklistplayers-match-requestresult.json shows Multi0=sonnet5 (anthropic/claude-sonnet-5) and Multi1=gpt6luna (openai/gpt-6-luna) as requested
passchecklistno-silent-outagerun.log shows continuous per-tick progress with 0 LLM failures across 57 turns each

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →