evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-sonnet5-gpt6luna-g2

GPT-6 Luna vs Claude Sonnet 5

Claude Sonnet 5 destroyed the enemy base after 6.3 game-minutes · 8.7 min wall clock

LOSS france · west
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$800Value lost$14,800 Units killed / lost2 / 18Buildings killed / lost2 / 8 Peak army$1,900Orders issued39 Decision turns (failed)48 (0)Mean latency3.0sModel cost$0.016
WIN ukraine · east
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$14,800Value lost$800 Units killed / lost18 / 2Buildings killed / lost8 / 2 Peak army$5,400Orders issued111 Decision turns (failed)48 (0)Mean latency4.2sModel cost$0.306

Evaluation

passchecks 7/7 verdict computed from the gating checks

StatusKindCheckDetail
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passcode_checkoutputs-existall present
passcode_checkgame-has-winnersonnet5 won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 48, 'Multi1': 48}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/48', 'Multi1': '0/48'}
passcode_checkvideo-decodesduration 64.7s
passcode_checkvideo-completevideo ends at tick 9401 of 9401
passcode_checklimits-respectedended after 8.7 wall-minutes (budget 30.0 min): base destroyed

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →