evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-jev-gpt6luna-g4

GPT-6 Luna vs Jev Router

Jev Router destroyed the enemy base after 9.3 game-minutes · 16.9 min wall clock

LOSS england · west
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$6,300Value lost$23,600 Units killed / lost19 / 35Buildings killed / lost1 / 10 Peak army$2,700Orders issued54 Decision turns (failed)70 (0)Mean latency3.3sModel cost$0.025
WIN germany · east
Jev Router
typesafe/jev-router
Value destroyed$22,200Value lost$4,900 Units killed / lost35 / 19Buildings killed / lost9 / 0 Peak army$9,550Orders issued80 Decision turns (failed)70 (13)Mean latency6.0sModel cost$0.128

Evaluation

passchecks 10/10 verdict computed from the gating checks

StatusKindCheckDetail
passstepenvironmentEnvironment check (snapshot, engine, content, key)
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passstepnotesAppend Notes to summary.md
passcode_checkoutputs-existall present
passcode_checkgame-has-winnerjev won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 70, 'Multi1': 70}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '0/70', 'Multi1': '10/70'}
passcode_checkvideo-decodesduration 95.2s
passcode_checkvideo-completevideo ends at tick 13974 of 13974
passcode_checklimits-respectedended after 16.9 wall-minutes (budget 30.0 min): base destroyed
passchecklistnotes-explain-winsummary.md Notes section cites final thoughts from both models explaining the decisive Jev Router push
passchecklistplayers-match-requestresult.json shows Multi0=gpt6luna(openai/gpt-6-luna) and Multi1=jev(typesafe/jev-router) exactly as requested
passchecklistno-silent-outageevery decision tick in run.log has a latency entry; Jev Router had 10/70 failed turns (14%) flagged by llm-turns-valid check, no long silent gaps

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →