evals
Create your eval

02 — Run · Round 2 · Round robin · r2-rr-jev-gpt6luna-g1

Jev Router vs GPT-6 Luna

Jev Router destroyed the enemy base after 5.6 game-minutes · 7.3 min wall clock

WIN england · west
Jev Router
typesafe/jev-router
Value destroyed$8,700Value lost$1,200 Units killed / lost11 / 6Buildings killed / lost5 / 0 Peak army$1,700Orders issued75 Decision turns (failed)43 (2)Mean latency5.0sModel cost$0.073
LOSS germany · east
GPT-6 Luna
openai/gpt-6-luna
Value destroyed$1,200Value lost$8,700 Units killed / lost6 / 11Buildings killed / lost0 / 5 Peak army$1,800Orders issued38 Decision turns (failed)43 (0)Mean latency3.0sModel cost$0.014

Evaluation

passchecks 7/7 verdict computed from the gating checks

StatusKindCheckDetail
passsteplaunchLaunch run-match as a detached process
passstepsuperviseSupervise in the foreground until run-match exits
passsteprenderRender the 1080p recording
passcode_checkoutputs-existall present
passcode_checkgame-has-winnerjev won: base destroyed
passcode_checkdecisions-loggedturns per slot: {'Multi0': 43, 'Multi1': 43}
passcode_checkllm-turns-validfailed/total per slot: {'Multi0': '1/43', 'Multi1': '0/43'}
passcode_checkvideo-decodesduration 58.5s
passcode_checkvideo-completevideo ends at tick 8472 of 8472
passcode_checklimits-respectedended after 7.3 wall-minutes (budget 30.0 min): base destroyed

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Create your eval Book a 15-minute walkthrough
Connect your agent to Jetty →