evals
Run evals on Jetty

02 — Run · Round 1 · Round robin · rr-jev-gpt56terra-g3

Jev Router vs GPT-5.6 Terra

GPT-5.6 Terra destroyed the enemy base after 5.2 game-minutes · 7.8 min wall clock

LOSS ukraine · west
Jev Router
typesafe/jev-router
Value destroyed$1,400Value lost$9,400 Units killed / lost14 / 16Buildings killed / lost0 / 5 Peak army$1,900Orders issued75 Decision turns (failed)39 (0)Mean latency5.5sModel cost$0.092
WIN russia · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$9,600Value lost$1,600 Units killed / lost18 / 16Buildings killed / lost5 / 0 Peak army$2,600Orders issued64 Decision turns (failed)39 (0)Mean latency3.3sModel cost$0.266

Evaluation

passchecks 7/7 verdict computed from the gating checks

Round-1 run: written before the runbook moved to evals v2, so these are the older self-reported checks.

StatusKindCheckDetail
passcode_checkresult-jsonresult json
passcode_checkhas-winnerhas winner
passcode_checkdecisions-loggeddecisions logged
passcode_checkreplay-savedreplay saved
passcode_checkvideo-renderedvideo rendered
passcode_checkno-llm-outageno llm outage
passcode_checkvideo-completevideo complete

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Get Started Book a 15-minute walkthrough
Connect your agent to Jetty →