evals
Run evals on Jetty

02 — Run · Round 1 · vs OpenRA AI · ladder-gpt56terra-beginner-g1

GPT-5.6 Terra vs OpenRA AI (Beginner)

OpenRA AI (Beginner) won on score at the time limit (34,750 to 27,700) after 20.0 game-minutes · 22.8 min wall clock

LOSS ukraine · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$7,600Value lost$27,600 Units killed / lost31 / 87Buildings killed / lost4 / 0 Peak army$4,650Orders issued126 Decision turns (failed)151 (0)Mean latency3.8sModel cost$1.211
WIN ukraine · east
OpenRA AI (Beginner)
openra/beginner-ai
Value destroyed$27,600Value lost$7,600 Units killed / lost87 / 31Buildings killed / lost0 / 4 Peak army$0Orders issued51 Decision turns (failed)0 (0)Mean latency–Model cost$0.000

Evaluation

passchecks 7/7 verdict computed from the gating checks

Round-1 run: written before the runbook moved to evals v2, so these are the older self-reported checks.

StatusKindCheckDetail
passcode_checkresult-jsonresult json
passcode_checkhas-winnerhas winner
passcode_checkdecisions-loggeddecisions logged
passcode_checkreplay-savedreplay saved
passcode_checkvideo-renderedvideo rendered
passcode_checkno-llm-outageno llm outage
passcode_checkvideo-completevideo complete

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Ready to stop guessing if your outputs are good?

Get Started Book a 15-minute walkthrough
Connect your agent to Jetty →