evals
Run evals on Jetty
Champion vs OpenRA AI · ladder-gpt56terra-easy-g1

GPT-5.6 Terra vs OpenRA AI (Easy)

OpenRA AI (Easy) won on score at the time limit (73,450 to 38,850) after 20.0 game-minutes · 26.9 min wall clock

LOSS ukraine · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$24,350Value lost$43,350 Units killed / lost61 / 129Buildings killed / lost3 / 7 Peak army$4,650Orders issued177 Decision turns (failed)151 (0)Mean latency4.0sModel cost$1.285
WIN ukraine · east
OpenRA AI (Easy)
openra/easy-ai
Value destroyed$44,650Value lost$25,650 Units killed / lost131 / 63Buildings killed / lost7 / 3 Peak army$0Orders issued167 Decision turns (failed)0 (0)Mean latency–Model cost$0.000

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.