evals
Run evals on Jetty
Champion vs OpenRA AI · ladder-gpt56terra-normal-g1

GPT-5.6 Terra vs OpenRA AI (Normal)

OpenRA AI (Normal) destroyed the enemy base after 10.4 game-minutes · 12.7 min wall clock

LOSS germany · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$1,050Value lost$24,700 Units killed / lost3 / 59Buildings killed / lost0 / 9 Peak army$2,700Orders issued81 Decision turns (failed)78 (0)Mean latency3.5sModel cost$0.572
WIN ukraine · east
OpenRA AI (Normal)
openra/normal-ai
Value destroyed$25,000Value lost$1,350 Units killed / lost61 / 5Buildings killed / lost9 / 0 Peak army$0Orders issued134 Decision turns (failed)0 (0)Mean latency–Model cost$0.000

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.