evals
Run evals on Jetty
Champion vs OpenRA AI · ladder-gpt56terra-medium-g2

OpenRA AI (Medium) vs GPT-5.6 Terra

OpenRA AI (Medium) destroyed the enemy base after 5.8 game-minutes · 8.1 min wall clock

WIN england · west
OpenRA AI (Medium)
openra/medium-ai
Value destroyed$12,000Value lost$300 Units killed / lost23 / 3Buildings killed / lost6 / 0 Peak army$0Orders issued62 Decision turns (failed)0 (0)Mean latency–Model cost$0.000
LOSS ukraine · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$300Value lost$12,000 Units killed / lost3 / 23Buildings killed / lost0 / 6 Peak army$2,200Orders issued51 Decision turns (failed)44 (0)Mean latency3.1sModel cost$0.299

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.