evals
Run evals on Jetty
Champion vs OpenRA AI · ladder-gpt56terra-medium-g1

GPT-5.6 Terra vs OpenRA AI (Medium)

OpenRA AI (Medium) destroyed the enemy base after 10.8 game-minutes · 14.6 min wall clock

LOSS ukraine · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$7,200Value lost$33,000 Units killed / lost14 / 98Buildings killed / lost2 / 12 Peak army$3,000Orders issued127 Decision turns (failed)81 (0)Mean latency4.0sModel cost$0.642
WIN germany · east
OpenRA AI (Medium)
openra/medium-ai
Value destroyed$33,000Value lost$7,200 Units killed / lost98 / 14Buildings killed / lost12 / 2 Peak army$0Orders issued152 Decision turns (failed)0 (0)Mean latency–Model cost$0.000

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.