evals
Run evals on Jetty
Round robin · rr-sonnet5-gpt56terra-g1

Claude Sonnet 5 vs GPT-5.6 Terra

GPT-5.6 Terra destroyed the enemy base after 7.8 game-minutes · 12.2 min wall clock

LOSS ukraine · west
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$7,700Value lost$13,200 Units killed / lost43 / 29Buildings killed / lost0 / 9 Peak army$2,700Orders issued92 Decision turns (failed)59 (0)Mean latency4.7sModel cost$0.399
WIN england · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$13,200Value lost$7,700 Units killed / lost29 / 43Buildings killed / lost9 / 0 Peak army$5,800Orders issued91 Decision turns (failed)59 (0)Mean latency3.7sModel cost$0.455

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.