evals
Run evals on Jetty
Pilot cup · cup1-third-g2

GPT-5.6 Terra vs Gemini 3.8 Flash

GPT-5.6 Terra destroyed the enemy base after 5.9 game-minutes · 6.2 min wall clock

The recording stops at 3.3 of 5.9 game-minutes: this pilot game hit a since-fixed bridge bug (building sales were sent outside the lockstep order queue), so its replay drifts out of sync at the first sale. Results and stats come from the live game.

Pilot game (run locally, before the Jetty runs) Replay (.orarep) Decision log Runtime config
WIN ukraine · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$8,600Value lost$2,400 Units killed / lost13 / 18Buildings killed / lost5 / 0 Peak army$3,400Orders issued60 Decision turns (failed)45 (0)Mean latency3.1sModel cost$0.311
LOSS ukraine · east
Gemini 3.8 Flash
google/gemini-3.8-flash
Value destroyed$2,400Value lost$8,600 Units killed / lost18 / 13Buildings killed / lost0 / 5 Peak army$1,800Orders issued30 Decision turns (failed)45 (0)Mean latency2.8sModel cost$0.147

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.