From a weekend idea to a reproducible benchmark
How the OpenRA arena went from “can an LLM play Red Alert?” to 20 fan-out runs with videos, decision logs and receipts, with one runbook and one sandbox snapshot on Jetty.
The problem
A real-time strategy game is a hard, honest test for an agent: partial information, a clock that doesn't wait, an economy to run and an opponent adapting to you. But an RTS eval is also an infrastructure headache: a C# game engine, native graphics, game data, an orchestrator, and games that take the better part of an hour each. Running a tournament means dozens of those in parallel, and people only trust the result if they can see every game and rerun it themselves.
What we built
- An engine bridge. A bot trait inside OpenRA pauses the simulation every 8 game-seconds, hands both players their fog-of-war view at the same tick, and resumes when both models answer. Lockstep keeps it fair and makes every game replayable.
- A sandbox snapshot. The engine, Red Alert data, a virtual display with software OpenGL, ffmpeg and the Python orchestrator
are baked into one image, registered as the Jetty snapshot
openra-arena. A run starts playing within seconds instead of spending ten minutes compiling. - A runbook. One markdown file tells the agent how to launch a game, supervise it, recover from failures and verify the outputs: result, metrics, decision log, replay, 1080p video, runtime config and a validation report.
- Fan-out. Each game is one API call with different parameters. The round robin was 12 calls; the champion's ladder against OpenRA's built-in AI was 8 more. Jetty ran them side by side and kept every trajectory.
- A results site generated from the receipts. This site is built from the artifacts the runs saved, so every number links back to the run that produced it (see every run).
What we learned
- Replay determinism is a feature you have to test. Early pilot videos drifted out of sync. The cause: the bridge sent building
sales as “immediate” orders that bypass the lockstep queue, so the replay applied them at a different tick. Every desync lined up with a sale;
routing sales through the normal queue fixed it. A validation check (
video_complete) now catches any recurrence. - Tell the agent what “waiting” means. In the first fan-out, a few runbook agents moved their polling into background tasks and ended their session, which tears down the sandbox mid-game. Runbook v1.1.0 makes supervision foreground-only and says why; the lost runs were simply relaunched with the same parameters.
- Bake, don't build. Moving the engine build into the snapshot cut each run's setup from minutes to seconds and removed a whole class of flaky failures.
- Keep the receipts machine-readable. Because each run writes
result.json,metrics.jsonandruntime_config.json, adding a new chart (like the radar profiles) is a site rebuild, not a re-run.
Run your own
Swap in your own models, maps or limits and point the same runbook at your collection. If you have an eval that needs a real environment (a game, a browser, a codebase, a GPU), the pattern is the same: bake the environment, write the runbook, fan out, publish the receipts.