01 — Author
Tell Jetty
the job to be done.
One runbook plus a sandbox snapshot. Jetty scores every run against the runbook's checks.
The job
OpenRA is an open-source engine for Command & Conquer: Red Alert. Each run is one 1v1 skirmish on Singles, a two-player map, with random factions, one MCV and $5,000 each. A player wins by destroying every enemy building. Games end after a 30-minute wall-clock budget, when the side with the lower score (assets value plus value of enemy units and buildings destroyed) surrenders.
Lockstep decisions and fairness
A C# bot trait inside the engine pauses the simulation every 200 ticks (8 game-seconds), serialises what each player can see (its units, buildings, production, economy, and enemies currently visible through fog of war) and sends both views at the same tick. Both models are called in parallel and the game resumes only when both have answered, so neither sees the other's move first and model latency never costs game time. Latency does cost wall-clock time, so under the wall-clock budget a slow pair plays fewer game-minutes than a fast one.
Each turn a model gets a text briefing (economy, power, what it can build with prices, every unit and building with IDs and cells,
visible enemies, the result of its previous commands, and a memory note it wrote last turn) and replies with JSON: its reasoning, a
note for next turn, and commands such as build, place, train, attack_move,
attack, deploy, rally, repair and sell. Invalid commands are rejected with an
explanation the model sees next turn. Models run through OpenRouter with low reasoning effort where supported, up to
4000 output tokens and a 120-second timeout; a failed reply is a turn with no orders.
How a run is checked (evals v2)
Deterministic code checks and a reviewer checklist, taken from the runbook. The agent writes them into
validation_report.json in Jetty's evals-v2 shape and Jetty computes the verdict: any failing check fails the run.
| Code check | What it verifies |
|---|---|
| outputs-exist | Every required output file exists and is non-empty ```bash cd {{results_dir}} && for f in result.json metrics.json decisions.jsonl replay.orarep match.mp4 thumb.jpg runtime_config.json summary.md; do test -s "$f" || { echo "missing $f"; exit 1; }; done ``` ### game-has-winner — The game finished with a winner ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); assert r.get('winner'), 'no winner'; print(r['winner'], r['end_reason'])" ``` ### decisions-logged — Every model player has decision turns in decisions.jsonl ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); s={json.loads(l)['slot'] for l in open('{{results_dir}}/decisions.jsonl')}; need={k for k,v in r['providers'].items() if v!='cpu'}; assert need<=s, need-s" ``` ### llm-turns-valid — At most 20% of each model's turns failed ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); c,f=r['llm_calls'],r['llm_failures']; bad={s:(f[s],c[s]) for s in c if f[s]>max(3,0.2*c[s])}; assert not bad, bad" ``` ### video-decodes — match.mp4 decodes ```bash test "$(ffprobe -v error -show_entries format=duration -of csv=p=0 {{results_dir}}/match.mp4 | cut -d. -f1)" -gt 0 ``` ### video-complete — The recording covers the whole game ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/render.json')); assert not r.get('truncated'), r" ``` ### limits-respected — The game ended within the wall-clock budget ```bash python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); assert r['wall_seconds']/60 <= {{max_wall_minutes}}+8, r['wall_seconds']" ``` ## Checklist Confirm each item by inspection; a failed item fails the run. - [ ] `notes-explain-win`: summary.md has a Notes section explaining how the game was won, citing both models' own reasoning - [ ] `players-match-request`: the players and models recorded in result.json are the ones requested in `{{player_a}}` and `{{player_b}}` - [ ] `no-silent-outage`: run.log shows no long stretch of failed or empty model turns that the code checks did not already flag ## Write Validation Report Write `{{results_dir}}/validation_report.json` in the evals-v2 shape with `arena-report`. It runs every code check above, records one `step` entry per step you pass in, and one `checklist` entry per item you pass in (status, a colon, then a one-line reason). Jetty computes the verdict from these checks. |
| Checklist item | Reviewer confirms |
|---|---|
| notes-explain-win | summary.md has a Notes section explaining how the game was won, citing both models' own reasoning |
| players-match-request | the players and models recorded in result.json are the ones requested in `{{player_a}}` and `{{player_b}}` |
| no-silent-outage | run.log shows no long stretch of failed or empty model turns that the code checks did not already flag |
The environment: sandbox snapshot openra-arena
4 vCPU / 8 GB. Debian bookworm with the .NET 8 runtime, OpenRA built from the OpenRA-RL fork at 9b271c1 plus the arena
patch, Red Alert freeware data, Xvfb with Mesa software OpenGL (games play with the real renderer so replays stay deterministic), ffmpeg,
Python and the orchestrator. Recordings re-simulate the replay in a capture mode: 1920×1080, 30 fps, 6× speed.
Players
A preset ID, openrouter:<model>|<Display name> for any OpenRouter chat model, or cpu:<level>
for OpenRA's scripted AI.
| Spec | Name | Model |
|---|---|---|
| jev | Jev Router | typesafe/jev-router |
| sonnet5 | Claude Sonnet 5 | anthropic/claude-sonnet-5 |
| gpt56terra | GPT-5.6 Terra | openai/gpt-5.6-terra |
| gpt56luna | GPT-5.6 Luna | openai/gpt-5.6-luna |
| cpu-beginner | OpenRA AI (Beginner) | openra/beginner-ai |
| cpu-easy | OpenRA AI (Easy) | openra/easy-ai |
| cpu-medium | OpenRA AI (Medium) | openra/medium-ai |
| cpu-normal | OpenRA AI (Normal) | openra/normal-ai |
Run it on your own Jetty
Set JETTY_TOKEN to an API key for your collection and COLLECTION to its name. The
openra-arena snapshot is available to Jetty runbooks.
# 1. Create the task in your collection (once)
curl -s -X POST https://flows-api.jetty.io/api/v1/tasks/$COLLECTION \
-H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" \
-d "$(jq -n --rawfile rb RUNBOOK.md --slurpfile wf task-workflow.json \
'{name: "openra-arena-match", workflow: ($wf[0] | .init_params.instruction = $rb)}')"
# 2. Your collection needs OPENROUTER_API_KEY (Settings → Environment, or:)
curl -s -X PATCH https://flows-api.jetty.io/api/v1/collections/$COLLECTION/environment \
-H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" \
--data-binary '{"environment_variables": {"OPENROUTER_API_KEY": "sk-or-..."}}'
# 3. Run a game: any OpenRouter model vs another, or vs OpenRA's AI
curl -s -X POST https://flows-api.jetty.io/api/v1/run/$COLLECTION/openra-arena-match \
-H "Authorization: Bearer $JETTY_TOKEN" \
-F 'init_params={"vars": {"player_a": "openrouter:anthropic/claude-sonnet-5|Claude Sonnet 5",
"player_b": "cpu:normal", "game_id": "my-first-game", "max_wall_minutes": "30"}}'
The runbook
---
version: "2.0.0"
evals_version: 2
evaluation: programmatic
agent: claude-code
model: anthropic/claude-sonnet-4.6
model_provider: openrouter
snapshot: openra-arena
primary_outputs:
- match.mp4
- summary.md
- result.json
secrets:
OPENROUTER_API_KEY:
env: OPENROUTER_API_KEY
description: "OpenRouter key used by the two competing models (and by this runbook's agent)"
required: true
---
# OpenRA Arena: one LLM-vs-LLM Red Alert game
## Objective
Play **one full 1v1 game of Command & Conquer: Red Alert (OpenRA)** between two decision models and deliver the
replay, a 1080p recording, the full decision log, scored metrics and an evals-v2 validation report to `{{results_dir}}`.
Each player is an LLM (any OpenRouter model) or OpenRA's built-in AI. The engine pauses every 8 seconds of game time,
sends both players their own fog-of-war view **at the same tick**, and resumes only when both have answered, so model
latency never costs game time. A game ends when one side loses all its buildings, or after `{{max_wall_minutes}}`
minutes of real (wall-clock) time, when the side with the lower score (assets value + value of enemy units and
buildings destroyed) surrenders. `{{max_ticks}}` is only a safety cap on game time (25 ticks = 1 game-second).
The `openra-arena` snapshot ships everything prebuilt: the patched OpenRA engine at `/opt/openra`, Red Alert game data at
`/opt/openra-content`, the orchestrator at `/opt/decision-ra`, Xvfb + Mesa (software OpenGL) and ffmpeg. The entry
points are `run-match` (plays, records and scores a game) and `arena-report` (writes the validation report). Do not build
or install anything.
## REQUIRED OUTPUT FILES
- `{{results_dir}}/result.json`: winner, end reason, final per-player stats, model cost, latency and failure counts
- `{{results_dir}}/metrics.json`: per-player metrics (value destroyed and lost, peak army, orders, latency, cost)
- `{{results_dir}}/decisions.jsonl`: every decision turn (the briefing each model saw, its reasoning, its commands and their results)
- `{{results_dir}}/replay.orarep`: OpenRA replay
- `{{results_dir}}/match.mp4`: 1920x1080 recording of the whole game at 6x speed
- `{{results_dir}}/thumb.jpg`: thumbnail from the recording
- `{{results_dir}}/runtime_config.json`: players, models, map, limits, engine commit and patch hash, host
- `{{results_dir}}/summary.md`: result table plus the Notes you add in Step 5
- `{{results_dir}}/validation_report.json`: evals-v2 report (`"version": 2`, `checks[]`), written in the last section
## Parameters
| Parameter | Template Variable | Default | Description |
|-----------|-------------------|---------|-------------|
| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |
| West player | `{{player_a}}` | `jev` | Player in the west slot (see Player specs) |
| East player | `{{player_b}}` | `sonnet5` | Player in the east slot |
| Game ID | `{{game_id}}` | `game` | Label for this game (used in logs and the result) |
| Wall-clock limit | `{{max_wall_minutes}}` | `30` | Real-time minutes before the game is decided on score |
| Max ticks | `{{max_ticks}}` | `90000` | Safety cap on game time (90000 ticks = 60 game-minutes) |
Player specs:
- `jev`, `sonnet5`, `gpt56terra`, `gpt56luna`, `gemini38f`: roster presets (Jev Router, Claude Sonnet 5, GPT-5.6 Terra,
GPT-5.6 Luna, Gemini 3.8 Flash)
- `openrouter:<model>` or `openrouter:<model>|<Display Name>`: any OpenRouter chat model
- `cpu:beginner`, `cpu:easy`, `cpu:medium`, `cpu:normal`: OpenRA's built-in skirmish AI, weakest to strongest
## Dependencies
- `openra-arena` snapshot: engine at `/opt/openra` (`yxc20089/OpenRA@9b271c1` + `patches/openra-arena.patch`), game data at
`/opt/openra-content`, `run-match` and `arena-report` on PATH
- `OPENROUTER_API_KEY` collection environment variable (not needed when both players are `cpu:*`)
## Steps
### Step 1: Environment Setup
```bash
mkdir -p {{results_dir}}
test -x /usr/local/bin/run-match && test -x /usr/local/bin/arena-report && test -f /opt/openra/bin/OpenRA.dll \
&& test -d /opt/openra-content/ra/v2 && echo "PASS: openra-arena snapshot present" || echo "FAIL: not running on the openra-arena snapshot"
test -n "${OPENROUTER_API_KEY:-}" && echo "PASS: OPENROUTER_API_KEY set" || echo "WARN: OPENROUTER_API_KEY missing (only cpu:* players will work)"
nproc; free -g | head -2
```
If the snapshot check fails, stop: record `environment` as failed in the validation report (last section) and explain
in `summary.md`. Do not try to build OpenRA inside the sandbox.
### Step 2: Launch the game as a detached process
A game takes up to `{{max_wall_minutes}}` minutes plus a few minutes to render, longer than one shell call may block.
Start it with `nohup` (a plain shell command, not an agent background task):
```bash
cd {{results_dir}}
nohup run-match --a "{{player_a}}" --b "{{player_b}}" --game-id "{{game_id}}" \
--max-wall-minutes {{max_wall_minutes}} --max-ticks {{max_ticks}} --out {{results_dir}} > {{results_dir}}/run.log 2>&1 &
echo $! > /tmp/run-match.pid
sleep 60; tail -5 {{results_dir}}/run.log
```
Within a minute `run.log` shows `Starting ...` and one line per decision turn:
`[game] tick 1201 | <A>: assets ... | <B>: assets ... | 6.2s`.
### Step 3: Supervise until the game ends (foreground only)
> **CRITICAL:** the sandbox, and the game with it, is destroyed the moment your session ends. Stay in this step until
> `run-match` has exited. Do NOT use background tasks, `run_in_background`, `tail -f`, cron or any other asynchronous
> watcher, and do NOT end your turn "to check back later": there is no later.
Run this exact command. It blocks for at most 9 minutes and prints the latest progress line:
```bash
PID=$(cat /tmp/run-match.pid)
for i in $(seq 1 54); do
if ! kill -0 "$PID" 2>/dev/null; then echo "RUN-MATCH EXITED"; break; fi
sleep 10
done
tail -2 {{results_dir}}/run.log
```
Repeat it until the output contains `RUN-MATCH EXITED` (expect 4 to 6 repetitions). Never kill a game that is still
logging turns.
### Step 4: Handle failures (max 1 retry)
Read the preliminary `{{results_dir}}/validation_report.json` that `run-match` wrote.
- `game-has-winner` failed (engine crash or lost connection): check `engine.log` and `run.log`, then relaunch Step 2
**once** with `--game-id "{{game_id}}-retry"` and supervise it again.
- `video-decodes` failed but the game has a winner: rerun only the render with
`cd /opt/decision-ra && DISPLAY=:99 .venv/bin/python -m arena.render /tmp/arena/{{game_id}}`, then copy `match.mp4`
and `thumb.jpg` into `{{results_dir}}`.
- Never replay a game that already has a winner just to get a different result.
### Step 5: Notes
Append a `## Notes` section to `{{results_dir}}/summary.md`: anything unusual in `run.log` (model errors, retries,
truncated video) and two or three sentences on how the game was won, citing the last few `thoughts` of each player from
`decisions.jsonl`.
## Code Checks
Each check is a deterministic command; a failing check fails the run. `arena-report` runs the same checks and records
them, so run these for your own verification and then write the report as described below.
### outputs-exist — Every required output file exists and is non-empty
```bash
cd {{results_dir}} && for f in result.json metrics.json decisions.jsonl replay.orarep match.mp4 thumb.jpg runtime_config.json summary.md; do test -s "$f" || { echo "missing $f"; exit 1; }; done
```
### game-has-winner — The game finished with a winner
```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); assert r.get('winner'), 'no winner'; print(r['winner'], r['end_reason'])"
```
### decisions-logged — Every model player has decision turns in decisions.jsonl
```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); s={json.loads(l)['slot'] for l in open('{{results_dir}}/decisions.jsonl')}; need={k for k,v in r['providers'].items() if v!='cpu'}; assert need<=s, need-s"
```
### llm-turns-valid — At most 20% of each model's turns failed
```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); c,f=r['llm_calls'],r['llm_failures']; bad={s:(f[s],c[s]) for s in c if f[s]>max(3,0.2*c[s])}; assert not bad, bad"
```
### video-decodes — match.mp4 decodes
```bash
test "$(ffprobe -v error -show_entries format=duration -of csv=p=0 {{results_dir}}/match.mp4 | cut -d. -f1)" -gt 0
```
### video-complete — The recording covers the whole game
```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/render.json')); assert not r.get('truncated'), r"
```
### limits-respected — The game ended within the wall-clock budget
```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); assert r['wall_seconds']/60 <= {{max_wall_minutes}}+8, r['wall_seconds']"
```
## Checklist
Confirm each item by inspection; a failed item fails the run.
- [ ] `notes-explain-win`: summary.md has a Notes section explaining how the game was won, citing both models' own reasoning
- [ ] `players-match-request`: the players and models recorded in result.json are the ones requested in `{{player_a}}` and `{{player_b}}`
- [ ] `no-silent-outage`: run.log shows no long stretch of failed or empty model turns that the code checks did not already flag
## Write Validation Report
Write `{{results_dir}}/validation_report.json` in the evals-v2 shape with `arena-report`. It runs every code check above,
records one `step` entry per step you pass in, and one `checklist` entry per item you pass in (status, a colon, then a
one-line reason). Jetty computes the verdict from these checks.
```bash
arena-report --results {{results_dir}} --iterations 1 \
--steps '{"environment": "pass", "launch": "pass", "supervise": "pass", "render": "pass", "notes": "pass"}' \
--checklist '{"notes-explain-win": "pass: <why>", "players-match-request": "pass: <why>", "no-silent-outage": "pass: <why>"}'
```
Use `fail: <reason>` for any step or checklist item that did not hold, and `--iterations 2` if you used the retry in
Step 4. The command prints every check and the verdict. Do not hand-edit the file and do not invent a different format:
it must stay `{"version": 2, "verdict", "overall_passed", "iterations", "checks": [...]}`.
## Tips
- Cost is dominated by the two competing models (roughly $0.05 to $3 per game per model), not by this runbook's agent.
- `decisions.jsonl` rows contain the full text briefing each model saw, which is useful for debugging strange play.
- To watch the replay with full controls, open `replay.orarep` in an OpenRA build matching `runtime_config.json`.
## Changelog
- 2.0.0: Evals v2 shape (Code Checks, Checklist, Write Validation Report with `arena-report`); games end on a
`max_wall_minutes` real-time budget (default 30) instead of a 20-game-minute cap; GPT-5.6 Luna preset.
- 1.1.0: Supervision must be foreground-only: agents that moved polling into background tasks ended their session and
lost the game.
- 1.0.0: Initial version: one game per run, 20-game-minute cap, lockstep decisions every 8 game-seconds.
Improve: what changed between rounds
Every round's lessons go back into the runbook, the Improve step of the loop. The changelog is part of the runbook above; the biggest changes: v1.1.0 made supervision foreground-only after agents ended their sessions mid-game, and v2.0.0 moved to Jetty's evals-v2 checks and a 30-minute wall-clock budget per game.