---
version: "2.0.0"
evals_version: 2
evaluation: programmatic
agent: claude-code
model: anthropic/claude-sonnet-4.6
model_provider: openrouter
snapshot: openra-arena
primary_outputs:
  - match.mp4
  - summary.md
  - result.json
secrets:
  OPENROUTER_API_KEY:
    env: OPENROUTER_API_KEY
    description: "OpenRouter key used by the two competing models (and by this runbook's agent)"
    required: true
---

# OpenRA Arena: one LLM-vs-LLM Red Alert game

## Objective

Play **one full 1v1 game of Command & Conquer: Red Alert (OpenRA)** between two decision models and deliver the
replay, a 1080p recording, the full decision log, scored metrics and an evals-v2 validation report to `{{results_dir}}`.

Each player is an LLM (any OpenRouter model) or OpenRA's built-in AI. The engine pauses every 8 seconds of game time,
sends both players their own fog-of-war view **at the same tick**, and resumes only when both have answered, so model
latency never costs game time. A game ends when one side loses all its buildings, or after `{{max_wall_minutes}}`
minutes of real (wall-clock) time, when the side with the lower score (assets value + value of enemy units and
buildings destroyed) surrenders. `{{max_ticks}}` is only a safety cap on game time (25 ticks = 1 game-second).

The `openra-arena` snapshot ships everything prebuilt: the patched OpenRA engine at `/opt/openra`, Red Alert game data at
`/opt/openra-content`, the orchestrator at `/opt/decision-ra`, Xvfb + Mesa (software OpenGL) and ffmpeg. The entry
points are `run-match` (plays, records and scores a game) and `arena-report` (writes the validation report). Do not build
or install anything.

## REQUIRED OUTPUT FILES

- `{{results_dir}}/result.json`: winner, end reason, final per-player stats, model cost, latency and failure counts
- `{{results_dir}}/metrics.json`: per-player metrics (value destroyed and lost, peak army, orders, latency, cost)
- `{{results_dir}}/decisions.jsonl`: every decision turn (the briefing each model saw, its reasoning, its commands and their results)
- `{{results_dir}}/replay.orarep`: OpenRA replay
- `{{results_dir}}/match.mp4`: 1920x1080 recording of the whole game at 6x speed
- `{{results_dir}}/thumb.jpg`: thumbnail from the recording
- `{{results_dir}}/runtime_config.json`: players, models, map, limits, engine commit and patch hash, host
- `{{results_dir}}/summary.md`: result table plus the Notes you add in Step 5
- `{{results_dir}}/validation_report.json`: evals-v2 report (`"version": 2`, `checks[]`), written in the last section

## Parameters

| Parameter | Template Variable | Default | Description |
|-----------|-------------------|---------|-------------|
| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |
| West player | `{{player_a}}` | `jev` | Player in the west slot (see Player specs) |
| East player | `{{player_b}}` | `sonnet5` | Player in the east slot |
| Game ID | `{{game_id}}` | `game` | Label for this game (used in logs and the result) |
| Wall-clock limit | `{{max_wall_minutes}}` | `30` | Real-time minutes before the game is decided on score |
| Max ticks | `{{max_ticks}}` | `90000` | Safety cap on game time (90000 ticks = 60 game-minutes) |

Player specs:

- `jev`, `sonnet5`, `gpt56terra`, `gpt56luna`, `gemini38f`: roster presets (Jev Router, Claude Sonnet 5, GPT-5.6 Terra,
  GPT-5.6 Luna, Gemini 3.8 Flash)
- `openrouter:<model>` or `openrouter:<model>|<Display Name>`: any OpenRouter chat model
- `cpu:beginner`, `cpu:easy`, `cpu:medium`, `cpu:normal`: OpenRA's built-in skirmish AI, weakest to strongest

## Dependencies

- `openra-arena` snapshot: engine at `/opt/openra` (`yxc20089/OpenRA@9b271c1` + `patches/openra-arena.patch`), game data at
  `/opt/openra-content`, `run-match` and `arena-report` on PATH
- `OPENROUTER_API_KEY` collection environment variable (not needed when both players are `cpu:*`)

## Steps

### Step 1: Environment Setup

```bash
mkdir -p {{results_dir}}
test -x /usr/local/bin/run-match && test -x /usr/local/bin/arena-report && test -f /opt/openra/bin/OpenRA.dll \
  && test -d /opt/openra-content/ra/v2 && echo "PASS: openra-arena snapshot present" || echo "FAIL: not running on the openra-arena snapshot"
test -n "${OPENROUTER_API_KEY:-}" && echo "PASS: OPENROUTER_API_KEY set" || echo "WARN: OPENROUTER_API_KEY missing (only cpu:* players will work)"
nproc; free -g | head -2
```

If the snapshot check fails, stop: record `environment` as failed in the validation report (last section) and explain
in `summary.md`. Do not try to build OpenRA inside the sandbox.

### Step 2: Launch the game as a detached process

A game takes up to `{{max_wall_minutes}}` minutes plus a few minutes to render, longer than one shell call may block.
Start it with `nohup` (a plain shell command, not an agent background task):

```bash
cd {{results_dir}}
nohup run-match --a "{{player_a}}" --b "{{player_b}}" --game-id "{{game_id}}" \
  --max-wall-minutes {{max_wall_minutes}} --max-ticks {{max_ticks}} --out {{results_dir}} > {{results_dir}}/run.log 2>&1 &
echo $! > /tmp/run-match.pid
sleep 60; tail -5 {{results_dir}}/run.log
```

Within a minute `run.log` shows `Starting ...` and one line per decision turn:
`[game] tick  1201 | <A>: assets ... | <B>: assets ... | 6.2s`.

### Step 3: Supervise until the game ends (foreground only)

> **CRITICAL:** the sandbox, and the game with it, is destroyed the moment your session ends. Stay in this step until
> `run-match` has exited. Do NOT use background tasks, `run_in_background`, `tail -f`, cron or any other asynchronous
> watcher, and do NOT end your turn "to check back later": there is no later.

Run this exact command. It blocks for at most 9 minutes and prints the latest progress line:

```bash
PID=$(cat /tmp/run-match.pid)
for i in $(seq 1 54); do
  if ! kill -0 "$PID" 2>/dev/null; then echo "RUN-MATCH EXITED"; break; fi
  sleep 10
done
tail -2 {{results_dir}}/run.log
```

Repeat it until the output contains `RUN-MATCH EXITED` (expect 4 to 6 repetitions). Never kill a game that is still
logging turns.

### Step 4: Handle failures (max 1 retry)

Read the preliminary `{{results_dir}}/validation_report.json` that `run-match` wrote.

- `game-has-winner` failed (engine crash or lost connection): check `engine.log` and `run.log`, then relaunch Step 2
  **once** with `--game-id "{{game_id}}-retry"` and supervise it again.
- `video-decodes` failed but the game has a winner: rerun only the render with
  `cd /opt/decision-ra && DISPLAY=:99 .venv/bin/python -m arena.render /tmp/arena/{{game_id}}`, then copy `match.mp4`
  and `thumb.jpg` into `{{results_dir}}`.
- Never replay a game that already has a winner just to get a different result.

### Step 5: Notes

Append a `## Notes` section to `{{results_dir}}/summary.md`: anything unusual in `run.log` (model errors, retries,
truncated video) and two or three sentences on how the game was won, citing the last few `thoughts` of each player from
`decisions.jsonl`.

## Code Checks

Each check is a deterministic command; a failing check fails the run. `arena-report` runs the same checks and records
them, so run these for your own verification and then write the report as described below.

### outputs-exist — Every required output file exists and is non-empty

```bash
cd {{results_dir}} && for f in result.json metrics.json decisions.jsonl replay.orarep match.mp4 thumb.jpg runtime_config.json summary.md; do test -s "$f" || { echo "missing $f"; exit 1; }; done
```

### game-has-winner — The game finished with a winner

```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); assert r.get('winner'), 'no winner'; print(r['winner'], r['end_reason'])"
```

### decisions-logged — Every model player has decision turns in decisions.jsonl

```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); s={json.loads(l)['slot'] for l in open('{{results_dir}}/decisions.jsonl')}; need={k for k,v in r['providers'].items() if v!='cpu'}; assert need<=s, need-s"
```

### llm-turns-valid — At most 20% of each model's turns failed

```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); c,f=r['llm_calls'],r['llm_failures']; bad={s:(f[s],c[s]) for s in c if f[s]>max(3,0.2*c[s])}; assert not bad, bad"
```

### video-decodes — match.mp4 decodes

```bash
test "$(ffprobe -v error -show_entries format=duration -of csv=p=0 {{results_dir}}/match.mp4 | cut -d. -f1)" -gt 0
```

### video-complete — The recording covers the whole game

```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/render.json')); assert not r.get('truncated'), r"
```

### limits-respected — The game ended within the wall-clock budget

```bash
python3 -c "import json; r=json.load(open('{{results_dir}}/result.json')); assert r['wall_seconds']/60 <= {{max_wall_minutes}}+8, r['wall_seconds']"
```

## Checklist

Confirm each item by inspection; a failed item fails the run.

- [ ] `notes-explain-win`: summary.md has a Notes section explaining how the game was won, citing both models' own reasoning
- [ ] `players-match-request`: the players and models recorded in result.json are the ones requested in `{{player_a}}` and `{{player_b}}`
- [ ] `no-silent-outage`: run.log shows no long stretch of failed or empty model turns that the code checks did not already flag

## Write Validation Report

Write `{{results_dir}}/validation_report.json` in the evals-v2 shape with `arena-report`. It runs every code check above,
records one `step` entry per step you pass in, and one `checklist` entry per item you pass in (status, a colon, then a
one-line reason). Jetty computes the verdict from these checks.

```bash
arena-report --results {{results_dir}} --iterations 1 \
  --steps '{"environment": "pass", "launch": "pass", "supervise": "pass", "render": "pass", "notes": "pass"}' \
  --checklist '{"notes-explain-win": "pass: <why>", "players-match-request": "pass: <why>", "no-silent-outage": "pass: <why>"}'
```

Use `fail: <reason>` for any step or checklist item that did not hold, and `--iterations 2` if you used the retry in
Step 4. The command prints every check and the verdict. Do not hand-edit the file and do not invent a different format:
it must stay `{"version": 2, "verdict", "overall_passed", "iterations", "checks": [...]}`.

## Tips

- Cost is dominated by the two competing models (roughly $0.05 to $3 per game per model), not by this runbook's agent.
- `decisions.jsonl` rows contain the full text briefing each model saw, which is useful for debugging strange play.
- To watch the replay with full controls, open `replay.orarep` in an OpenRA build matching `runtime_config.json`.

## Changelog

- 2.0.0: Evals v2 shape (Code Checks, Checklist, Write Validation Report with `arena-report`); games end on a
  `max_wall_minutes` real-time budget (default 30) instead of a 20-game-minute cap; GPT-5.6 Luna preset.
- 1.1.0: Supervision must be foreground-only: agents that moved polling into background tasks ended their session and
  lost the game.
- 1.0.0: Initial version: one game per run, 20-game-minute cap, lockstep decisions every 8 game-seconds.
