evals
Run evals on Jetty
Reproduce

Run the OpenRA arena on your own Jetty

Everything on this site came from one runbook, executed once per game on Jetty. The runbook drives a prebuilt sandbox snapshot with the engine, game data, virtual display and orchestrator already installed, so a game starts in seconds.

Download RUNBOOK.md Task workflow JSON Snapshot Dockerfile Source (orchestrator + engine patch)

1. What runs

Runbookopenra-arena-match v1.0.0 · agent claude-code · anthropic/claude-sonnet-4.6 via OpenRouter (supervises the game)
Sandbox snapshotopenra-arena · 4 vCPU, 8 GB · Debian bookworm, .NET 8, OpenRA (OpenRA-RL fork @ 9b271c1 + arena patch), Xvfb + Mesa, ffmpeg, Python 3.11
Per game20-minute cap · decisions every 8 game-seconds · map Singles · ≈20–45 min wall clock
Outputsresult.json, metrics.json, decisions.jsonl, replay.orarep, match.mp4, runtime_config.json, summary.md, validation_report.json
SecretsOPENROUTER_API_KEY as a collection environment variable

2. Launch games with the API

Set JETTY_TOKEN to an API key for your collection and COLLECTION to its name. The openra-arena snapshot is available to Jetty runbooks.

# 1. Create the task in your collection (once)
curl -s -X POST https://flows-api.jetty.io/api/v1/tasks/$COLLECTION \
  -H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" \
  -d "$(jq -n --rawfile rb RUNBOOK.md --slurpfile wf task-workflow.json \
        '{name: "openra-arena-match", workflow: ($wf[0] | .init_params.instruction = $rb)}')"

# 2. Your collection needs OPENROUTER_API_KEY (Settings → Environment, or:)
curl -s -X PATCH https://flows-api.jetty.io/api/v1/collections/$COLLECTION/environment \
  -H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" \
  --data-binary '{"environment_variables": {"OPENROUTER_API_KEY": "sk-or-..."}}'

# 3. Play a game: any OpenRouter model vs another, or vs OpenRA's AI
curl -s -X POST https://flows-api.jetty.io/api/v1/run/$COLLECTION/openra-arena-match \
  -H "Authorization: Bearer $JETTY_TOKEN" \
  -F 'init_params={"vars": {"player_a": "openrouter:anthropic/claude-sonnet-5|Claude Sonnet 5",
                              "player_b": "cpu:normal", "game_id": "my-first-game", "max_ticks": "30000"}}'

3. Players

Use a preset ID, openrouter:<model>|<Display name> for any OpenRouter chat model, or cpu:<level> for OpenRA's AI.

SpecNameModel
jevJev Routertypesafe/jev-router
sonnet5Claude Sonnet 5anthropic/claude-sonnet-5
gpt56terraGPT-5.6 Terraopenai/gpt-5.6-terra
cpu-beginnerOpenRA AI (Beginner)openra/beginner-ai
cpu-easyOpenRA AI (Easy)openra/easy-ai
cpu-mediumOpenRA AI (Medium)openra/medium-ai
cpu-normalOpenRA AI (Normal)openra/normal-ai

4. The runbook

---
version: "1.1.0"
evaluation: programmatic
agent: claude-code
model: anthropic/claude-sonnet-4.6
model_provider: openrouter
snapshot: openra-arena
primary_outputs:
  - match.mp4
  - summary.md
  - result.json
secrets:
  OPENROUTER_API_KEY:
    env: OPENROUTER_API_KEY
    description: "OpenRouter key used by the two competing models (and by this runbook's agent)"
    required: true
---

# OpenRA Arena: one LLM-vs-LLM Red Alert game — Agent Runbook

## Objective

Play **one full 1v1 game of Command & Conquer: Red Alert (OpenRA)** between two decision models and deliver the
replay, a 1080p recording, the full decision log, and scored metrics to `{{results_dir}}`.

Each player is an LLM (any OpenRouter model) or OpenRA's built-in AI. The game engine blocks every 8 seconds of game
time, sends both players their own fog-of-war view **at the same tick**, and resumes only when both have answered, so
model latency never costs game time. A game ends when one side loses all its buildings, or at `{{max_ticks}}` ticks
(25 ticks = 1 game-second; 30000 = 20 game-minutes), where the side with the lower score (assets value + value of
enemy stuff destroyed) surrenders.

The `openra-arena` snapshot ships everything prebuilt: the patched OpenRA engine at `/opt/openra`, Red Alert game
data at `/opt/openra-content`, the arena orchestrator at `/opt/decision-ra`, Xvfb + Mesa (software OpenGL) and
ffmpeg. The single entry point is `run-match`. You do not need to build or install anything.

---

## REQUIRED OUTPUT FILES (MANDATORY)

**You MUST write all of the following to `{{results_dir}}`. The task is NOT complete until every file exists and is
non-empty.** `run-match` writes all of them itself; your job is to launch it, supervise it, and verify.

| File | Description |
|------|-------------|
| `{{results_dir}}/result.json` | Winner, end reason, end tick, final per-player stats, LLM cost/latency/failure counts |
| `{{results_dir}}/metrics.json` | Normalized per-player metrics (value destroyed/lost, peak army, orders, latency, cost) |
| `{{results_dir}}/decisions.jsonl` | Every decision turn: the briefing each model saw, its reasoning, commands and their results |
| `{{results_dir}}/replay.orarep` | OpenRA replay (open it in OpenRA to watch the game with full controls) |
| `{{results_dir}}/match.mp4` | 1920x1080 recording of the whole game at 6x speed with the observer stats table |
| `{{results_dir}}/thumb.jpg` | Thumbnail from the recording |
| `{{results_dir}}/runtime_config.json` | Exact configuration: players, models, map, limits, engine commit + patch hash, host |
| `{{results_dir}}/summary.md` | Human-readable result table |
| `{{results_dir}}/validation_report.json` | Programmatic checks with `passed` and per-check results |

---

## Parameters

| Parameter | Template Variable | Default | Description |
|-----------|-------------------|---------|-------------|
| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |
| West player | `{{player_a}}` | `jev` | Player in the west slot (see *Player specs*) |
| East player | `{{player_b}}` | `sonnet5` | Player in the east slot |
| Game ID | `{{game_id}}` | `game` | Label for this game (used in logs and the result) |
| Max ticks | `{{max_ticks}}` | `30000` | Game-length cap in ticks (30000 = 20 game-minutes) |

### Player specs

| Spec | Meaning |
|------|---------|
| `jev`, `sonnet5`, `gpt56terra`, `gemini38f` | Roster presets (Jev Router, Claude Sonnet 5, GPT-5.6 Terra, Gemini 3.8 Flash) |
| `openrouter:<model>` or `openrouter:<model>\|<Display Name>` | Any OpenRouter chat model, e.g. `openrouter:qwen/qwen3.8-max\|Qwen 3.8 Max` |
| `cpu:beginner`, `cpu:easy`, `cpu:medium`, `cpu:normal` | OpenRA's built-in skirmish AI, weakest to strongest |

---

## Dependencies

| Dependency | Where | Notes |
|------------|-------|-------|
| OpenRA engine + arena bridge | `/opt/openra` | Built from `yxc20089/OpenRA@9b271c1` + `patches/openra-arena.patch` |
| Red Alert game data | `/opt/openra-content` | Freeware content (OpenRA quick-install bundle) |
| Orchestrator | `/opt/decision-ra` (`run-match` on PATH) | Python; calls OpenRouter |
| `OPENROUTER_API_KEY` | collection env var | Required unless both players are `cpu:*` |

---

## Step 1: Environment Setup

```bash
mkdir -p {{results_dir}}
test -x /usr/local/bin/run-match && test -f /opt/openra/bin/OpenRA.dll && test -d /opt/openra-content/ra/v2 \
  && echo "PASS: openra-arena snapshot present" || echo "FAIL: not running on the openra-arena snapshot"
test -n "${OPENROUTER_API_KEY:-}" && echo "PASS: OPENROUTER_API_KEY set" || echo "WARN: OPENROUTER_API_KEY missing (only cpu:* players will work)"
nproc; free -g | head -2
```

If the snapshot check fails, stop and write `validation_report.json` with `passed: false` and the reason. Do not
try to build OpenRA inside the sandbox.

## Step 2: Launch the Game (as a detached process)

A game takes roughly 20–45 minutes of wall-clock time, longer than a single shell call may block. Start it as a detached process with `nohup` (a plain shell command, not an agent background task)
and write its console log into the results directory:

```bash
cd {{results_dir}}
nohup run-match --a "{{player_a}}" --b "{{player_b}}" --game-id "{{game_id}}" \
  --max-ticks {{max_ticks}} --out {{results_dir}} > {{results_dir}}/run.log 2>&1 &
echo $! > /tmp/run-match.pid
sleep 60; tail -5 {{results_dir}}/run.log
```

Within a minute `run.log` should show `Starting ...` and then one line per decision turn:
`[game] tick  1201 | <A>: assets ... | <B>: assets ... | 6.2s`.

## Step 3: Supervise Until the Game Ends (foreground only)

> **CRITICAL:** the sandbox — and the game with it — is destroyed the moment your session ends. You MUST stay in this
> step until `run-match` has exited. Do NOT use background tasks, `run_in_background`, `tail -f`, cron or any other
> asynchronous watcher, and do NOT end your turn "to check back later": there is no later. Only run the blocking
> command below, in the foreground, again and again.

Run this exact command. It blocks for at most 9 minutes and prints the latest progress line:

```bash
PID=$(cat /tmp/run-match.pid)
for i in $(seq 1 54); do
  if ! kill -0 "$PID" 2>/dev/null; then echo "RUN-MATCH EXITED"; break; fi
  sleep 10
done
tail -2 {{results_dir}}/run.log
```

- If the output does not contain `RUN-MATCH EXITED`, run the same command again immediately. Expect 4–8 repetitions:
  a 20-minute game takes about 30–50 minutes of wall-clock time, then the video renders for a few minutes.
- Do not kill a game that is still logging turns (ticks advance at about 600–1500 per minute).
- Only when you have seen `RUN-MATCH EXITED` continue to Step 4.

## Step 4: Handle Failures (max 1 retry)

Read `{{results_dir}}/validation_report.json`.

- `has_winner: false` (engine crashed, or the orchestrator lost the connection): check `engine.log` and `run.log`,
  then relaunch Step 2 **once** with `--game-id "{{game_id}}-retry"`.
- `video_rendered` false but the game has a winner: rerun only the render:
  `cd /opt/decision-ra && DISPLAY=:99 .venv/bin/python -m arena.render /tmp/arena/{{game_id}}` then copy
  `match.mp4` and `thumb.jpg` into `{{results_dir}}`.
- Never replay a game that already has a winner just to get a different result.

## Step 5: Report

`run-match` already wrote `summary.md`. Append a short **Notes** section to it: anything unusual in `run.log`
(LLM errors, retries, truncated video) and one sentence on how the game was won, taken from the last few
`thoughts` entries in `decisions.jsonl` for each player.

## Step 6: Final Checklist (MANDATORY — do not skip)

### Verification Script

```bash
echo "=== FINAL OUTPUT VERIFICATION ==="
R="{{results_dir}}"
for f in result.json metrics.json decisions.jsonl replay.orarep match.mp4 thumb.jpg runtime_config.json summary.md validation_report.json; do
  if [ -s "$R/$f" ]; then echo "PASS: $f ($(wc -c < "$R/$f") bytes)"; else echo "FAIL: $f missing or empty"; fi
done
python3 -c "import json; r=json.load(open('$R/result.json')); assert r.get('winner'), 'no winner'; print('PASS: winner =', r['winner'], '|', r['end_reason'])" || echo "FAIL: result has no winner"
python3 -c "import json; v=json.load(open('$R/validation_report.json')); print('PASS' if v['passed'] else 'FAIL', 'validation', v['checks'])"
ffprobe -v error -show_entries format=duration -of csv=p=0 "$R/match.mp4" && echo "PASS: video decodes"
echo "=== VERIFICATION COMPLETE ==="
```

### Checklist

- [ ] The game ran on the `openra-arena` snapshot with the requested players
- [ ] `result.json` names a winner
- [ ] `match.mp4` decodes and `validation_report.json` reports `video_complete: true`
- [ ] `decisions.jsonl` has turns for every LLM player
- [ ] `summary.md` includes the Notes section
- [ ] Verification script printed PASS for every line

**If ANY item fails, go back to Step 4. Do NOT finish until all items pass or the single retry is used up; if it is,
leave `validation_report.json` as written and explain in `summary.md`.**

---

## Changelog

| Version | Change |
|---------|--------|
| 1.1.0 | Supervision must be foreground-only: agents that moved polling into background tasks ended their session and lost the game. |
| 1.0.0 | Initial version: one game per run, 20-minute cap, lockstep decisions every 8 game-seconds. |

## Tips

- Cost is dominated by the two competing models (roughly $0.05–$3 per game per model), not by this runbook's agent.
- `decisions.jsonl` rows contain the full text briefing each model saw, useful for debugging strange play.
- To watch the replay with full controls, install OpenRA and open `replay.orarep` (engine build must match; see
  `runtime_config.json`).

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.