{
  "init_params": {
    "agent": "claude-code",
    "model": "anthropic/claude-sonnet-4.6",
    "snapshot": "openra-arena",
    "instruction": "---\nversion: \"2.0.0\"\nevals_version: 2\nevaluation: programmatic\nagent: claude-code\nmodel: anthropic/claude-sonnet-4.6\nmodel_provider: openrouter\nsnapshot: openra-arena\nprimary_outputs:\n  - match.mp4\n  - summary.md\n  - result.json\nsecrets:\n  OPENROUTER_API_KEY:\n    env: OPENROUTER_API_KEY\n    description: \"OpenRouter key used by the two competing models (and by this runbook's agent)\"\n    required: true\n---\n\n# OpenRA Arena: one LLM-vs-LLM Red Alert game\n\n## Objective\n\nPlay **one full 1v1 game of Command & Conquer: Red Alert (OpenRA)** between two decision models and deliver the\nreplay, a 1080p recording, the full decision log, scored metrics and an evals-v2 validation report to `{{results_dir}}`.\n\nEach player is an LLM (any OpenRouter model) or OpenRA's built-in AI. The engine pauses every 8 seconds of game time,\nsends both players their own fog-of-war view **at the same tick**, and resumes only when both have answered, so model\nlatency never costs game time. A game ends when one side loses all its buildings, or after `{{max_wall_minutes}}`\nminutes of real (wall-clock) time, when the side with the lower score (assets value + value of enemy units and\nbuildings destroyed) surrenders. `{{max_ticks}}` is only a safety cap on game time (25 ticks = 1 game-second).\n\nThe `openra-arena` snapshot ships everything prebuilt: the patched OpenRA engine at `/opt/openra`, Red Alert game data at\n`/opt/openra-content`, the orchestrator at `/opt/decision-ra`, Xvfb + Mesa (software OpenGL) and ffmpeg. The entry\npoints are `run-match` (plays, records and scores a game) and `arena-report` (writes the validation report). Do not build\nor install anything.\n\n## REQUIRED OUTPUT FILES\n\n- `{{results_dir}}/result.json`: winner, end reason, final per-player stats, model cost, latency and failure counts\n- `{{results_dir}}/metrics.json`: per-player metrics (value destroyed and lost, peak army, orders, latency, cost)\n- `{{results_dir}}/decisions.jsonl`: every decision turn (the briefing each model saw, its reasoning, its commands and their results)\n- `{{results_dir}}/replay.orarep`: OpenRA replay\n- `{{results_dir}}/match.mp4`: 1920x1080 recording of the whole game at 6x speed\n- `{{results_dir}}/thumb.jpg`: thumbnail from the recording\n- `{{results_dir}}/runtime_config.json`: players, models, map, limits, engine commit and patch hash, host\n- `{{results_dir}}/summary.md`: result table plus the Notes you add in Step 5\n- `{{results_dir}}/validation_report.json`: evals-v2 report (`\"version\": 2`, `checks[]`), written in the last section\n\n## Parameters\n\n| Parameter | Template Variable | Default | Description |\n|-----------|-------------------|---------|-------------|\n| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |\n| West player | `{{player_a}}` | `jev` | Player in the west slot (see Player specs) |\n| East player | `{{player_b}}` | `sonnet5` | Player in the east slot |\n| Game ID | `{{game_id}}` | `game` | Label for this game (used in logs and the result) |\n| Wall-clock limit | `{{max_wall_minutes}}` | `30` | Real-time minutes before the game is decided on score |\n| Max ticks | `{{max_ticks}}` | `90000` | Safety cap on game time (90000 ticks = 60 game-minutes) |\n\nPlayer specs:\n\n- `jev`, `sonnet5`, `gpt56terra`, `gpt56luna`, `gemini38f`: roster presets (Jev Router, Claude Sonnet 5, GPT-5.6 Terra,\n  GPT-5.6 Luna, Gemini 3.8 Flash)\n- `openrouter:<model>` or `openrouter:<model>|<Display Name>`: any OpenRouter chat model\n- `cpu:beginner`, `cpu:easy`, `cpu:medium`, `cpu:normal`: OpenRA's built-in skirmish AI, weakest to strongest\n\n## Dependencies\n\n- `openra-arena` snapshot: engine at `/opt/openra` (`yxc20089/OpenRA@9b271c1` + `patches/openra-arena.patch`), game data at\n  `/opt/openra-content`, `run-match` and `arena-report` on PATH\n- `OPENROUTER_API_KEY` collection environment variable (not needed when both players are `cpu:*`)\n\n## Steps\n\n### Step 1: Environment Setup\n\n```bash\nmkdir -p {{results_dir}}\ntest -x /usr/local/bin/run-match && test -x /usr/local/bin/arena-report && test -f /opt/openra/bin/OpenRA.dll \\\n  && test -d /opt/openra-content/ra/v2 && echo \"PASS: openra-arena snapshot present\" || echo \"FAIL: not running on the openra-arena snapshot\"\ntest -n \"${OPENROUTER_API_KEY:-}\" && echo \"PASS: OPENROUTER_API_KEY set\" || echo \"WARN: OPENROUTER_API_KEY missing (only cpu:* players will work)\"\nnproc; free -g | head -2\n```\n\nIf the snapshot check fails, stop: record `environment` as failed in the validation report (last section) and explain\nin `summary.md`. Do not try to build OpenRA inside the sandbox.\n\n### Step 2: Launch the game as a detached process\n\nA game takes up to `{{max_wall_minutes}}` minutes plus a few minutes to render, longer than one shell call may block.\nStart it with `nohup` (a plain shell command, not an agent background task):\n\n```bash\ncd {{results_dir}}\nnohup run-match --a \"{{player_a}}\" --b \"{{player_b}}\" --game-id \"{{game_id}}\" \\\n  --max-wall-minutes {{max_wall_minutes}} --max-ticks {{max_ticks}} --out {{results_dir}} > {{results_dir}}/run.log 2>&1 &\necho $! > /tmp/run-match.pid\nsleep 60; tail -5 {{results_dir}}/run.log\n```\n\nWithin a minute `run.log` shows `Starting ...` and one line per decision turn:\n`[game] tick  1201 | <A>: assets ... | <B>: assets ... | 6.2s`.\n\n### Step 3: Supervise until the game ends (foreground only)\n\n> **CRITICAL:** the sandbox, and the game with it, is destroyed the moment your session ends. Stay in this step until\n> `run-match` has exited. Do NOT use background tasks, `run_in_background`, `tail -f`, cron or any other asynchronous\n> watcher, and do NOT end your turn \"to check back later\": there is no later.\n\nRun this exact command. It blocks for at most 9 minutes and prints the latest progress line:\n\n```bash\nPID=$(cat /tmp/run-match.pid)\nfor i in $(seq 1 54); do\n  if ! kill -0 \"$PID\" 2>/dev/null; then echo \"RUN-MATCH EXITED\"; break; fi\n  sleep 10\ndone\ntail -2 {{results_dir}}/run.log\n```\n\nRepeat it until the output contains `RUN-MATCH EXITED` (expect 4 to 6 repetitions). Never kill a game that is still\nlogging turns.\n\n### Step 4: Handle failures (max 1 retry)\n\nRead the preliminary `{{results_dir}}/validation_report.json` that `run-match` wrote.\n\n- `game-has-winner` failed (engine crash or lost connection): check `engine.log` and `run.log`, then relaunch Step 2\n  **once** with `--game-id \"{{game_id}}-retry\"` and supervise it again.\n- `video-decodes` failed but the game has a winner: rerun only the render with\n  `cd /opt/decision-ra && DISPLAY=:99 .venv/bin/python -m arena.render /tmp/arena/{{game_id}}`, then copy `match.mp4`\n  and `thumb.jpg` into `{{results_dir}}`.\n- Never replay a game that already has a winner just to get a different result.\n\n### Step 5: Notes\n\nAppend a `## Notes` section to `{{results_dir}}/summary.md`: anything unusual in `run.log` (model errors, retries,\ntruncated video) and two or three sentences on how the game was won, citing the last few `thoughts` of each player from\n`decisions.jsonl`.\n\n## Code Checks\n\nEach check is a deterministic command; a failing check fails the run. `arena-report` runs the same checks and records\nthem, so run these for your own verification and then write the report as described below.\n\n### outputs-exist \u2014 Every required output file exists and is non-empty\n\n```bash\ncd {{results_dir}} && for f in result.json metrics.json decisions.jsonl replay.orarep match.mp4 thumb.jpg runtime_config.json summary.md; do test -s \"$f\" || { echo \"missing $f\"; exit 1; }; done\n```\n\n### game-has-winner \u2014 The game finished with a winner\n\n```bash\npython3 -c \"import json; r=json.load(open('{{results_dir}}/result.json')); assert r.get('winner'), 'no winner'; print(r['winner'], r['end_reason'])\"\n```\n\n### decisions-logged \u2014 Every model player has decision turns in decisions.jsonl\n\n```bash\npython3 -c \"import json; r=json.load(open('{{results_dir}}/result.json')); s={json.loads(l)['slot'] for l in open('{{results_dir}}/decisions.jsonl')}; need={k for k,v in r['providers'].items() if v!='cpu'}; assert need<=s, need-s\"\n```\n\n### llm-turns-valid \u2014 At most 20% of each model's turns failed\n\n```bash\npython3 -c \"import json; r=json.load(open('{{results_dir}}/result.json')); c,f=r['llm_calls'],r['llm_failures']; bad={s:(f[s],c[s]) for s in c if f[s]>max(3,0.2*c[s])}; assert not bad, bad\"\n```\n\n### video-decodes \u2014 match.mp4 decodes\n\n```bash\ntest \"$(ffprobe -v error -show_entries format=duration -of csv=p=0 {{results_dir}}/match.mp4 | cut -d. -f1)\" -gt 0\n```\n\n### video-complete \u2014 The recording covers the whole game\n\n```bash\npython3 -c \"import json; r=json.load(open('{{results_dir}}/render.json')); assert not r.get('truncated'), r\"\n```\n\n### limits-respected \u2014 The game ended within the wall-clock budget\n\n```bash\npython3 -c \"import json; r=json.load(open('{{results_dir}}/result.json')); assert r['wall_seconds']/60 <= {{max_wall_minutes}}+8, r['wall_seconds']\"\n```\n\n## Checklist\n\nConfirm each item by inspection; a failed item fails the run.\n\n- [ ] `notes-explain-win`: summary.md has a Notes section explaining how the game was won, citing both models' own reasoning\n- [ ] `players-match-request`: the players and models recorded in result.json are the ones requested in `{{player_a}}` and `{{player_b}}`\n- [ ] `no-silent-outage`: run.log shows no long stretch of failed or empty model turns that the code checks did not already flag\n\n## Write Validation Report\n\nWrite `{{results_dir}}/validation_report.json` in the evals-v2 shape with `arena-report`. It runs every code check above,\nrecords one `step` entry per step you pass in, and one `checklist` entry per item you pass in (status, a colon, then a\none-line reason). Jetty computes the verdict from these checks.\n\n```bash\narena-report --results {{results_dir}} --iterations 1 \\\n  --steps '{\"environment\": \"pass\", \"launch\": \"pass\", \"supervise\": \"pass\", \"render\": \"pass\", \"notes\": \"pass\"}' \\\n  --checklist '{\"notes-explain-win\": \"pass: <why>\", \"players-match-request\": \"pass: <why>\", \"no-silent-outage\": \"pass: <why>\"}'\n```\n\nUse `fail: <reason>` for any step or checklist item that did not hold, and `--iterations 2` if you used the retry in\nStep 4. The command prints every check and the verdict. Do not hand-edit the file and do not invent a different format:\nit must stay `{\"version\": 2, \"verdict\", \"overall_passed\", \"iterations\", \"checks\": [...]}`.\n\n## Tips\n\n- Cost is dominated by the two competing models (roughly $0.05 to $3 per game per model), not by this runbook's agent.\n- `decisions.jsonl` rows contain the full text briefing each model saw, which is useful for debugging strange play.\n- To watch the replay with full controls, open `replay.orarep` in an OpenRA build matching `runtime_config.json`.\n\n## Changelog\n\n- 2.0.0: Evals v2 shape (Code Checks, Checklist, Write Validation Report with `arena-report`); games end on a\n  `max_wall_minutes` real-time budget (default 30) instead of a 20-game-minute cap; GPT-5.6 Luna preset.\n- 1.1.0: Supervision must be foreground-only: agents that moved polling into background tasks ended their session and\n  lost the game.\n- 1.0.0: Initial version: one game per run, 20-game-minute cap, lockstep decisions every 8 game-seconds.\n",
    "vars": {
      "results_dir": "/app/results",
      "player_a": "jev",
      "player_b": "sonnet5",
      "game_id": "game",
      "max_wall_minutes": "30",
      "max_ticks": "90000"
    },
    "file_paths": [],
    "runbook_evals_version": 2,
    "model_provider": "openrouter"
  },
  "steps": [
    "run"
  ],
  "step_configs": {
    "run": {
      "activity": "runbook",
      "agent_path": "init_params.agent",
      "model_path": "init_params.model",
      "snapshot_path": "init_params.snapshot",
      "instruction_path": "init_params.instruction",
      "template_variables_path": "init_params.vars",
      "files_path": "init_params.file_paths",
      "cpus": 4,
      "memory": "8G",
      "timeout_sec": 7200,
      "langfuse_tracing": true,
      "model_provider_path": "init_params.model_provider"
    }
  }
}