evals
Run evals on Jetty
OpenRA arena · Command & Conquer: Red Alert · updated 2026-09-28

Can a language model run a war?

Frontier models command a full Red Alert army in real time: build a base, run an economy, scout, and attack through fog of war. Every game is a Jetty runbook run in an identical sandbox, with a replay, a video and every decision the model made.

Leader
GPT-5.6 Terra
6–2 in the round robin
Games
20
12 round robin · 8 vs built-in AI
Game time played
5.0 h
20-minute cap per game
Model spend
$21
OpenRouter, all games
Headline results
Round robin
GPT-5.6 Terra wins, 6–2

Elo 1553; 3–1 vs Jev Router, 3–1 vs Claude Sonnet 5.

Cost
Jev Router plays for $0.10 a game

about 5× cheaper than Claude Sonnet 5 ($0.49), and 4–4 in the round robin.

vs the machine
OpenRA's built-in AI went 8–0

against the champion, including Beginner, the weakest setting. LLM commanders are not close yet.

Stage 1 · Round robin

Leaderboard

Each pair of models played 4 games, alternating starting positions. Ranked by wins, then Elo (K=32, in game order). "Destroyed" and "Lost" are the resource value of units and buildings; K/D is the ratio of the two.

#ModelW–LWin %EloDestroyed / gameLost / gameK/D valuePeak armyLatencyCost / game
1 GPT-5.6 Terra
openai/gpt-5.6-terra
6–2 75% 1553 $13,400 $11,512 1.16 $4,825 3.7s $0.46
2 Jev Router
typesafe/jev-router
4–4 50% 1488 $10,619 $8,131 1.31 $3,144 5.0s $0.10
3 Claude Sonnet 5
anthropic/claude-sonnet-5
2–6 25% 1459 $12,012 $16,388 0.73 $3,256 4.7s $0.49
Profiles

How each model plays

Eight metrics per model, each scaled 0–100 against the best model on that metric (for Survival and Speed, lower loss and lower latency score higher). The shape tells you the style: a rusher spikes on Firepower and Decisiveness, a turtle on Economy and Survival.

All models

Click a name to hide or show it.

GPT-5.6 Terra

6–2 · Elo 1553

Jev Router

4–4 · Elo 1488

Claude Sonnet 5

2–6 · Elo 1459

Head to head

Who beats whom

Read across a row: that model's record against each column.

Row beat columnGPT-5.6 TerraJev RouterClaude Sonnet 5
GPT-5.6 Terra—3–13–1
Jev Router1–3—3–1
Claude Sonnet 51–31–3—

Army value over the game

Mean value of each model's army at each game minute, across its round-robin games.

Stage 2 · Champion vs the machine

GPT-5.6 Terra vs OpenRA's built-in AI

The round-robin winner played OpenRA's own skirmish AI at every difficulty, two games per level (one from each side of the map). The built-in AI is a scripted bot with perfect micro-management and no thinking time, a useful yardstick for how far LLM commanders still have to go.

Value destroyed per game, by AI difficulty

Mean over the games at each level. Hover a bar for the record.

DifficultyRecordGPT-5.6 Terra destroyedAI destroyedGames
Beginner0–2$7,075$27,050G1 G2
Easy0–2$18,900$44,950G1 G2
Medium0–2$3,750$22,500G1 G2
Normal0–2$8,275$42,100G1 G2
Every game

Watch the games

Each game page has the full 1080p recording (6× speed) next to both models' reasoning, synced to the video, plus the replay file and the Jetty trajectory.

Claude Sonnet 5 vs GPT-5.6 Terra · GPT-5.6 Terra destroyed the enemy base · Watch with the decision timeline →

Round robin

Champion vs OpenRA AI

Pilot cup (local, before the Jetty runs: 14-minute cap, Gemini 3.8 Flash included)
FAQ

OpenRA arena: questions and answers

Which LLM is best at Command & Conquer: Red Alert?

In the OpenRA arena round robin (updated 2026-09-28), GPT-5.6 Terra ranked first with a 6–2 record and Elo 1553. Full standings: GPT-5.6 Terra 6–2; Jev Router 4–4; Claude Sonnet 5 2–6. Each pair played 4 games, alternating starting positions, on the map Singles with a 20-minute game-time cap.

Can an LLM beat OpenRA's built-in AI?

Not yet. The round-robin champion, GPT-5.6 Terra, went 0–8 against OpenRA's scripted skirmish AI across all four difficulty levels (Beginner 0–2, Easy 0–2, Medium 0–2, Normal 0–2), including the weakest, Beginner.

How much does it cost for an LLM to play a game of Red Alert?

Average model spend per game through OpenRouter: GPT-5.6 Terra $0.46; Jev Router $0.10; Claude Sonnet 5 $0.49. Jev Router was the cheapest. A model is called once every 8 game-seconds, roughly 50 to 150 calls per game.

How does the OpenRA arena keep games fair between models?

The engine pauses every 8 game-seconds and gives both models their own fog-of-war view at the same tick, then resumes only when both have answered, so a slower model never loses game time. Every model gets the same prompt, commands, map and time limit, and starting positions alternate between games.

Can I run the OpenRA LLM arena myself?

Yes. The arena is a Jetty runbook (openra-arena-match) that runs on a prebuilt sandbox snapshot (openra-arena). Create the task in your Jetty collection, add an OPENROUTER_API_KEY, and launch one API call per game with any OpenRouter model, or cpu:<level> for OpenRA's AI, as each player. Instructions: evaljetty.com/openra/run.html.

What data is published for each game?

Every game has a 1080p video, the OpenRA replay file, the full decision log (what each model saw, its reasoning and its commands), result and metrics JSON, and the exact runtime configuration. Aggregates: evaljetty.com/openra/data/leaderboard.json, leaderboard.csv and games.json.

Method
How the arena works

Lockstep decisions every 8 game-seconds, fog of war, a 20-minute cap, identical prompts and tools for every model.

Read the method →
Reproduce
Run it on your own Jetty

One runbook, one prebuilt sandbox snapshot, one API call per game. Bring any OpenRouter model.

Runbook & config → All runs →
Case study
How we built it on Jetty

One snapshot, one runbook, parallel fan-out, and what broke along the way.

Read the case study →

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.