---
title: "Evals by Jetty: a launchpad for agent benchmarks and evals"
url: https://evaljetty.com/
description: "Ground your team in benchmarks you care about. Worked examples of how to author, run, investigate and improve agentic workloads on Jetty, with every runbook, run and evaluation published."
updated: 2026-09-28
publisher: Jetty (https://jetty.io)
---

A launchpad for agent benchmarks and evals

# Ground your team in benchmarks you care about.

Public agent benchmarks are great for providing a general level of performance, but they don’t know your context or your task. What matters is the task that you’re building and how it performs. This website provides worked examples of how you can author, run, investigate and improve agentic workloads.

[See the worked examples](https://evaljetty.com/#evals)[Start on Jetty](https://jetty.io/?utm_source=evaljetty&utm_medium=referral&utm_campaign=hub)

Your workload

[Author](https://evaljetty.com/openra/author.html)[Run](https://evaljetty.com/openra/runs/index.html)[Investigate](https://evaljetty.com/openra/investigations/index.html)[Improve](https://evaljetty.com/openra/author.html#changelog)

![](https://evaljetty.com/assets/pelicans/confused.svg)

Worked examples

## Two benchmarks, *built the way you'd build yours.*

Each one is a real task with its runbook, every run with its evaluation, and the questions the runs answer. Copy the runbook, swap in your task and your models.

[New · Real-time strategy

### OpenRA arena

Frontier LLMs command full Command & Conquer: Red Alert armies against each other and OpenRA's built-in AI. Economy, scouting and combat in real time; every decision logged, every game on video.

runbook · runs with evaluations · 3 investigations](https://evaljetty.com/openra/index.html)[Updated · Generative SVG

### Pelican-on-a-bicycle benchmark

Coding agents draw a pelican riding a bicycle in pure SVG, then critique and redraw it over several rounds. Now with current models and one external judge, so scores compare across agents.

runbook · 98 runs · 3 investigations](https://evaljetty.com/pelicans/index.html)

The loop

## Author, run, investigate, *then improve.*

General leaderboards tell you how a model does on average. This loop tells you how it does on your task.

[01 · Author

### Write down the task you care about

One plain-language runbook: the objective, the steps, and the code checks and checklist that decide pass or fail, in a pinned sandbox.](https://evaljetty.com/openra/author.html)[02 · Run

### Run it on your models

One run per case, fanned out in parallel. Each keeps its inputs, outputs, logs, cost and the evaluation Jetty computed.](https://evaljetty.com/openra/runs/index.html)[03 · Investigate & improve

### Answer your own questions

Which model fits your task, is it worth the cost, are the evals doing anything. Then fold the answers back into the runbook.](https://evaljetty.com/openra/investigations/index.html)

---
Source: https://evaljetty.com/ · Evals by Jetty · Run your own evals: https://jetty.io/?utm_source=evaljetty&utm_medium=llms
