evals
Run evals on Jetty
Evals by Jetty

Open evals you can rerun yourself.

Each eval here is a Jetty runbook: plain-language instructions, a pinned sandbox, and a trajectory for every run. Read the results, watch the runs, then run the same runbook on your own account with your own models.

Runbooks

Instructions, not scripts

A runbook tells an agent what done looks like and how to check it. Jetty runs it in a sandbox and keeps the receipts.

Receipts

Every run is inspectable

Inputs, the agent's full log, the artifacts it produced and the runtime config, linked from every result on this site.

Reproducible

Bring your own models

Copy the runbook into your collection, add a provider key, and rerun any eval against the models you care about.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.