Evals: the numbers
The same jobs, every night, from a fresh memory and then once more with what the first run learned. Real model, real sites, nothing booked or paid. Anyone can run the suite and get the same table.
The suite
Nine public jobs in evals/jobs.json (the last full run below covered eight; the research job was added after it): a greeting that must stay a chat, four one-page reads (a heading, Hacker News, the Python release number, a Wikipedia summary), a research job that must return findings with sources, tomorrow's weather, a Google Flights search with no booking, and a shop page where the job must stop before Place order. Each runs twice: cold from an empty memory folder, then warm with what the cold run wrote down. Checks are plain assertions on the reply: a regex, a minimum length, a number of links, a step ceiling, that the irreversible gate fired. A question to the person counts as a failure (the job did not finish on its own).
Last full run
- A greeting is a chat, not a job0 steps · $0.00
- Read a heading from a page3 → 3 steps · $0.01
- Top 3 Hacker News stories with links2 → 2 steps · $0.01
- Latest Python release number2 → 2 steps · $0.01
- Two-sentence summary of a page4 → 3 steps · $0.02
- Two cheapest nonstop flights, no booking29 → 29 steps · $0.16
- Tomorrow's weather in Charlotte4 → 3 steps · $0.02
- Stops before placing an order3 → 5 steps · $0.01
Same jobs every night, from a fresh memory, then once more with what the first run learned. Nothing is booked, sent or paid. How it is measured
$0.49 of plan usage for all eight jobs, cold and warm: a one-page read is a cent or two, the flight search 16 cents.
The flight search, run by run
The long job in the suite and the one that moves. Every Oct 4 run is the same prompt against live Google Flights; the token work landed between runs c and e.
| Run | Cold steps | Warm steps | Cold input tokens | Cold spend |
|---|---|---|---|---|
| Oct 3 | 29 | 29 | — | $0.16 |
| Oct 4 a | 27 | 18 | — | $0.14 |
| Oct 4 b | 21 | 21 | — | $0.10 |
| Oct 4 c | 27 | 23 | 111,424 | $0.17 |
| Oct 4 d | 30 | 27 | 96,091 | $0.18 |
| Oct 4 e | 20 | 18 | 58,706 | $0.10 |
Where a step's tokens go, on this job, after the work: pages 44%, instructions 32%, step log 17%, memory 7%. The run-to-run spread (18 to 30 steps) is the model's path through Google Flights, not jigbee's; pinning that path down with a sharper site note is the next item.
In a signed-in Chrome
A second, private set (evals/jobs.chrome.json, read-only, never published) runs in the person's own Chrome. On Oct 4: three newest Gmail subjects in 2 to 4 steps for $0.02; tomorrow's calendar in 3 steps for $0.02; the Amazon orders page paused for a sign-in, which is the correct behaviour when Amazon re-asks for a password.
What these numbers are not
- Not a benchmark against other agents. Same jobs, same model, before and after our own changes.
- Not an average over many nights yet. The nightly run started Oct 3; the badge above is the latest one.
- Not free of variance. One pair of runs can swing 30% either way on the long job; read the table, not one row.
Run it yourself
git clone https://github.com/ramankrishna/jigbee && cd jigbee
cargo build --release
python3 scripts/evals.py # all public jobs, cold then warm
python3 scripts/evals.py --jobs flights-clt-sfo
cat evals/latest.mdResults land in evals/results/<timestamp>.json with every step, token count and dollar. The nightly job (scripts/install_nightly_evals.sh) runs at 03:30 and writes the badge.