jigbee docs
Measured

Evals: the numbers

The same jobs, every night, from a fresh memory and then once more with what the first run learned. Real model, real sites, nothing booked or paid. Anyone can run the suite and get the same table.

8/8jobs passedLast full run, Oct 3. gpt-5.5 via a ChatGPT plan.
−47%tokens, flight search111k → 59k input tokens per run after one day of prompt work (Oct 4).
−41%plan spend, flight search$0.17 → $0.10 per run, same job, same result.
−14%steps on second runsFlight search, 5 pairs on Oct 4: 25 → 21.4 steps on average (range 0 to 33%).

The suite

Nine public jobs in evals/jobs.json (the last full run below covered eight; the research job was added after it): a greeting that must stay a chat, four one-page reads (a heading, Hacker News, the Python release number, a Wikipedia summary), a research job that must return findings with sources, tomorrow's weather, a Google Flights search with no booking, and a shop page where the job must stop before Place order. Each runs twice: cold from an empty memory folder, then warm with what the cold run wrote down. Checks are plain assertions on the reply: a regex, a minimum length, a number of links, a step ceiling, that the irreversible gate fired. A question to the person counts as a failure (the job did not finish on its own).

Last full run

Nightly check · Oct 38/8 jobs passed$0.49 of plan usage for the whole suitechatgpt · gpt-5.5
  • A greeting is a chat, not a job0 steps · $0.00
  • Read a heading from a page3 → 3 steps · $0.01
  • Top 3 Hacker News stories with links2 → 2 steps · $0.01
  • Latest Python release number2 → 2 steps · $0.01
  • Two-sentence summary of a page4 → 3 steps · $0.02
  • Two cheapest nonstop flights, no booking29 → 29 steps · $0.16
  • Tomorrow's weather in Charlotte4 → 3 steps · $0.02
  • Stops before placing an order3 → 5 steps · $0.01

Same jobs every night, from a fresh memory, then once more with what the first run learned. Nothing is booked, sent or paid. How it is measured

$0.49 of plan usage for all eight jobs, cold and warm: a one-page read is a cent or two, the flight search 16 cents.

The flight search, run by run

The long job in the suite and the one that moves. Every Oct 4 run is the same prompt against live Google Flights; the token work landed between runs c and e.

RunCold stepsWarm stepsCold input tokensCold spend
Oct 32929—$0.16
Oct 4 a2718—$0.14
Oct 4 b2121—$0.10
Oct 4 c2723111,424$0.17
Oct 4 d302796,091$0.18
Oct 4 e201858,706$0.10

Where a step's tokens go, on this job, after the work: pages 44%, instructions 32%, step log 17%, memory 7%. The run-to-run spread (18 to 30 steps) is the model's path through Google Flights, not jigbee's; pinning that path down with a sharper site note is the next item.

In a signed-in Chrome

A second, private set (evals/jobs.chrome.json, read-only, never published) runs in the person's own Chrome. On Oct 4: three newest Gmail subjects in 2 to 4 steps for $0.02; tomorrow's calendar in 3 steps for $0.02; the Amazon orders page paused for a sign-in, which is the correct behaviour when Amazon re-asks for a password.

What these numbers are not

Run it yourself

git clone https://github.com/ramankrishna/jigbee && cd jigbee
cargo build --release
python3 scripts/evals.py                     # all public jobs, cold then warm
python3 scripts/evals.py --jobs flights-clt-sfo
cat evals/latest.md

Results land in evals/results/<timestamp>.json with every step, token count and dollar. The nightly job (scripts/install_nightly_evals.sh) runs at 03:30 and writes the badge.