Benchmarks
Every number on this page is computed from the per-task rows under plots/ (the output of kleene bench csv), so a new run that lands in the repository lands here. How the harness works is in the benchmark doc; the reading is in the write-up.
Does writing the work as SQL make a model solve more for less? Three modes on the same model, tools and budgets answer two comparisons: frozen is Kleene with the playbook off, learning is Kleene with the playbook on and adoption gated by replay, and plain is a tool-calling agent with no SQL. Learning against frozen isolates the playbook; frozen against plain isolates the abstraction.
Pass rate against cost
One panel per pack. Colour is the mode, a circle is Claude Opus 5.5 and a diamond is Claude Haiku 4.5. Hover or tap a mark for its numbers; click a mark to pin the tooltip.
Calls per task
The playbook's effect shows in calls. On the four first packs with Claude Opus 5.5, learning takes fewer calls per task than frozen on every pack; on legal the second half of the learning run took 7.5 calls a task where the first took 13.7. The plain agent answers each of these in one call, because every context fits in one prompt.
Every cell
| pack | mode | model | solved | $ / task | total $ | calls / task | tokens / task | s / task |
|---|
What the numbers say
- The first four packs are saturated. Opus 5.5 solves every task of terminal, oolong-like, finance-synthetic and legal-synthetic in every mode, so they compare cost only. The plain agent wins there: every context fits in one prompt, and the SQL abstraction pays for a catalog, a planner and rendered results on every turn.
- The playbook cuts calls. Learning against frozen: terminal 2.2 to 1.3, oolong 4.9 to 3.9, finance 3.5 to 2.4, legal 12.7 to 10.6 calls a task. Only oolong and legal adopted entries; on terminal and finance the gate rejected every candidate, so their difference is run-to-run variance.
- The gate is the cost. Each candidate is replayed on three earlier tasks in both arms. Those replays are recorded in
playbook_evals, not here: about $34 of the Opus run's $44 went to the gate. At twenty tasks a pack it spends more than the playbook saves. - The harder packs separate on accuracy. On Haiku 4.5, coding solves 2 of 12 steps in either Kleene mode, always the first step of an episode. On memo-rubric learning beat frozen on both axes, 11/20 against 8/20 at 3.8 against 20 calls a task, though part of the frozen cost was a schema failure mode the adapter has since fixed, and learning shares the frozen run's memo.
- Not run yet: logbook-hard in any mode (a frozen task took about 500 calls at Haiku's pace), coding in plain mode, and the harder packs on Opus 5.5. The caveats in full are in write-up 4.1.
The packs
| pack | tasks | what the model gets, what it must produce | oracle | stands in for |
|---|---|---|---|---|
terminal | 6 | a workspace and an instruction (create a file, count lines, rename, sum a CSV column); files left behind and a one-row FINAL | shell | Terminal-Bench |
oolong-like | 20 | dated meeting notes naming projects and hours; one row per project with its total | exact | OOLONG-style aggregation |
finance-synthetic | 20 | a synthetic income statement over several years with distractor notes; one number (a sum, a growth rate, a ratio) | number | FinanceBench |
legal-synthetic | 20 | a synthetic contract with planted clause categories among near-misses; the set of categories present | exact | CUAD, Harvey LAB |
logbook-hard | 30 | about 1,650 log entries (30k tokens) in four phrasings with corrections; one of five aggregate questions | exact | OOLONG at a size that cannot be read once |
memo-rubric | 20 | a 40-section services agreement with planted terms, a rejected proposal and a superseding amendment; a short memo | judge | Harvey LAB drafting |
coding | 12 | a small Python project with visible tests; implement, extend, then fix a bug, in three-step episodes | shell (hidden tests) | SWE-bench |
The synthetic packs are frozen generator output, so two people see the same tasks in the same order. The external datasets they imitate are not redistributed; OOLONG, Harvey LAB and contract redlining import on demand with bench import-*.
The five plots
What bench plot wrote beside each run's rows, for the models selected above.
Claude Opus 5.5, 27 Sep 2026. root on Opus 5.5 at high effort, worker and judge on Sonnet 5, proxy on Haiku 4.5. Rows: plots/evals.csv.
Claude Haiku 4.5, 1 Oct 2026. every solver alias on Haiku 4.5, judge on Sonnet 5.5; the sweep was stopped before logbook-hard finished. Rows: plots/haiku-2026-10-01/evals.csv.
Reproduce
A learning run replays candidates against earlier tasks, so budget about $45 for a full sweep of the first four packs at replay_sample 3, and use bench run --resume <run-id> after an interruption. Every model call is memoised, so a second run of the same pack in the same store is free, which is also why a clean comparison wants a fresh store per mode.