kleene@sql

Benchmarks

Every number on this page is computed from the per-task rows under plots/ (the output of kleene bench csv), so a new run that lands in the repository lands here. How the harness works is in the benchmark doc; the reading is in the write-up.

Does writing the work as SQL make a model solve more for less? Three modes on the same model, tools and budgets answer two comparisons: frozen is Kleene with the playbook off, learning is Kleene with the playbook on and adoption gated by replay, and plain is a tool-calling agent with no SQL. Learning against frozen isolates the playbook; frozen against plain isolates the abstraction.

model
mode
pack

Pass rate against cost

One panel per pack. Colour is the mode, a circle is Claude Opus 5.5 and a diamond is Claude Haiku 4.5. Hover or tap a mark for its numbers; click a mark to pin the tooltip.

Calls per task

The playbook's effect shows in calls. On the four first packs with Claude Opus 5.5, learning takes fewer calls per task than frozen on every pack; on legal the second half of the learning run took 7.5 calls a task where the first took 13.7. The plain agent answers each of these in one call, because every context fits in one prompt.

Every cell

packmodemodelsolved$ / tasktotal $calls / tasktokens / tasks / task

What the numbers say

The packs

packtaskswhat the model gets, what it must produceoraclestands in for
terminal6a workspace and an instruction (create a file, count lines, rename, sum a CSV column); files left behind and a one-row FINALshellTerminal-Bench
oolong-like20dated meeting notes naming projects and hours; one row per project with its totalexactOOLONG-style aggregation
finance-synthetic20a synthetic income statement over several years with distractor notes; one number (a sum, a growth rate, a ratio)numberFinanceBench
legal-synthetic20a synthetic contract with planted clause categories among near-misses; the set of categories presentexactCUAD, Harvey LAB
logbook-hard30about 1,650 log entries (30k tokens) in four phrasings with corrections; one of five aggregate questionsexactOOLONG at a size that cannot be read once
memo-rubric20a 40-section services agreement with planted terms, a rejected proposal and a superseding amendment; a short memojudgeHarvey LAB drafting
coding12a small Python project with visible tests; implement, extend, then fix a bug, in three-step episodesshell (hidden tests)SWE-bench

The synthetic packs are frozen generator output, so two people see the same tasks in the same order. The external datasets they imitate are not redistributed; OOLONG, Harvey LAB and contract redlining import on demand with bench import-*.

The five plots

What bench plot wrote beside each run's rows, for the models selected above.

Claude Opus 5.5, 27 Sep 2026. root on Opus 5.5 at high effort, worker and judge on Sonnet 5, proxy on Haiku 4.5. Rows: plots/evals.csv.

Learning curve, Claude Opus 5.5
Learning curve: task number against rolling-5 accuracy, one line per run; a learning line that rises while frozen stays flat is the playbook working
Cost parity, Claude Opus 5.5
Cost parity: cumulative dollars against cumulative accuracy; at equal spend, which mode is higher
Calls against difficulty, Claude Opus 5.5
Calls against difficulty: task rating against calls made, solved and failed as two series
Estimate accuracy, Claude Opus 5.5
Estimate accuracy: the planner's estimated calls against the calls a statement made; on the diagonal is exact
Plan space, Claude Opus 5.5
Plan space: relations joined against join orders considered, one point per statement

Claude Haiku 4.5, 1 Oct 2026. every solver alias on Haiku 4.5, judge on Sonnet 5.5; the sweep was stopped before logbook-hard finished. Rows: plots/haiku-2026-10-01/evals.csv.

Learning curve, Claude Haiku 4.5
Learning curve: task number against rolling-5 accuracy, one line per run; a learning line that rises while frozen stays flat is the playbook working
Cost parity, Claude Haiku 4.5
Cost parity: cumulative dollars against cumulative accuracy; at equal spend, which mode is higher
Calls against difficulty, Claude Haiku 4.5
Calls against difficulty: task rating against calls made, solved and failed as two series
Estimate accuracy, Claude Haiku 4.5
Estimate accuracy: the planner's estimated calls against the calls a statement made; on the diagonal is exact
Plan space, Claude Haiku 4.5
Plan space: relations joined against join orders considered, one point per statement

Reproduce

sh
# a pack, one mode, with a fresh store per mode $ kleene bench run tasks/oolong-like --mode frozen $ kleene bench run tasks/oolong-like --mode learning $ kleene bench run tasks/oolong-like --mode plain # the table, the rows and the plots $ kleene bench report $ kleene bench csv > plots/<model-date>/evals.csv $ kleene bench plot plots/<model-date> # the README's results table and Pareto plots from every committed CSV $ kleene bench results plots/evals.csv plots/haiku-2026-10-01/evals.csv --out plots/results --readme README.md

A learning run replays candidates against earlier tasks, so budget about $45 for a full sweep of the first four packs at replay_sample 3, and use bench run --resume <run-id> after an interruption. Every model call is memoised, so a second run of the same pack in the same store is free, which is also why a clean comparison wants a fresh store per mode.