Technical overview
How a model's SQL becomes a planned, budgeted graph of calls. The long form is in Architecture, the dialect reference and the write-up.
The claim
An agent's work can be written as queries. Retrieval is a scan or a grep table function, judgement is a boolean call function in a WHERE, delegation is a lateral table call that opens a child session, search is a recursive CTE with a beam LIMIT, and the answer is a relation. Once the work is a query, three things fall out that a tool-calling loop cannot offer: a planner, budgets as query semantics, and learning as tables.
The algebra
Ordinary operators (σ, π, ⋈, ∪, γ, μ) plus a call kind on any operator that evaluates a call expression:
| Kind | Shape | Cost |
|---|---|---|
| λf | scalar map-call: llm_bool(...), a prompt function in a projection | one call per distinct input |
| κg | expand-call: CROSS JOIN LATERAL g(...) | output cardinality is the branching factor |
| ρd | delegation: rlm(...), spawn(...) | the child's whole plan |
| σllm, ⋈llm | a predicate or join condition containing calls | calls × selectivity, observed once the predicate has been seen |
Cost is calls × per-alias tokens and dollars. Eight rewrite rules work on it: cheap-first over filters and join conditions, memo dedupe, batching (BATCH n on a prompt function prices ceil(rows / n) calls), cascade through a cheap proxy, semi-join for EXISTS, beam-limited recursion, join ordering over call predicates, and budget refusal.
Illustrative numbers in the shape the renderer prints. EXPLAIN needs no provider; its output is part of the interface and tested, see kleene explain.
One turn of a session
A session is a loop. The harness sends the model a cached system prefix (the CallSQL rules, the catalog, the budget) plus the transcript; the model replies with CallSQL in one fence; the harness parses it, annotates every operator with the calls it implies, prices the plan, refuses it if the remaining budget cannot pay, executes it, renders the rows back into the transcript, and repeats until the model writes FINAL.
sequenceDiagram
autonumber
participant M as Model
participant H as Harness
participant S as kleene-sql
participant A as kleene-algebra
participant X as kleene-exec
participant L as LiveSink
participant D as DuckDB
H->>M: system prefix + transcript
M-->>H: CallSQL in one sql fence
H->>S: parse, validate, resolve names and types
S-->>H: LogicalPlan
H->>A: annotate with call kinds, run rewrite rules
A-->>H: CallPlan with rows, calls, tokens, dollars
alt estimate exceeds remaining budget
H-->>M: refused, with the plan
else within budget
H->>X: execute(plan)
loop every scalar, table, tool or child call
X->>L: call
L->>D: memo lookup
alt hit
D-->>L: cached result
else miss
L->>L: route, call provider or tool, price, charge budget
L->>D: memo store, TraceEvent
end
L-->>X: rows
end
X-->>H: batches
H->>D: persist session
H-->>M: rendered rows + footer
end
Note over M,H: repeat until FINAL Delegation: rlm and spawn
Child sessions are the same loop one level deeper, with a role, a budget slice and their own table namespace. rlm(q, ctx) opens one child per input row; spawn('reviewer', task) opens a child with that agent's tools and budget. A child's FINAL comes back to the parent as a row. Budgets have calls, tokens, dollars, depth and wall clock; a child gets the parent's remaining slice intersected with its role's budget, and its spending rolls up.
sequenceDiagram
participant R as Root session (depth 0)
participant P as Provider
participant C1 as Child 1 (depth 1, worker)
participant C2 as Child 2 (depth 1, worker)
R->>P: turn: peek at ctx, partition it
P-->>R: CREATE TABLE parts AS SELECT ...
R->>P: turn
P-->>R: SELECT * FROM parts CROSS JOIN LATERAL rlm(question, chunk)
par one child per row, concurrently
R->>C1: task + chunk, budget slice
C1->>P: turns until FINAL
C1-->>R: (answer, detail, session)
and
R->>C2: task + chunk, budget slice
C2->>P: turns until FINAL
C2-->>R: (answer, detail, session)
end
R->>P: turn: aggregate the children's answers
P-->>R: FINAL FROM (SELECT ...) Recursion with a beam
Recursive CTEs run by semi-naive evaluation: the recursive term sees only the previous round's delta. UNION terminates when a round adds nothing new. A trailing ORDER BY ... LIMIT k inside the recursive term keeps the best k new rows of each round, which is a beam of width k: search as a query.
Processes: CLI, daemon and TUI
kleene run, repl, explain, learn and bench run the harness in one process. The TUI talks to an engine daemon over a Unix socket and a JSONL protocol with cursors, so a client that reconnects resumes from where it left off and kleene attach can watch the same events headless.
sequenceDiagram
participant U as kleene (TUI)
participant Dm as kleene daemon
participant Hs as Harness + store + provider
U->>Dm: connect .kleene/daemon.sock (starts one if none listens)
U->>Dm: Subscribe { after: cursor }
Dm-->>U: Hello, replayed Events
U->>Dm: StartRun (a task typed at the prompt)
Dm->>Hs: run
Hs-->>Dm: streamed text, trace events
Dm-->>U: CallDelta ... TurnFinished
U->>Dm: /sql → Submit, /trace and /board → Query
Dm-->>U: Table
Dm-->>U: RunFinished Crates
A Rust workspace. Every crate depends on kleene-core and nothing depends on the binary.
flowchart TB
CLI[kleene · CLI binary]
TUI[kleene-tui · ratatui client, setup wizard]
DAEMON[kleene-daemon · Unix socket, JSONL protocol]
HARNESS[kleene-harness · sessions, turn loop, LiveSink, learn, bench]
EXEC[kleene-exec · operators, semi-naive recursion, concurrent calls]
ALGEBRA[kleene-algebra · call kinds, cost model, rules, EXPLAIN]
SQL[kleene-sql · sqlparser to LogicalPlan]
LLM[kleene-llm · Anthropic, OpenAI-compatible, router, replay]
TOOLS[kleene-tools · files, grep, shell, git, web_search]
STORE[kleene-store · DuckDB: tables, memo, trace, sessions]
TRACE[kleene-trace · TraceEvent, sinks]
CORE[kleene-core · Value, Schema, Catalog, Budget, CallKind]
CLI --> TUI
CLI --> DAEMON
CLI --> HARNESS
TUI --> DAEMON
DAEMON --> HARNESS
HARNESS --> EXEC
HARNESS --> LLM
HARNESS --> TOOLS
HARNESS --> STORE
HARNESS --> TRACE
EXEC --> ALGEBRA
ALGEBRA --> SQL
SQL -.-> CORE
EXEC -.-> CORE
LLM -.-> CORE
TOOLS -.-> CORE
STORE -.-> CORE
TRACE -.-> CORE What the benchmarks measured
Four packs ran against Claude Opus 5.5 on 27 September 2026 in learning, frozen and plain-agent modes. Every task was solved in every mode, so those packs compare cost only: the playbook cuts calls per task, the replay gate costs more than it saves at twenty tasks a pack, and the one-call tool-calling baseline is cheapest wherever the context fits in a prompt. The harder packs (coding, memo-rubric, logbook-hard) do separate on accuracy. The numbers, charts and the reading are on the benchmarks page; the harness is described in the benchmark doc and the write-up.
Read on
- CallSQL dialect: every statement and function the model can write.
- Architecture: the path of one statement, the store and trace tables, the daemon protocol, the planner, the testing strategy.
- Write-up: the claim, the algebra, what was measured and the honest gaps.
