kleene@sql

relational algebra for recursive model calls

Write declarative SQL. Kleene compiles joins, recursion, predicates and aggregation into an execution graph of language-model calls, recursive sub-sessions and tool calls. It plans and costs that graph before spending a token, runs it, and lets you query the trace with the same SQL. Watch it happen in a terminal UI.

kleene — ~/notes — 80×24
$ kleene kleene 0.1.0 · one stream, a prompt, slash commands · /setup to add a key > Which project consumed the most hours in total? The notes are in demos/oolong/corpus.txt ── turn 1 ─────────────────────────────────────────── 0 calls · $0.0000 No context table was given, so the notes come in through the workspace tools. CREATE TABLE ctx AS SELECT c.ordinal, c.text FROM read('demos/oolong/corpus.txt') r CROSS JOIN LATERAL chunks(r.text, 200) c; SELECT COUNT(*), MIN(ordinal), MAX(ordinal) FROM ctx; 60 | 0 | 59 1 row ── turn 2 ────────────────────────────────────────── 67 calls · $0.1100 Only some notes mention hours. A cheap proxy filters those, then a worker extracts project and hours as JSON. CREATE TABLE hours AS SELECT ordinal, llm_json('Extract project and hours from: ' || text, '{"project":"string","hours":"number"}') AS h FROM ctx WHERE mentions_hours(text); 28 rows ↳ worker cc3970 d1 Resolve which project 'the migration' refers to in notes 14, 22 and 41 │ ── turn 1 ──────────────────────────────────── 0 calls · $0.0000 │ SELECT ordinal, text FROM ctx WHERE ordinal IN (14, 22, 41); │ FINAL FROM (SELECT 'Osprey' AS project); │ ■ final · 1 turns · 2 calls · 3100 tok · $0.0040 ── turn 3 ────────────────────────────────────────────────────── writing The hours table has 28 rows across six projects. Summing per project and picking the top one: FINAL FROM (SELECT h.project, SUM(h.hours) AS total_hours FROM hours GROUP BY 1 ORDER BY 2 DESC LIMIT 1); 1 row >

A session: the model writes CallSQL turn by turn, every statement shows its rows and its cost, child sessions nest under the turn that opened them. Every call is memoised, so the second run is free.

A planner for calls

Predicates that call a model have a cost the optimizer can see. Join order, conjunct order, semi-joins for EXISTS, cascades through cheap proxies and beam-limited recursion are rewrite rules with a cost model. EXPLAIN shows the estimate before the spend.

Budgets as semantics

Calls, tokens, dollars, depth and wall clock are dimensions of a budget that child sessions inherit and slice. A statement the remaining budget cannot pay for is refused, with its plan.

Learning as tables

What the harness learns (playbook SQL, ratings, generator dials, sampled selectivities) is rows in DuckDB: inspectable, replay-gated and revertible. Query the trace with the same SQL you ran.

Recursive language models

A recursive language model (RLM) is a model that does not read its context. The root model sees only the task; the context sits in a variable it can peek at, grep, partition and hand to sub-calls, and a sub-call can be another model with the same powers one level down. Zhang, Kraska and Khattab (arXiv:2512.24601) showed that a root with a REPL and llm_query / rlm_query handles contexts far past the window: on OOLONG at 132k tokens an RLM beat GPT-5 by 114% at cost parity, and on BrowseComp-Plus it held 100% at a thousand documents. The strategies the model discovers on its own are peek, grep, partition-and-map, summarise, and solve in code.

Kleene's bet is that those strategies are queries. A peek is a LIMIT, grep is a table function, partition-and-map is GROUP BY and a LATERAL rlm(...), summarise is an aggregate, and the answer is a relation. Once they are written as SQL the engine can plan them, price them, memoise them, run them concurrently and refuse the ones the budget cannot pay for. Depth is a budget dimension too: reproductions found that depth two degrades (format collapse, runaway sub-calls), so the harness caps it and shows every child under the turn that opened it.

Sources and the rest of the reading are in the research digest.

rlm in one statement
-- the root never reads the corpus; sixty notes become six children CREATE TABLE parts AS SELECT ordinal / 10 AS part, string_agg(text, '\n' ORDER BY ordinal) AS chunk FROM ctx GROUP BY 1; SELECT r.answer, r.session FROM parts CROSS JOIN LATERAL rlm('Which project logged the most hours here?', chunk) r; -- each child is the same loop at depth 1 with a slice of the budget; -- its FINAL comes back as the row, its trace is queryable afterwards: SELECT depth, role, outcome, calls, dollars FROM trace_sessions ORDER BY started_at;

Why build this

A learning journey

Kleene started as a question worth a few evenings: if a recursive language model's strategies are peek, grep, partition and map, what happens when the model writes them as queries and an engine plans them? Every milestone answered one part of that, from a differential-tested relational core to a planner that prices model calls. The design and its dead ends are in the plan, and the honest gaps in the write-up.

Shared with the community

Everything is in the open: MIT licensed Rust, the dialect reference, the architecture, the benchmark packs with their fixtures and plots, and a technical report. If the idea is right, others should be able to reproduce it, break it and build on it; if it is wrong, the trace tables will say so in SQL.

Nobody uses SQL like this yet

SQL is usually where the data waits for the program. Since SQL:1999 added WITH RECURSIVE it has been Turing complete (Andrew Gierth ran a cyclic tag system in one PostgreSQL query in 2010), yet almost nobody writes programs in it. Kleene does: the query is the program, a model call is a function, a child session is a recursive call, and the executor runs it to a fixpoint. The pay-off is that forty years of query optimisation apply to agent workflows, and EXPLAIN tells you which complexity fragment you are in (CQ, FO or REC) before anything runs.

One query, priced before it runs

VERIFY is one model call per distinct candidate. The NOT EXISTS is an anti-semi-join that stops on the first refuting counterexample. The planner reorders the predicates by cost and tells you how many calls that is before you spend them.

candidates.sql
SELECT candidate FROM possibilities WHERE VERIFY(candidate) AND NOT EXISTS ( SELECT 1 FROM counterexamples ce WHERE REFUTE(candidate, ce.text) );

Delegation is a lateral join

rlm(question, chunk) opens one child session per input row, with a slice of the budget and its own table namespace. spawn('reviewer', task) opens a child with that agent's tools. A child's FINAL comes back to the parent as a row.

partition-and-map.sql
CREATE TABLE parts AS SELECT ordinal / 10 AS part, string_agg(text, '\n') AS chunk FROM ctx GROUP BY 1; SELECT r.answer, r.detail FROM parts CROSS JOIN LATERAL rlm('Which project logged the most hours?', chunk) r;

Install

sh
# curl: detects OS and architecture, verifies the SHA-256, installs to ~/.local/bin $ curl -fsSL https://raw.githubusercontent.com/MarcusElwin/kleene/main/install.sh | sh # Homebrew $ brew install MarcusElwin/kleene/kleene # From source (compiles DuckDB the first time, about ten minutes) $ cargo install --git https://github.com/MarcusElwin/kleene kleene # Check the engine without a model $ kleene repl -c "SELECT 1 + 1 AS two" 2 1 row

Prebuilt binaries cover macOS (Apple silicon, Intel) and Linux (x86_64, aarch64). Then kleene setup for a provider (Anthropic, OpenAI, or any compatible endpoint), or just export ANTHROPIC_API_KEY. The full reference is in the CLI docs.

Status

Pre-release. Every milestone of the plan is on main: the relational core with recursive CTEs by semi-naive evaluation, the call algebra with memo, batching and budgets, the RLM harness with rlm and spawn, the daemon and the Catppuccin TUI, the planner, the continual learning loop and the benchmark harness. Rust, model-agnostic, no SDK and no gateway required, MIT licensed.

[x] M0 scaffold   [x] M1 relational core   [x] M2 call algebra   [x] M3 RLM harness
[x] M4 daemon and TUI   [x] M5 planner   [x] M6 continual loop   [x] M7 benchmarks   [ ] v0.1.0 release