kleene@sql

Architecture

How Kleene is put together: the crates, the path a statement takes from the model’s reply to rows in DuckDB, and the processes that run it. The design rationale is in PLAN.md; the language is in DIALECT.md; this document is the map of the code as built.

One paragraph

A session is a loop: the harness sends the model a cached system prefix (the CallSQL rules, the catalog of tables, functions, tools and agents, the budget) plus the transcript; the model replies with CallSQL inside one ```sql fence; the harness parses it, annotates every operator with the model and tool calls it implies, prices that plan, refuses it if the remaining budget cannot pay, executes it, renders the rows back into the transcript, and repeats until the model writes FINAL. Child sessions for rlm(...) and spawn(...) are the same loop one level deeper with a role, a budget slice and their own table namespace. Every table, memo entry and trace event lives in one DuckDB file, so the trace is queryable with the same SQL, and the learning loop’s state (tasks, ratings, playbook) is tables too.

Crates

kleene            CLI binary: run, resume, repl, explain, trace, tui, attach, daemon, learn, bench
├── kleene-tui    ratatui client over the daemon protocol
├── kleene-daemon Unix-socket server, JSONL protocol with cursors, client
└── kleene-harness
    │                sessions and the turn loop, LiveSink (memo → provider → budget → trace),
    │                Repl, prompt rendering, agent roles, the continual loop (learn/), bench
    ├── kleene-exec      operators, three-valued logic, semi-naive recursion, concurrent calls
    │   └── kleene-algebra   LogicalPlan → CallPlan: call kinds, cost model, rules, EXPLAIN
    │       └── kleene-sql   sqlparser → validated subset → name/type resolution → LogicalPlan
    ├── kleene-llm       Provider trait, Anthropic and OpenAI-compatible adapters, router, replay
    ├── kleene-tools     table functions and CALL tools with volatility labels, workspace jail
    ├── kleene-store     DuckDB: session tables, memo, trace tables, catalog metadata
    └── kleene-trace     TraceEvent, Tracer, sinks (store, fan-out)
kleene-core       shared interface types: Value, Batch, Schema, Catalog, Budget, CallKind, ids
kleene-difftest   proptest generator + DuckDB oracle for the relational core

Every crate depends on kleene-core and nothing else depends on the binary. kleene-core changes deliberately and first (see CLAUDE.md).

CrateOwnsKey types
kleene-coreThe vocabulary every crate sharesValue, DataType, Schema, Batch, Catalog, FunctionDef, CallKind, Volatility, Budget, BudgetUsage, RunId/SessionId/StatementId
kleene-sqlParsing and planning to a positional logical plan; every rejection is an SqlError with a hintplan_sql(), LogicalPlan, Statement, SqlError, render_error()
kleene-algebraCall kinds on operators, estimates, rewrite rules, plan search, EXPLAIN textannotate(), CallPlan, Rule, Rewrite, explain()
kleene-execExecuting a plan against a CallSinkexecute(), execute_statement(), CallSink, MemorySink, ExecError
kleene-llmTalking to models over the wire, no SDKsProvider, CompletionRequest, AnthropicProvider, OpenAiCompatProvider, RoutedProvider, RouterConfig, Pricing, ReplayProvider, RecordingProvider, provider_from_env()
kleene-toolsTools as table functions and statementsToolRegistry, standard_tools(), ToolContext, catalog_entries()
kleene-storeDuckDB behind a mutex on the blocking poolDuckDbStore, DuckDbTraceSink, MemoryStore
kleene-traceThe event modelTraceEvent, Tracer, TraceSink, FanoutSink
kleene-harnessSessions, the REPL, prompts, roles, learning, benchmarksHarness, HarnessConfig, Session, Repl, LiveSink, StoreSink, AgentRole, learn::Learn
kleene-daemonThe engine as a serverDaemon, Client, ClientRequest, ServerMessage, Cursor, EventLog
kleene-tuiThe terminal client (one stream, slash commands) and the setup wizardApp, Entry, Model, Theme, setup::SetupApp, headless()
kleeneThe CLImain.rs only
kleene-difftestProperty-based equivalence with DuckDBgenerator, run, compare

The path of one statement

model reply ──extract_sql──▶ CallSQL text
                                 │
                                 ▼
  kleene-sql   sqlparser 0.62 AST ─▶ subset validation ─▶ CallSQL extensions
                  (CREATE FUNCTION … AS PROMPT, CREATE AGENT, CALL, SET, FINAL, EXPLAIN)
                  ─▶ name and type resolution against the Catalog ─▶ LogicalPlan
                                 │
                                 ▼
  kleene-algebra  annotate(LogicalPlan, Catalog, Stats) ─▶ CallPlan
                     every operator carries: rows, calls, tokens, dollars, depth,
                     call kinds (λ scalar, κ table, ρ recursive, tool), fences
                     rules run to a fixpoint: cheap-first, semi-join, cascade,
                     beam recursion, join ordering (once, before the loop), fences
                     explain() renders the plan, the alternatives and the plan space
                                 │
                                 ▼
  harness Repl       budget check: estimate > remaining ⇒ refuse with the plan
                                 │
                                 ▼
  kleene-exec     execute(plan, ExecContext) ─▶ stream of Batches
                     materialising operators, semi-naive recursion with beams,
                     rows evaluated concurrently up to call_concurrency
                        │ scalar_call / table_call / tool_call / child session
                        ▼
  harness LiveSink   memo lookup ─▶ router picks a candidate ─▶ adapter over reqwest
                     ─▶ usage priced ─▶ budget charged ─▶ TraceEvent emitted
                     ─▶ memo stored   (rlm/spawn: ChildRunner opens a child session)
                                 │
                                 ▼
  harness render     rows (truncated to 20 × 200 chars by default) + a footer
                     (calls, tokens, dollars, remaining budget) back into the transcript

kleene repl and kleene explain run the same path without a model turn around it. EXPLAIN stops after annotation; EXPLAIN ANALYZE runs and prints actuals beside estimates.

The catalog

Catalog is the session’s view of the world: tables with declared CallSQL types, builtin and user functions with their CallKind and Volatility, tools, and declared agents. The planner resolves names against it, the annotator prices calls from it, the prompt renderer groups it by call kind and volatility for the model. A child session sees a restricted catalog (the role’s tools only), which is the security boundary for delegation.

Call kinds and volatility

Each function or tool is IMMUTABLE, STABLE or VOLATILE, and the planner trusts the label: rules move predicates freely across immutable and stable operators and never across a volatile one (a fence). Prompt functions default to immutable (same prompt, same answer, so the memo applies), SQL bodies to stable, shell bodies and every side-effecting tool to volatile. Volatile tools are only allowed under CALL, never in a FROM.

Sessions and delegation

Harness::run(task, context)
  └─ root Session (depth 0, role "root", full catalog, whole budget)
       turn 1: system prefix + task ─▶ model ─▶ SQL ─▶ Repl ─▶ rendered rows
       turn 2: … + transcript ─▶ …
       ├─ CROSS JOIN LATERAL rlm(q, ctx)   ─▶ one child per input row, depth 1, worker tier
       ├─ CROSS JOIN LATERAL spawn('reviewer', task) ─▶ child with the role's tools and budget
       └─ FINAL FROM (query)               ─▶ the answer relation
  • Prefix caching. The system prefix is deterministic for the same inputs and marked cacheable for providers that support it, so a long transcript costs new tokens only for what changed.
  • Namespaces. A child’s tables are cgs_<id>__name in the shared store; ctx is preloaded one row per paragraph. Its FINAL comes back to the parent as (answer, detail JSON) for rlm and (answer, detail JSON, session) for spawn; a child that never reached FINAL returns one row with a NULL answer and the outcome in detail, so the parent’s statement survives.
  • Budgets. Budget has calls, tokens, dollars, depth and wall clock. A child gets the parent’s remaining slice intersected with its role’s budget; its spending rolls up. Depth is enforced at the call site.
  • Persistence and resume. kleene_sessions is written after every turn (transcript, functions, agents, settings, usage). kleene resume continues a root session that hit its turn cap or was interrupted.
  • Cancellation. LiveSink::cancel_statement is checked before every model and tool call; the model sees a cancelled statement as an error and goes on.
  • Observation. HarnessConfig.observer is an Observer with hooks for session start and end, each turn, and each model call’s start, streamed text and finish. The session’s own reply goes through LiveSink::complete_streaming, which uses the provider’s streaming path and forwards every text delta to the observer; statement calls use the buffered path. kleene run’s printer and the daemon (which turns deltas into CallDelta messages for the TUI) are both observers.

Store and trace

One DuckDB file (.kleene/run.duckdb by default) holds everything:

TablesWritten byRead by
user tables, ctx, cgs_*__*statementsstatements
kleene_columnsstoreplanner (declared CallSQL types; JSON is stored as VARCHAR)
memoLiveSinkLiveSink (keyed by model and prompt fingerprint, across runs)
trace_runs, trace_sessions, trace_statements, trace_calls, trace_tool_calls, trace_rounds, trace_finalDuckDbTraceSink from a background taskkleene trace, /trace in the TUI, EXPLAIN’s sampled selectivity
kleene_sessionsharness after every turnkleene resume
tasks, task_ratings, solver_ratings, generator_state, playbook, playbook_evals, attempts, view trace_taskslearnlearn board/report/playbook, /board in the TUI
evals, bench_runsbenchbench report/curve/csv

TraceEvents flow through a Tracer to a sink. In-process that is the store sink; under the daemon a FanoutSink mirrors them into the event log as well, so the TUI and the trace tables tell the same story. The store flushes the async writer before answering a Query, so /trace never lags.

Processes

kleene run / repl / explain / learn / bench      one process, in-process harness

kleene daemon ──── .kleene/daemon.sock ────┬─ kleene tui      (ratatui client)
   Daemon: EventLog with Cursor{generation,seq}   ├─ kleene attach   (JSONL, headless)
   Harness + store + provider                     └─ any JSONL client

kleene with no arguments is the unified experience: the TUI is one scrolling stream with a prompt, and a task typed there becomes a StartRun on the daemon. The reply streams back as CallDelta messages and each finished turn as a TurnFinished carrying the reply, its SQL and every statement’s rendered result, so the client shows the conversation without reading the store. /sql is a Submit, /trace and /board are Query.

The protocol (version 3) is newline-delimited JSON over a Unix socket. Requests: Subscribe { after }, StartRun, Submit (REPL statements for a client-owned session), Query (DuckDB SQL over the store, with a tag), Cancel, CancelRun, ListRuns, Reload (re-read the keys; /setup sends it after saving), Detach. Replies: Hello, Event (a trace event with its cursor), CallDelta, TurnFinished, RunAccepted, RunFinished, Submitted, Table, Runs, Ok, Error. A client that reconnects sends its last cursor and gets replay from there; a stale generation replays everything. kleene tui starts a daemon in the background if none is listening.

Model layer

Provider is one trait (complete, streaming optional). Two adapters speak the wire formats directly over reqwest: the Anthropic Messages API and the OpenAI-compatible chat API (OpenAI, local servers, gateways). A third, OpenResponsesProvider, is compiled in with the gateway feature of kleene-llm and speaks the Open Responses API (POST /responses) of Aura or any other such gateway, taking cost from usage.cost_usd. RoutedProvider resolves an alias (root, worker, proxy, judge, or your own) to an ordered list of (provider, model) candidates with default options, skips candidates whose circuit breaker is open, prices usage from a Pricing table, and is what the harness holds. ReplayProvider serves recorded fixtures for tests and offline benchmarks; RecordingProvider writes them.

Configuration is ProviderSettings: the config file kleene setup writes (~/.config/kleene/config.toml, owner-readable) with the environment layered on top field by field, so KEY=... kleene run still wins. Besides the model providers it names the service behind the web_search tool ([web_search]: Brave, Tavily, Exa or Linkup and a key), which the CLI turns into a kleene_tools::WebSearchBackend on the ToolContext. The setup wizard itself lives in kleene-tui::setup as a pure state machine with a renderer, shared by kleene setup, the TUI’s first start and its /setup command; saving from /setup sends the daemon Reload, which rebuilds the provider and the search backend for runs started from then on. See CLI.md.

Planner

kleene-algebra runs rules over the annotated plan until nothing changes. Join ordering runs once first: it collects the relations under a tree of inner and cross joins, then does dynamic programming over subsets with branch-and-bound, pricing each split with the predicates that first become applicable there (pure conjuncts narrow the pairs, call conjuncts pay one call per surviving pair). Cheap-first orders conjuncts to mirror the executor’s left-to-right short-circuit. Cascade rewrites an oracle predicate with a declared proxy into a band. Semi-join turns [NOT] EXISTS with a call into a semi or anti join that stops at the first (counter)example. Beam recursion simulates the fixpoint round by round so a trailing ORDER BY … LIMIT k caps every frontier. Estimates use the store’s row counts, sampled selectivities (observed pass rates of boolean call predicates in this session), per-alias cost factors and pricing.

The continual loop and benchmarks

harness::learn keeps the loop’s state in the store. Generators (sat3, graph, puzzle, corpus, repo, statements, contracts) produce tasks deterministically from a dial and a seed, each with a code oracle (verify). The curriculum picks the pending task nearest even odds for the current solver rating (Bradley-Terry over tasks and solver configurations), runs it through Harness, judges FINAL, moves both ratings and the generator’s dial. SQL that solved a task becomes a playbook candidate and is adopted only after a replay eval with memo bypassed wins on solves and cost. learn::bench runs task packs (tasks/*/pack.json) under learning, frozen and plain modes; plain is a tool-calling agent on the same provider, tools and budget, the baseline for cost parity.

Testing strategy

  • The relational core is tested differentially against DuckDB (kleene-difftest): proptest generates schemas, data and queries in the CallSQL/DuckDB intersection and compares multisets, plus a curated corpus of shrunk regressions.
  • Model-facing text (EXPLAIN, rendered tables, error messages) is part of the interface and asserted verbatim.
  • Anything that needs a model uses scripted or replay providers; no test talks to a network.
  • Crates that link DuckDB keep one integration-test binary each (tests/all/main.rs) because every such binary carries the engine.

Where to look

QuestionFile
What SQL is acceptedcrates/kleene-sql/src/planner.rs, extensions.rs
How a plan is priced and rewrittencrates/kleene-algebra/src/annotate.rs, rules.rs
What the model reads each turncrates/kleene-harness/src/prompt.rs
The turn loop and childrencrates/kleene-harness/src/session.rs
Memo, budget, trace per callcrates/kleene-harness/src/live.rs
Wire formatscrates/kleene-llm/src/adapters/
Tool semantics and the workspace jailcrates/kleene-tools/src/tools/, paths.rs
Store schemacrates/kleene-store/src/duckdb.rs
Daemon protocolcrates/kleene-daemon/src/lib.rs
TUI layout, keys, themecrates/kleene-tui/src/ui.rs, lib.rs, theme.rs
Provider config file and the setup wizardcrates/kleene-llm/src/config.rs, crates/kleene-tui/src/setup.rs
The learning loopcrates/kleene-harness/src/learn/
CLI wiringcrates/kleene/src/main.rs