Nihilai Collective Logo

lithify

Nihilai Collective

Journaled agent context · Time travel · O(1) forks · Compaction · Pure Rust on redb

What It Is

lithify is a context store for long-running LLM agents. You write a plain State type and a Delta enum; lithify journals every delta and every side-effect result before it is used, so the agent survives kill -9 and resumes at the exact step without re-executing anything whose result was already committed. The same journal gives you the state as of any event, a fork from any point that costs one row, compaction that keeps what was compacted away, and a results matrix across every branch.

The model is git's: an append-only log of content-addressed events, branch refs that point at a sequence number, snapshots as trees, compaction as squash, replay as checkout.

Sediment is deposited layer on layer; under pressure it compacts and lithifies into rock, which keeps the layers. That is the whole design.

Why

Anyone who has run an agent for longer than a coffee break has hit four problems. The process dies, and resuming means a hand-written checkpoint format that silently re-executes whatever happened since the last save: an email sent twice, a model call billed twice. Nobody knows what the model saw at step 37; the state that produced the prompt is recorded, if at all, as a pickle. Experiments are expensive: asking what would have happened with a different instruction at step 37 means paying for steps 1 through 36 again, so the experiment runs once, with n = 1. And the context fills up, gets truncated or summarised by an ad hoc rule, and what was lost is gone.

Each of these is usually solved separately, by hand, in every agent framework, and the solutions do not compose. lithify treats them as one problem with one well-understood shape. Durability is the log. Time travel is materialisation at an arbitrary sequence number, which recovery already needs. A fork is a row that points into the log. Compaction is a snapshot, which replay already uses. None of these adds a storage concept; the library is six redb tables.

How It Works

Every change is an event appended to a branch, and there are four kinds: Delta (a consumer delta, applied on replay), Effect (the recorded result of a model call, tool call or file read, read back instead of re-executed), Compact (the state replaced wholesale by a compactor, with a snapshot at the same sequence number) and Observe (a measurement for the results matrix). Each event carries blake3(parent_hash ‖ canonical(event)), so a branch head's hash identifies everything that has ever happened on it.

A branch is a row: parent, fork point, head. Forking at sequence s inserts one row and copies nothing. branch.effect(key, || …) looks the key up through the fork chain, so a sweep forked at step 37 shares steps 1–37's model calls, tool calls and file reads and pays only for its own tail. A Compactor is handed the branch, so the summarisation call it makes is itself journaled; the pre-compaction history stays addressable, and "what did the agent know before it forgot?" is branch.at(seq - 1).

storage: six redb tables
branches u64 -> BranchMeta children u64 ->> u64 multimap: parent -> children events (u64, u64) -> Record (branch, seq) effects (u64, [u8;32]) -> u64 (branch, blake3(key)) -> seq snapshots (u64, u64) -> State by_hash [u8;32] -> (u64, u64)

redb was chosen for discipline as much as convenience: pure Rust, ACID, one writer with MVCC readers, an in-memory backend on the same code path, and no query language. Every join is written by hand, and the first time one would need a shuffle, that friction is the signal the data model is wrong. Analysis lives downstream: the matrix exports to CSV for DuckDB or Polars.

Quick Start

A context that survives kill -9, time-travels and forks
use lithify::{Journal, State}; use serde::{Deserialize, Serialize}; #[derive(Default, Clone, Serialize, Deserialize)] struct Ctx { messages: Vec<String> } #[derive(Serialize, Deserialize)] enum Delta { Say(String) } impl State for Ctx { type Delta = Delta; fn apply(&mut self, d: &Delta) { match d { Delta::Say(s) => self.messages.push(s.clone()) } } fn size(&self) -> usize { self.messages.iter().map(|m| m.len()).sum() } } fn main() -> lithify::Result<()> { // resumes if run.redb already exists let journal = Journal::<Ctx>::open("run.redb")?; let ctx = journal.root(); // journaled before it is returned; a committed result is never re-run let reply: String = ctx.effect("llm:step-1", || call_model("hello"))?; ctx.apply(Delta::Say(reply))?; let then = ctx.at(1)?; // the state as of event 1 let variant = ctx.fork_at(1, None, serde_json::json!({ "temperature": 0.2 }))?; Ok(()) }
Run the tests and the demo
git clone https://github.com/nihilai-collective/lithify.git cd lithify cargo test # includes a kill -9 / resume test scripts/demo.sh /path/to/a/source/tree # the pitch in one minute

The demo runs the bundled auditor agent, which audits a source tree with a budgeted, compacting context. It SIGKILLs the agent four times and finishes the run, shows the state as of event 40, lists events with their content hashes, and forks a five-budget sweep from nineteen files in, printing recall per budget. A deterministic mock model is the default; --model openai points it at any /v1/chat/completions endpoint.

Measured Against LangGraph

The same agent was written twice, once on lithify and once on LangGraph 1.2 with SqliteSaver, and run against one corpus (the redb 4.3.0 source tree, 74 files) with one deterministic mock model, so the only thing that differs is the persistence layer. Both produce identical results: 151 model calls, 2 compactions, 83 findings, recall 1.0, and identical recall-versus-budget curves across a five-point sweep. The harness is sound. Then each agent was run as a child process with 20 ms of injected model latency and SIGKILLed at random, eight times per trial, three trials each.

0
Tool calls redone
lithify · 3 trials × 8 SIGKILLs
65
Tool calls redone
LangGraph graph API · same trials
~57 µs
Per fork
in memory · 2000-event history

Read this carefully, because the headline is not "lithify never re-executes". Both systems redo roughly one model call per kill: the call that was in flight. That is a floor no client-side journal can get under; a call that was sent but whose response was not committed is indistinguishable, after a crash, from one never sent. What differs is the blast radius. lithify's unit of durability is the effect, so the in-flight call is redone and nothing else is. LangGraph's graph API checkpoints when a node returns, so a kill mid-step redoes the read, the plan and every tool call in that step.

Capability lithify LangGraph graph API + SqliteSaver
Unit of durabilityOne effect / deltaOne node
In-flight model call redone on crashYes (the floor)Yes (the floor)
Other work in the step redone on crashNoYes
Time-travel granularityEventNode
Fork costOne rowOne checkpoint write
What a fork inheritsState and effectsState
Cross-branch results querymatrix()Read the sqlite file by hand
Compaction primitiveCompactor + snapshotNone (user code)
Wall time, 74 files, mock model405 ms172 ms

Two things in LangGraph's favour, stated plainly. The wall-time gap is real: lithify fsyncs every event (755 commits in that run against LangGraph's 79 checkpoints). Against real model latency it vanishes; against a mock it dominates, and Options::lazy_deltas exists to drop the fsync for deltas and observations while keeping effects immediate. And LangGraph's functional API (@entrypoint / @task) persists task results individually and would be expected to narrow the tool-call column to zero; the comparison is against the graph API because that is what most LangGraph agents are written in. The model-call column would not move.

Features

  • Resume after kill -9 at the exact step; committed effects are never re-executed
  • Time travel at single-event granularity with branch.at(seq)
  • O(1) forks that share the prefix, including its effects
  • Compaction journaled with a snapshot; the pre-compaction history stays addressable
  • Content-addressed history: a blake3 hash chain per branch
  • Results matrix: observe() across branches, joined to each branch's parameters, CSV export
  • Your own State and Delta types; no opinion about models, prompts or tools
  • Six redb tables, about 800 lines of library
  • Pure Rust: no C toolchain, no system SQLite
  • On-disk and in-memory backends on one code path
  • Property, replay, fork and SIGKILL tests
  • MIT or Apache-2.0, at your option

Status

lithify is a proof of concept, and the API will change. The white paper states eight falsifiable hypotheses; replay equivalence, effect idempotency, fork cost and crash blast radius are tested and hold. The interesting open one is compaction with a real LLM compactor, where a larger budget producing a worse summary is plausible and would be a finding about the compactor rather than the journal.

⚠ Known Limits
redb is single-process: two processes opening one journal is undefined. The state is serialised whole at every snapshot, which is fine at agent scale and wasteful for megabyte contexts. Effect keys are the consumer's responsibility; two calls with one key on one history see one execution, which is the feature and the footgun. Renaming a Delta variant will hurt until versioned deltas with upcasting exist.

Where it is going: a #[derive(Delta)] once a second consumer shows what is common, coordination-free branch merges for lattice-shaped deltas, and a shared multi-agent corpus that agents compact into and pull knowledge back out of. Durability is where it had to start, because everything else is built on it.

Read the white paper →
Repository
Library, the auditor agent, the LangGraph baseline, and the scripts to reproduce every number on this page.
Full Comparison
The lithify vs LangGraph runs in detail: sweep, crash trials, time travel, and lines of code.
All Projects
Everything else the Nihilai Collective builds.