What It Is
lithify is a context store for long-running LLM agents.
You write a plain State type and a Delta enum; lithify journals every
delta and every side-effect result before it is used, so the agent survives
kill -9 and resumes at the exact step without re-executing
anything whose result was already committed. The same journal gives you the state as of any
event, a fork from any point that costs one row, compaction that keeps what was compacted
away, and a results matrix across every branch.
The model is git's: an append-only log of content-addressed events, branch refs that point at
a sequence number, snapshots as trees, compaction as squash, replay as checkout.
Sediment is deposited layer on layer; under pressure it compacts and lithifies into rock, which
keeps the layers. That is the whole design.
Why
Anyone who has run an agent for longer than a coffee break has hit four problems.
The process dies, and resuming means a hand-written
checkpoint format that silently re-executes whatever happened since the last save: an email
sent twice, a model call billed twice. Nobody knows what the model
saw at step 37; the state that produced the prompt is recorded, if at all, as a pickle.
Experiments are expensive: asking what would have happened
with a different instruction at step 37 means paying for steps 1 through 36 again, so the
experiment runs once, with n = 1. And the context fills up,
gets truncated or summarised by an ad hoc rule, and what was lost is gone.
Each of these is usually solved separately, by hand, in every agent framework, and the
solutions do not compose. lithify treats them as one problem with one well-understood shape.
Durability is the log. Time travel is materialisation at an arbitrary sequence number, which
recovery already needs. A fork is a row that points into the log. Compaction is a snapshot,
which replay already uses. None of these adds a storage concept; the library is
six redb tables.
How It Works
Every change is an event appended to a branch, and there are four kinds:
Delta (a consumer delta, applied on replay),
Effect (the recorded result of a model call, tool call or
file read, read back instead of re-executed),
Compact (the state replaced wholesale by a compactor, with a
snapshot at the same sequence number) and
Observe (a measurement for the results matrix). Each event
carries blake3(parent_hash ‖ canonical(event)), so a branch head's hash identifies
everything that has ever happened on it.
A branch is a row: parent, fork point, head. Forking at sequence s inserts one row and
copies nothing. branch.effect(key, || …) looks the key up through the fork chain, so
a sweep forked at step 37 shares steps 1–37's model calls, tool calls
and file reads and pays only for its own tail. A Compactor is handed the
branch, so the summarisation call it makes is itself journaled; the pre-compaction history stays
addressable, and "what did the agent know before it forgot?" is branch.at(seq - 1).
storage: six redb tables
branches u64 -> BranchMeta
children u64 ->> u64 multimap: parent -> children
events (u64, u64) -> Record (branch, seq)
effects (u64, [u8;32]) -> u64 (branch, blake3(key)) -> seq
snapshots (u64, u64) -> State
by_hash [u8;32] -> (u64, u64)
redb was chosen for discipline as much as convenience: pure Rust, ACID, one writer with MVCC
readers, an in-memory backend on the same code path, and no query
language. Every join is written by hand, and the first time one would need a shuffle,
that friction is the signal the data model is wrong. Analysis lives downstream: the matrix
exports to CSV for DuckDB or Polars.
Quick Start
A context that survives kill -9, time-travels and forks
use lithify::{Journal, State};
use serde::{Deserialize, Serialize};
#[derive(Default, Clone, Serialize, Deserialize)]
struct Ctx { messages: Vec<String> }
#[derive(Serialize, Deserialize)]
enum Delta { Say(String) }
impl State for Ctx {
type Delta = Delta;
fn apply(&mut self, d: &Delta) {
match d { Delta::Say(s) => self.messages.push(s.clone()) }
}
fn size(&self) -> usize { self.messages.iter().map(|m| m.len()).sum() }
}
fn main() -> lithify::Result<()> {
// resumes if run.redb already exists
let journal = Journal::<Ctx>::open("run.redb")?;
let ctx = journal.root();
// journaled before it is returned; a committed result is never re-run
let reply: String = ctx.effect("llm:step-1", || call_model("hello"))?;
ctx.apply(Delta::Say(reply))?;
let then = ctx.at(1)?; // the state as of event 1
let variant = ctx.fork_at(1, None, serde_json::json!({ "temperature": 0.2 }))?;
Ok(())
}
Run the tests and the demo
git clone https://github.com/nihilai-collective/lithify.git
cd lithify
cargo test # includes a kill -9 / resume test
scripts/demo.sh /path/to/a/source/tree # the pitch in one minute
The demo runs the bundled auditor agent, which audits a source
tree with a budgeted, compacting context. It SIGKILLs the agent four times and finishes the run,
shows the state as of event 40, lists events with their content hashes, and forks a five-budget
sweep from nineteen files in, printing recall per budget. A deterministic mock model is the
default; --model openai points it at any /v1/chat/completions endpoint.
Measured Against LangGraph
The same agent was written twice, once on lithify and once on
LangGraph 1.2 with SqliteSaver, and run against one corpus
(the redb 4.3.0 source tree, 74 files) with one deterministic mock model, so the only thing that
differs is the persistence layer. Both produce identical results: 151 model calls, 2 compactions,
83 findings, recall 1.0, and identical recall-versus-budget curves across a five-point sweep.
The harness is sound. Then each agent was run as a child process with 20 ms of injected model
latency and SIGKILLed at random, eight times per trial, three trials each.
0
Tool calls redone
lithify · 3 trials × 8 SIGKILLs
65
Tool calls redone
LangGraph graph API · same trials
~57 µs
Per fork
in memory · 2000-event history
Read this carefully, because the headline is not "lithify never re-executes". Both systems redo
roughly one model call per kill: the call that was in flight.
That is a floor no client-side journal can get under; a call that was sent but whose response
was not committed is indistinguishable, after a crash, from one never sent. What differs is the
blast radius. lithify's unit of durability is the effect, so the in-flight call is redone and
nothing else is. LangGraph's graph API checkpoints when a node returns, so a kill mid-step
redoes the read, the plan and every tool call in that step.
| Capability |
lithify |
LangGraph graph API + SqliteSaver |
| Unit of durability | One effect / delta | One node |
| In-flight model call redone on crash | Yes (the floor) | Yes (the floor) |
| Other work in the step redone on crash | No | Yes |
| Time-travel granularity | Event | Node |
| Fork cost | One row | One checkpoint write |
| What a fork inherits | State and effects | State |
| Cross-branch results query | matrix() | Read the sqlite file by hand |
| Compaction primitive | Compactor + snapshot | None (user code) |
| Wall time, 74 files, mock model | 405 ms | 172 ms |
Two things in LangGraph's favour, stated plainly. The wall-time gap is real: lithify fsyncs every
event (755 commits in that run against LangGraph's 79 checkpoints). Against real model latency it
vanishes; against a mock it dominates, and Options::lazy_deltas exists to drop the
fsync for deltas and observations while keeping effects immediate. And LangGraph's
functional API (@entrypoint / @task)
persists task results individually and would be expected to narrow the tool-call column to zero;
the comparison is against the graph API because that is what most LangGraph agents are written in.
The model-call column would not move.
Features
- Resume after kill -9 at the exact step; committed effects are never re-executed
- Time travel at single-event granularity with
branch.at(seq)
- O(1) forks that share the prefix, including its effects
- Compaction journaled with a snapshot; the pre-compaction history stays addressable
- Content-addressed history: a blake3 hash chain per branch
- Results matrix:
observe() across branches, joined to each branch's parameters, CSV export
- Your own
State and Delta types; no opinion about models, prompts or tools
- Six redb tables, about 800 lines of library
- Pure Rust: no C toolchain, no system SQLite
- On-disk and in-memory backends on one code path
- Property, replay, fork and SIGKILL tests
- MIT or Apache-2.0, at your option
Status
lithify is a proof of concept, and the API will change. The
white paper states eight falsifiable hypotheses; replay equivalence, effect idempotency, fork
cost and crash blast radius are tested and hold. The interesting open one is compaction with a
real LLM compactor, where a larger budget producing a worse summary is plausible and would be a
finding about the compactor rather than the journal.
⚠ Known Limits
redb is single-process: two processes opening one journal is undefined. The state is serialised
whole at every snapshot, which is fine at agent scale and wasteful for megabyte contexts. Effect
keys are the consumer's responsibility; two calls with one key on one history see one execution,
which is the feature and the footgun. Renaming a
Delta variant will hurt until
versioned deltas with upcasting exist.
Where it is going: a #[derive(Delta)] once a second consumer shows what is common,
coordination-free branch merges for lattice-shaped deltas, and a shared multi-agent corpus that
agents compact into and pull knowledge back out of. Durability is where it had to start, because
everything else is built on it.