The Log Is the Truth
In mid-August my weekly review email told me I had worked nine sessions that week. It was actually twelve. Nothing was broken in the sense of a crash or a bad deploy; the number was wrong because my vault kept two records. Every session ends with a handoff file, and every session was also supposed to get an entry in a running session log. The weekly counted the log. The log was written by a nightly repair loop that ran after the weekly fired, and on Saturdays the two never met. After three months of drift, the defect was obvious: I had two sources of truth for “a session happened,” and every new agent session was free to pick the wrong one.
system-o is an open-source framework I run my workspace on: a markdown vault, a deterministic nightly chain, and a loop layer where an LLM proposes maintenance edits that scripts verify and gate. Its spec already said the handoff was the authority for session completion and that summaries should be derived where practical. I had executed that rule for time tracking in early August and for the headline count on August 16. The session log itself was still a second write. Version 0.4, built on August 21, finishes the thought. The log now carries a declared generated region. A script reads every handoff on disk, live and archived, and regenerates that region from them: the handoff’s title, its link, its completion note or opening paragraph, and the minutes it recorded. A wrap writes the handoff and nothing else. Everything below the region’s end marker is frozen history plus hand-written entries for the rare session that wrote no handoff. The nightly loop cell that used to draft log entries with a language model now has no model in it at all; its repair is “regenerate,” and the conformance harness proves the property I actually care about: run it twice, get identical bytes.
Three days before I built that, Cursor published a post about rebuilding Git hosting for agent-driven development. Their problem was the same shape at a scale I will never see. The industry-standard design they inherited keeps several full replicas of each repository and reaches consensus across them on every push, so every operation runs at the speed of the slowest server and every repository is a pet that needs checksumming and hand repair. Their replacement, Continuity, puts a write-ahead log in object storage and makes that log the single source of truth. Replicas become warm caches that can be rebuilt from the log or thrown away; pushes stay consistent through compare-and-swap on the log head; reads verify against the log with a conditional fetch. They quote 120 pushes per second on standard S3 and north of 300 on the express tier, and the freedom to give a monorepo hundreds of read replicas while a throwaway agent repository gets one.
I am not comparing the engineering. I am comparing the decision, because it is the same decision: find the one artifact that is allowed to be true, make it small and replayable, and demote everything else to a cache you would delete without a second thought. Cursor’s log is a sequence of Git pack writes; mine is a folder of dated markdown files. Their correctness check is that a replica can be rebuilt from the log; mine is that regenerating the log twice changes nothing. The pets-versus-cattle framing transfers almost without translation. A session log I hand-maintain is a pet: it needs a guard, a repair loop, and a grace window. A session log I derive is cattle.
The part of Cursor’s post that I think matters most for anyone running agents is the half about volume. Their stated driver is more code, more pull requests, more CI runs, and an explosion of short-lived repositories created by agents rather than people. My vault is that world in miniature. Agents work in isolated worktrees, ideas get a thirty-day clock on a launch pad before they are promoted or buried, and a trash folder auto-purges after a month. What made that tier cheap to run was deciding early which records persist (a registry card, a tombstone, a handoff) and letting everything else be disposable. Continuity is that same choice made at the storage layer: durability lives in the log, and the thing you scale is the number of caches.
What I would tell someone building a vault, a pipeline, or a governance process from this: the instinct to keep a tidy secondary record is how you end up with two truths. Pick the log. Derive the rest. Then test that the derivation is deterministic, because a cache you cannot rebuild byte-for-byte is not a cache, it is a second pet.