Skip to content

· Olivier Moreau · AI Notes  · 15 min read

Wiki or RAG? The real answer is a file format

An LLM-maintained wiki beats retrieval for knowledge that compounds. But a memory written by an agent is only as trustworthy as what it records about itself — and that is exactly the gap the Open Knowledge Format was built to close.

Ask an assistant what you agreed with a client three months ago, and one of two things happens behind the scenes.

Either it goes digging. It searches every note, email and transcript it can reach, pulls out the handful of passages that look closest to your question, and assembles an answer on the spot. Or it opens a page it wrote weeks ago — a page it has kept current ever since — and reads you what it says.

The first is retrieval-augmented generation. The second is what Andrej Karpathy called an LLM wiki. The argument between them is usually framed as a choice of architecture. It is really a question of trust, and the most useful answer to arrive this year is not a retrieval technique at all. It is a file format.

Two ways to remember

Karpathy’s gist, published in April, describes the retrieval experience with uncomfortable precision:

“The LLM is rediscovering knowledge from scratch on every question. There’s no accumulation.”

Ask something subtle that needs five documents stitched together, and the model has to find and assemble the same fragments every single time. Nothing is built up.

The wiki inverts that. Instead of retrieving from raw documents at query time, the LLM “incrementally builds and maintains a persistent wiki — a structured, interlinked collection of markdown files that sits between you and the raw sources.” A new source isn’t just indexed for later. It is read, and its substance is folded into the pages that already exist: summaries revised, entity pages updated, contradictions with older claims noted.

The design has three layers, and they are worth keeping distinct:

LayerWho owns itWhat it is
Raw sourcesYouImmutable documents. The model reads them and never edits them. The source of truth.
The wikiThe modelGenerated markdown pages — concepts, entities, summaries — cross-linked and kept consistent.
The schemaYou and the modelA CLAUDE.md or AGENTS.md describing the structure, the conventions and the workflows.

Three operations run over it: ingest a source, query the wiki, and lint it for contradictions, stale claims and orphaned pages. Two special files make it navigable: an index.md that catalogues every page in one line each, and an append-only log.md recording what happened and when.

“The wiki is a persistent, compounding artifact.”

Why the wiki wins

Personal and team wikis have always failed for the same reason, and Karpathy names it:

“The tedious part of maintaining a knowledge base is not the reading or the thinking — it’s the bookkeeping.”

Humans abandon wikis because the maintenance grows faster than the value. A model doesn’t get bored, doesn’t forget a cross-reference, and can touch fifteen files in one pass. The economics that killed every internal wiki you have ever seen simply stop applying.

The strongest public demonstration so far is Cole Medin’s knowledge base, compiled from 198 long-form videos. Its builders report that 785 candidate concepts collapsed into 186 canonical pages, that 99% of the 2,567 quotes it cites were found verbatim in the transcripts at the claimed timestamps, and that it declined all ten deliberately out-of-scope trap questions in its QA run. Those are the builder’s own validation figures and deserve to be read that way — but the method is published and re-runnable.

The number that matters is the first one. Merging 785 overlapping ideas into 186 pages is the synthesis. A retrieval system would have to perform that merge again on every question, invisibly, with whatever fragments happened to rank highest that day. The wiki does it once, in the open, where you can read it, diff it and correct it.

Your brain already works this way

The pattern feels new, but the architecture is old. It is roughly how your own memory is organised.

Neuroscience has a well-supported account of this, known as complementary learning systems. The hippocampus learns fast: it captures specific episodes — this meeting, that conversation, what was said and when — almost as they happen. The neocortex learns slowly: over time it extracts what those episodes have in common and keeps it as general, structured knowledge. You know what a contract renewal involves without remembering every renewal you learned it from.

The two are kept apart for a reason. A store that absorbed each new experience straight into its structured knowledge would overwrite what it already knew. So new material is held raw first and folded in gradually — much of it during sleep, when the hippocampus replays recent experience and the cortex integrates it.

Map it across, and the wiki’s layers line up almost one to one:

Human memoryLLM wiki
Hippocampus — fast, specific, episodicRaw sources and daily logs — immutable, timestamped
Neocortex — slow, general, structuredThe wiki — synthesised, cross-linked pages
Consolidation during sleepThe compile pass, often run after the working day
Semantic memory: knowing something without the momentA concept page that merges many sources
Recalling one specific episode in detailSearching the raw sources for the exact wording

That last row matters later. Nobody gets through the day by replaying episodes, but nobody settles an argument about exactly what was said from a general impression either. The brain keeps both, and so should a memory system.

The analogy is a lens, not a model: brains don’t store markdown, and the biology is far messier than a table. But it explains why the wiki feels right — and it also predicts how the wiki goes wrong.

Compounding cuts both ways

Here is what the enthusiastic version of the story leaves out. A persistent, compounding artifact compounds its mistakes too.

When retrieval gets something wrong, the error lives in one answer and evaporates with it. When a wiki gets something wrong, the error is written down, linked to, and cited by the next page the agent writes. Running an agent-maintained memory in production, these are the failure patterns that actually showed up — none of them exotic:

  • The health check that never ran. The lint pass crashed before it wrote its report. No report looked exactly like no problems, so for weeks the knowledge base was believed to be clean. The first run that actually completed came back with a long list.
  • The link quota. The compiler was told every article needed at least two links, without being required to link to pages that existed. It obliged by inventing plausible names for articles nobody had written. A rule meant to make the graph dense made it partly fictional.
  • The empty checkout. A fresh working copy without the memory folder linted perfectly: zero articles checked, zero issues found. The most reassuring output possible, for the least coverage possible.
  • Every page looked equally sure of itself. An article compiled from one offhand remark in a meeting carried exactly the same authority as one backed by three signed documents. Nothing on the page said where it came from, who had confirmed it, or whether it was still true.

The first three are process bugs. They get fixed with tests, and they did. The fourth is structural: the page had nowhere to record its own credibility. No amount of better prompting fixes a format that has no field for the answer.

Human memory fails in exactly these two ways. Research on reconsolidation suggests that recalling a memory makes it briefly editable before it is stored again, which is one reason details picked up after the fact get folded in and later remembered as original. And psychologists have a name for the fourth failure: source amnesia — remembering a fact while forgetting where you learned it, and with it any sense of how far to trust it. An agent that rewrites a page on every ingest is doing reconsolidation at machine speed. A page with no provenance is source amnesia by design.

Search didn’t lose

It is worth reading Karpathy closely here, because the gist is often summarised as “no RAG needed”. It doesn’t say that. It says the index file works “surprisingly well at moderate scale (~100 sources, ~hundreds of pages) and avoids the need for embedding-based RAG infrastructure” — and that as the wiki grows, “you want proper search”, pointing to a local hybrid keyword-and-vector engine. Even Cole’s deliberately embedding-free bundle uses local embeddings in its validation suite, to prove no two pages are near-duplicates.

The practical split looks like this:

The questionBest served by
Something already synthesised: a decision, a relationship, a whyThe wiki page
Exact wording: the clause, the figure, what the client wroteSearch over the raw sources
Anything from today that hasn’t been compiled yetSearch
A corpus too large or too messy to compile economicallySearch
”Is anything missing from what we know?”Neither — that needs a recall test

So the honest answer to “wiki or RAG” is both: the wiki is the map, and search covers the territory the map doesn’t show yet. Which moves the real question somewhere more interesting. When a page and a search result disagree — or when a page is simply old — which one do you believe?

OKF: the layer that was missing

The Open Knowledge Format began in Google Cloud’s Knowledge Catalog samples and now lives in its own repository, at version 0.2. Its ambition is deliberately small: “a directory of markdown files with YAML frontmatter”. In the spec’s own words, “if you can cat a file, you can read OKF; if you can git clone a repo, you can ship it.” The only field every document must carry is type.

What makes v0.2 relevant to this argument is the premise it starts from:

“Increasingly, a knowledge corpus is not authored once and then read: it is continuously written and maintained by agents.”

From there it makes five questions answerable from frontmatter alone: where did this come from, how much should I trust it, is it still true, is it the current version, and was this number produced the way it was supposed to be. Here is what a single memory page looks like with those answers attached:

---
type: Decision
title: Acme renewal — 12% discount, capped at one year
description: Discount agreed for the 2027 renewal, conditional on a two-year term.
tags: [acme, pricing, renewal]
status: stable
generated: { by: compiler/claude-sonnet-4-6, at: 2026-09-10T17:42:00Z }
verified:
- { by: human:jmartin, at: 2026-09-11T09:15:00Z }
stale_after: 2027-01-31T00:00:00Z
sources:
- id: renewal-call
resource: /daily/2026-09-10.md
title: Notes from the renewal call
- id: signed-quote
resource: https://drive.example.com/acme-quote-2027
title: Signed renewal quote
author: human:acme-procurement
last_modified: 2026-09-11T08:02:00Z
---
The discount applies to the first year only.[^signed-quote] Acme asked for
two years at 12%; we declined on the call.[^renewal-call]
[^signed-quote]: Signed renewal quote
[^renewal-call]: Notes from the renewal call

Every one of those fields closes a gap from the list above. In the terms of the brain analogy, OKF gives the wiki the one thing human memory is notoriously bad at: remembering where each thing came from.

sources records signals, not a score. Each source can carry who authored it, when it last changed and how much it is used. OKF refuses to store a credibility rating, because “a score is subjective, unportable across consumers, and goes stale”. A page resting on a signed document and a page resting on a hallway remark now look different, and a consumer can decide what that difference is worth.

Claims point at their source by key, not by position. The footnote label is the id of a sources entry. That sounds pedantic until you remember who is editing these files: agents rewrite them constantly, and a citation that says “source number two” silently misattributes the moment the list is reordered.

Writing and confirming are different events. generated records who produced the current content; verified records who checked it against its sources. From that, a consumer derives a trust tier — unverified, machine-confirmed or human-reviewed — keyed on whether a human: actor appears. The page a model wrote at three in the morning and the page a colleague signed off are finally distinguishable without opening either.

Freshness is a comparison, not a guess. status is draft, stable or deprecated; stale_after is an absolute instant. A page is stale when now is past it. No relative TTL, no reasoning about when the page was read.

A missing link is a to-do, not a defect. The spec says consumers must tolerate broken links, because a link to a page that doesn’t exist “may simply represent not-yet-written knowledge”. That is the right reading of the link-quota failure: the problem was never unwritten targets, it was a rule that rewarded producing them.

Numbers can be attested. For figures that matter — revenue, margin, anything you’d put in front of a board — an Attested Computation pins the sanctioned query. The agent may only supply parameter values; it must not write or edit the computation itself, and a deterministic checker compares what actually ran against what was sanctioned. It is the difference between an agent that reports a number and one that can prove where the number came from.

Crucially, OKF lists “prescribing storage, serving, or query infrastructure” as a non-goal. It does not pick a side in the wiki-versus-RAG argument. It makes both sides honest. A retriever can rank a deprecated page lower. An agent reading the index can see that a page is unverified before it repeats the claim to you.

What changes when the page carries its own credibility

A few patterns follow directly, and they are the ones we have found worth keeping.

Freshness should rank, not hide. A stale or deprecated page should still come back from a search, just lower down. Silently filtering it out is worse than showing it with a warning, because the fact you can’t see is the one you can’t correct.

A source link should survive the link dying. Documents move and share links break. A one-line note alongside each source keeps the memory useful after the URL returns a 404. OKF permits exactly this kind of producer-defined extra field.

Writers must never invent a source. A page cites a document only if that document actually appeared in the material being compiled. Provenance you fabricated is worse than provenance you left blank, because blank at least reads as unverified.

Filed answers start at the bottom tier. Karpathy’s best idea is that good answers should be written back into the wiki so explorations compound. With OKF, an answer filed by the agent arrives as generated and unverified — useful immediately, trusted only once somebody confirms it.

Before you trust a memory an agent wrote

None of this argues against letting an agent maintain your knowledge. It argues for knowing what has to be true first, and the honest sign that it isn’t yet.

  • Every page names its sources. Not ready if a page can assert something without saying where it came from.
  • Writer and verifier are recorded separately. Not ready if you can’t tell a page a person confirmed from one the model wrote unsupervised.
  • Facts carry an expiry. Not ready if the only way to find a stale page is to act on it and discover it’s wrong.
  • Raw sources are immutable. Not ready if the agent can edit the document it is citing.
  • Health checks fail loudly. Not ready if “no report” and “no problems” look identical.
  • Checks prove they saw something. Not ready if your lint passes on an empty folder.
  • Structural rules don’t reward invention. Not ready if the compiler is judged on how many links it produces.
  • Search covers what the wiki doesn’t. Not ready if a question about this morning’s notes comes back empty because they haven’t been compiled.
  • Recall is tested, not just precision. Not ready if nobody has ever asked the knowledge base what it is missing.
  • Reported numbers run the sanctioned way. Not ready if the agent writes its own query for a figure you are going to repeat to a client.

The argument was in the wrong place

Wiki versus RAG is an argument about when synthesis happens: once, at write time, or every time, at read time. Both answers are right for different questions, and any memory worth having will end up using both.

What neither approach could tell you, on its own, was whether to believe the result. A retrieved chunk and a compiled page both arrive looking equally confident. That was survivable while humans wrote most of what we store, because we carried the context about who wrote what and how reliable they were.

Agents now write most of it. The knowledge has to carry that context itself — and for the first time, there is a plain-text, vendor-neutral place to put it.


Further reading

The pattern

  • LLM Wiki — Andrej Karpathy’s original idea file: raw sources, the wiki, the schema, and the ingest / query / lint loop.
  • claude-memory-compiler — Cole Medin’s implementation for Claude Code: hooks capture sessions, a compiler turns daily logs into cross-linked articles.

The format

  • Open Knowledge Format — the canonical repository for the OKF v0.2 specification, reference agent and sample bundles.

The evidence

  • Cole Medin AI Knowledge Base — an OKF bundle compiled from 198 videos, with its full build pipeline and validation results in docs/MAKING-OF.md.

The analogy

The tooling

  • qmd — the local hybrid BM25 / vector search engine for markdown that Karpathy suggests once a wiki outgrows its index.
Back to Blog

Related Posts

View All Posts »

Tool calling is a protocol, not a feature

Every provider's comparison table says "supports function calling." That's a checkbox on a spectrum. The interesting question is what your agent does on the turn the model answers in the wrong channel — and why it's your cheapest tier that gets there first.

Nobody read last night's diff

Fully autonomous coding pipelines are real and shipping. What decides whether yours produces software or defects isn't the model — it's whether your documents say what success looks like, precisely enough to be checked without you.