LLM Wiki
An LLM Wiki is a persistent, structured knowledge base of Markdown files that is built and maintained by a large language model (LLM) rather than by humans directly. The pattern was introduced by Andrej Karpathy in April 2026 through a 75-line GitHub Gist that went viral, generating over 17 million views and more than 15,000 GitHub stars within days. Its core principle is knowledge compilation: instead of retrieving raw document fragments at query time as in traditional Retrieval-Augmented Generation (RAG), the LLM reads each source once at ingestion, distills it into interlinked pages, and keeps those pages current over time, so that knowledge is compiled once and compounded rather than re-derived on every query. The canonical architecture has three layers — immutable raw sources, an LLM-maintained wiki layer, and a schema file defining conventions — driven by three operations: ingest, query, and lint.
The paradigm was independently validated within months by four converging implementations: Cognition's DeepWiki, which applied the pattern to GitHub codebase documentation from its April 27, 2025 launch; Factory's AutoWiki, which treats documentation as a build artifact regenerated in CI; LangChain's OpenWiki, which extended the pattern from codebases to personal knowledge through OpenWiki Brains; and Garry Tan's GBrain, first shown publicly as a prototype on 5 April 2026 and formally open-sourced under the MIT license on 10 April, adopted rapidly enough to pass 5,000 GitHub stars within a day, which scaled the pattern to a personal knowledge brain of more than 140,000 pages. Google Cloud released the Open Knowledge Format (OKF) in June 2026 to standardize the pattern as a vendor-neutral knowledge representation, and Anthropic and OpenAI productized the underlying memory-consolidation idea, with Anthropic's Dreams API exposing a documented asynchronous job that reorganises a memory store against up to a hundred past session transcripts.
Empirical work both validated and bounded the pattern. A four-query controlled experiment reported cumulative token use of 47K for the compounding regime against 305K for a long-context baseline, an 84.6% saving; the same study showed chunk-level RAG using only 13.6K tokens, so the paper's claim rests not on raw token minimization but on the argument that compounding tokens function as capital investment, converting query expenditure into persistent, inheritable knowledge artifacts. A preregistered comparison found the compiled wiki stronger at connecting findings across papers while RAG used fewer query tokens, concluding that no architecture was best at organizing evidence, supporting each claim with citations, and minimizing cost at once. The WiCER study documented a 53–60% catastrophic failure rate for blind compilation, and a May 2026 position paper reported that the wiki failed to outperform RAG in two tested domains, arguing it is best understood as a compiled semantic memory for an agent rather than a RAG replacement. The scale limit of the flat index has itself become an engineering target: ByteDance's Volcengine open-sourced [[organizations/openviking|OpenViking]], which replaces scalar index navigation with a virtual filesystem and L0/L1/L2 context tiers loaded on demand for wikis of hundreds to thousands of pages.
A July 2026 follow-up to the WeChat LLM-Wiki line removed the assumption that a wiki is compiled once and then queried. [[WikiLoop]] jointly learns construction and navigation from downstream task feedback, so that a proposed page edit is scored by its effect on answering rather than by a proxy for page quality; with a 9B-parameter backbone it reported 62.6 aggregate answer correctness on the AuthTrace benchmark, 6.3 points above the LLM-Wiki baseline, with the largest gains on multi-document queries.
A complementary line of empirical work reframed quality as a knowledge-integrity problem rather than a context-capacity one. A 2026 study found that recall accuracy in a long interaction fell from 57% to 21% once contradictory facts entered the history, and rose to 73% — above the contradiction-free baseline — when an external metabolism layer resolved contradictions during idle time, a pattern that held across eight models and even a million-token context window. Alongside it, a systematic literature review of RAG for enterprise knowledge management reported that fewer than 15% of studies address real-time integration, and the [[technology/wikikv|WikiKV]] storage system proposed encoding a wiki's hierarchy directly into its key space, with schema evolution and snapshot-consistent reads, rather than indexing flat documents.
A third strand narrowed what verification can deliver. The WeChat/Tencent AuthTrace benchmark, reported on its own page, showed across eight systems and two QA models that evidence recall is the strongest observed predictor of answer correctness (r = 0.96) and that most failures stem from missing evidence rather than flawed synthesis. AttriWiki then showed that whether an answer was drawn from provided context or from parametric memory can be predicted by a linear probe at up to 0.96 Macro-F1, and that attribution mismatches raise error rates by up to 70 percent. A systematic study of uncertainty estimators found their association with hallucination highly variable and often weak, undercutting confidence scores as a direct correctness signal, while formal-methods work demonstrated that Linear Temporal Logic auditing and runtime monitors detect violations of temporally extended constraints more reliably than LLM judges and reduce agent violation rates without materially harming task performance. An independent release-governance study reached a convergent conclusion, finding that evidence coverage was the primary discriminator of severe regressions and that deterministic structural gates and model-based content review detect complementary failure classes.
Deletion itself became a design question rather than an implementation detail. Against the prevailing assumption that compilation should converge on the most current possible summary, a July 2026 paper presented an append-only, agent-aware wiki template built to preserve "failure paths" — dead ends, walked-back claims, and abandoned iterations — on the argument that publications and shared code structurally discard the record that would stop later researchers from repeating the same failures. Its central case study is a two-author project whose retroactive audit revised two prior experiments' claimed 20-of-20 coverage down to 14 and 12 evidence-based answers, and then to 18 and 18 after a fix, with the original overstatement and its correction both retained.
By July 2026 the pattern had moved into enterprise products, with Qihoo 360's [[organizations/360-ai-knowledge-base|AI Knowledge Base 3.0]] built around an LLM Wiki knowledge engine, CLPS's Project Athena, WPS Comate Wiki, and Jiran's Process Wisdom Star, alongside the first corporate LLM Wiki training programs in South Korea and a dedicated Wiki AI Pre-Conference at Wikimania 2026; WPS's knowledge-base "AI Wiki" function went live by August 2026. Tencent's [[WeKnora]] brought an open-source platform to the same space, making Wiki Mode generally available at release v0.5.0 — agent-generated, interlinked Markdown pages with an interactive knowledge graph — and pairing it with multi-tenant role-based access control, per-tenant audit logs, and ingest scaled to 40,000-document knowledge bases, reaching roughly 21,600 stars by early September 2026. Dnotitia, a Seoul-based long-term-memory AI vendor, open-sourced AKB (Agent Knowledge Base) in a release dated 20 May 2026, an agent-native platform that accumulates agent work context under team-, role-, and project-level access control.[^c11]
In September 2026 the pattern received its first first-party engineering account from a major platform company. Meta described an internal "organizational second brain," built for a compliance domain, as an agent that acts as a secondary expert for a given domain[^c12], whose institutional knowledge lives in more than two hundred version-controlled Markdown files under a strict taxonomy rather than in model weights, with expert corrections compiled into verified, regression-tested edits that require no retraining; Meta reported individual assessment time falling from days to minutes[^c13] and stated that the industry has converged on a similar idea[^c14], naming Karpathy's LLM Wiki and Google's Open Knowledge Format as the independent signals of that convergence. A second commercial production instance appeared earlier the same year: Perplexity's Brain records a per-task context graph of connectors used, valid sources, user edits and failed attempts, then aggregates the data overnight to update a personal large language model wiki loaded into the agent's execution environment before the next task[^c15]. Perplexity reported 25 percent higher accuracy on repetitive tasks, 16 percent better recall and 13 percent lower cost, while conceding that these are internal early metrics rather than external benchmarks[^c16] and that the system does not make the underlying model itself smarter[^c17] — the gain is attributed to accumulated context, not model capability. The architectural reading that accompanied both launches was one of layering rather than replacement: retrieval demoted to a slower tier beneath a compiled cache[^c18].
A parallel research line moved the empirical question from whether compilation beats retrieval to which structure makes it pay. SearchWiki compiles a corpus into a three-layer typed wiki of document overviews, cross-document topic pages, and page-level source records, then trains a 9B agent by reinforcement learning to navigate it, reporting that learned navigation over structured corpora is a superior alternative to flat retrieval[^c19]. WikiSkill co-evolved agent skills with a persistent wiki and showed by ablation that persistent knowledge accumulation is critical for effective skill evolution, with evolved skills transferring across model families[^c20]. A third study treated the curated store itself as the trained object and found that its advantage grows with overlap with the training questions, matching a graph-based baseline at roughly one percent of the links per point of corpus covered[^c21]. Community implementations published their failure rates alongside their gains: a newsroom-style system that separates authoring from review reported 107 adopted rule amendments against 72 rejections, and found that 44 of the 70 defect classes that had ever received a fix recurred afterwards[^c22].
A parallel vocabulary formed around the same idea during 2026: the "brain stack" distinguishes a personal second brain, an organizational company brain, and a shared brain coordinating multiple agents, with a cluster of startups racing to define the organizational layer as a typed, versioned semantic graph that stays current without manual wiki edits — a formulation whose critics note that retrieval-only products cover perhaps 40 percent of the problem, leaving joins, entitlement enforcement, and point-in-time reproducibility unaddressed. Implementation-level experimentation continued in parallel: a tiered L1-rule and L2-wiki cache architecture with log-driven page eviction, deterministic pre-commit contradiction gates paired with cross-provider citation audits, an installable bootstrap skill that packages the pattern with its schema contract and an optional local BM25 search layer, and a zero-code GitHub Copilot deployment each demonstrated the pattern's portability across runtimes and toolchains. Community implementations range from personal Obsidian workflows and git-repository knowledge bases to hosted services driven by Claude over the Model Context Protocol. The emerging consensus is hybrid: compiled wikis serve stable, high-frequency, curated knowledge while RAG continues to serve long-tail, fast-changing corpora, with the two increasingly layered in production systems. That convergence ran in both directions — in August 2026 the open-source retrieval framework RAGFlow added Wiki, Graph, Tree, Page Index, Mind Map, and Timeline as document- and dataset-level knowledge compilation targets, folding wiki compilation into mainstream RAG tooling as one output format among several.
The organizational end of the pattern consolidated under its own name during 2026. In its Requests for Startups, Y Combinator named the company brain as a missing primitive, describing a system that pulls knowledge out of every fragmented source, structures it, keeps it current, and renders it as an executable skills file for agents — not a search tool and not a chatbot over documents, but a living map of how a company actually works. The framing rests on a shift in diagnosis rather than a new model: frontier model quality stopped being the bottleneck for enterprise AI while organizational context became it, so companies have data but not memory. Four systems built independently had already converged on the same architecture — Karpathy's filed-back wiki, Garry Tan's GBrain with its routing discipline, Hannah Stulberg's Team OS at DoorDash, and Ramp's internal Glass skills marketplace — and the pattern they share is a portable Markdown substrate, a routing layer that tells the agent where things live, skills as the unit of work, a habit of writing results back, and a separation of context from compute. Vendor counterparts such as DevRev's Computer Memory articulate the enterprise version, adding continuous bidirectional synchronization with systems of record and permissions enforced at every node in place of a flat repository that anyone with access can read.
Two further strands sharpened the picture during 2026. On quality, deterministic structural gates and model-based content review were shown to detect complementary failure classes, while an independent release-governance study found that evidence coverage was the primary discriminator of severe regressions. On architecture, a sustained critique argued that Markdown directories lack the referential integrity, schema, permissions, query language, and auditability of a database, and that the pattern should therefore be understood as a personal or agent-facing compilation layer — valuable where a document set is stable and frequently read, and dependent on human curation and verification wherever accuracy is consequential. On law and licensing, the wider AI copyright reckoning reached the pattern's raw materials: the Bartz v Anthropic settlement established a benchmark of roughly $3,000 per pirated work, the Wikimedia Foundation disclosed that OpenAI-operated agents had crawled its platform without authorisation, and a community open letter argued that the attribution and share-alike obligations of CC BY-SA cannot be satisfied by model training and output, even as permissive licences such as the CC0-dedicated awesome-llm-wiki repository and MIT-licensed implementations remained the community default. The most widely adopted personal toolchain reflects that division of labour directly, pairing an Obsidian vault of plain Markdown with an agent that maintains it, and reserving human judgment for sourcing, review, and adjudication of contradictions. Benchmarking of the leading personal implementation is likewise self-reported rather than independent: GBrain's published BrainBench figures of 49.1% precision and 97.9% recall at five are measured on a 240-page synthetic corpus with 145 queries, and the headline 31.4-point gain comes from comparing the system against its own graph-disabled variant rather than against other memory systems, so the figure describes the contribution of typed-edge traversal within one codebase rather than a cross-system ranking.
The agent-memory research community that surrounds the pattern also matured its measurement apparatus during 2026. LongMemEval, LoCoMo, and BEAM became the standard benchmarks for comparing memory architectures[^c1], each with its own scale, question taxonomy and scoring regime ([[research/agent-memory-benchmarks|Agent-Memory Benchmarks]]), with BEAM's ten-million-token scale identified as the setting that context-window expansion alone cannot solve[^c2]; Mem0 reported 92.5 on LoCoMo and 94.4 on LongMemEval at roughly 6,900 tokens per query[^c3]. A solo-developed open-source system, agentmemory, reported a real-retrieval record of 96.20 percent on LongMemEval[^c4], and its account of reaching that figure is as instructive as the number: four successive configurations plateaued at exactly the same score until non-determinism in the approximate-nearest-neighbour index was removed, so a meaningful share of the apparent headroom at the top of the leaderboard was measurement noise rather than capability[^c5]. Diagnostic work moved in the opposite direction, decomposing memory systems into summarization, storage, and retrieval so that a wrong answer can be attributed to a specific operation rather than to the system as a whole[^c6]. A benchmark built on real agent trajectories found that existing memory systems fail chiefly because they lack causal and objective structure and rely on lossy similarity-based retrieval[^c7], and a hierarchical alternative improved accuracy by an average of 9.97 percentage points over linear memory while cutting prompt-token usage by 32.8 percent, on the argument that effective long-horizon memory depends less on storing more than on deciding what stays active[^c8][^c9]. A parallel 2026 comparison placed RAG, the LLM Wiki pattern, and agentic search as three positions on a cost-latency-quality trade-off rather than successive replacements, projecting that 75 percent of enterprise applications would adopt hybrid architectures combining all three by the end of 2026[^c10].