{"article":{"slug":"graft-metatron-and-the-two-kinds-of-context-coding-agents-need","title":"Graft, Metatron, and the two kinds of context coding agents need","subtitle":null,"summary":"Pavel Kerbel contrasts Graft’s recoverable WHAT/WHERE code maps with Metatron’s reviewed WHY/WHY NOT engineering memory, arguing stronger models still need both layers—and proposing a factorial eval to prove it.","content_type":"essay","language":"en","canonical_url":"https://getmetatron.com/blog/graft-vs-metatron/","author":{"name":"Pavel Kerbel","url":"https://getmetatron.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Metatron","url":"https://getmetatron.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Developer Tools","slug":"developer-tools","url":"https://listedarticles.com/topics/developer-tools"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2075,"reading_minutes":9,"published_at":"2026-09-06T00:00:00.000Z","added_at":"2026-09-30T06:13:52.773Z","updated_at":"2026-09-30T06:13:52.773Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/graft-metatron-and-the-two-kinds-of-context-coding-agents-need","markdown_url":"https://listedarticles.com/articles/graft-metatron-and-the-two-kinds-of-context-coding-agents-need.md","example":false,"citation":"Pavel Kerbel, Metatron. \"Graft, Metatron, and the two kinds of context coding agents need.\" 6 Sept 2026. https://getmetatron.com/blog/graft-vs-metatron/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://getmetatron.com/blog/graft-vs-metatron/"},"body_markdown":"[Home/[Blog/Graft and Metatron\n    \n        Source note: I reviewed Graft 0.17.0 at\n        [`05760b0`,\n        Metatron 0.13.0 at\n        [`8cd80db`,\n        and both projects' current READMEs, code, examples, and published benchmark material on September 6, 2026.\n        Graft moves quickly; follow the linked README for the latest feature and benchmark claims.\n\n## Two projects, one shared observation\n\nEvery fresh coding-agent session pays an onboarding tax. The agent greps for a\n        name, opens a file, follows an import, finds a caller, backs out, and tries another\n        path. Some of that exploration is necessary. Much of it reconstructs facts that a\n        previous session — or a teammate — already reconstructed.\n\nGraft and Metatron both begin there: repository context should survive longer than\n        one agent run. But “context” is doing too much work in that sentence. A map of the\n        current implementation and a record of the team's engineering judgment are both\n        context, in the same way that a street map and a local's warning about a flood-prone\n        road are both navigation. One tells you what is there. The other tells you what\n        happened before.\n\n      \n        Graft is primarily WHAT / WHERE: symbols, calls, imports,\n        dependencies, affected files.\n\n        Metatron is primarily WHY / WHY NOT: rationale, constraints,\n        rejected approaches, lessons, conventions.\n      \n## What Graft is solving\n\nGraft is unusually serious about structural code intelligence. A plain\n        `graft build` uses tree-sitter to produce a per-symbol wiring graph and\n        per-file cards. The graph connects functions, classes, calls, imports, inheritance,\n        and implementations. Its query surface turns that structure into practical agent\n        operations: find relevant code, show a file's API without its bodies, trace callers\n        and dependencies, search by enclosing symbol, build a token-budgeted repository\n        map, and calculate the blast radius of a diff. An optional LSP layer adds\n        compiler-resolved edges where static syntax is not enough.\n\nThe optional `--deep` build adds model-written file and symbol summaries,\n        concept nodes, and small “crux” excerpts containing the lines that carry the logic.\n        Sources are content-hashed; structural refreshes track the working tree, including\n        uncommitted changes. The generated graph is plain markdown and JSON, so an agent can\n        use its ordinary file-reading habits rather than depending on a proprietary search\n        service. Graft can also expose the same capabilities as MCP tools.\n\nThat is a strong answer to a real problem. A call graph can tell an agent that an\n        authorization helper has twelve callers. A repository map can reveal that a small\n        file is a highly connected hub. Blast-radius analysis can keep a one-file patch from\n        ignoring four siblings. These are exactly the situations where naive grep-and-open\n        exploration is slow and brittle.\n\n        Graft also publishes a\n        [repository-specific PocketBase evaluation:\n        across ten architecture questions and five implementation tasks, both arms touched\n        the same files as the maintainers on all five implementations, while the Graft arm\n        used 21% less cost and 14% less wall-clock time. “Same files” is a localization proxy\n        rather than a test-suite verdict, but it is well matched to Graft's navigation claim.\n\n## What Metatron is solving\n\nMetatron stores a different unit: a reviewed engineering decision. Each decision\n        has a pattern, scope, rationale, confidence, and source references. The default\n        files-first mode keeps those decisions as markdown under `context/`, next\n        to the code. An agent is instructed to consult them before planning and to record\n        durable lessons it discovers while working.\n\nThe lifecycle is roughly:\n\n```\nconsult → work → discover → propose → review → canonical decision\n```\n\nWith the default pull-request review gate, an agent writes the proposed decision on\n        its working branch and human review of the PR is the curation act. Teams that want a\n        separate queue can stage it under `context/candidate/` before promotion.\n        Metatron's optional MCP mode adds relevance-ranked serving, feedback, and an explicit\n        candidate store, but the invariant is the same: nothing self-promotes into canonical\n        knowledge.\n\nMetatron can bootstrap candidates from structural signals and Git history, but that\n        is not the whole system. The important part is the write-back loop. An agent can\n        discover a constraint during a failed implementation, record it, and offer it for\n        review so that a future agent does not need to repeat the failure.\n\n## Recoverable context and non-recoverable context\n\nHere is the conceptual center of the comparison. Most of Graft's core facts are\n        recoverable from the current repository. Given enough time, tokens, and tool\n        calls, a sufficiently capable agent can parse the symbols, follow the imports, read\n        the bodies, inspect the tests, and reconstruct much of the architecture. Graft makes\n        that reconstruction faster, cheaper, and more dependable. “Recoverable” does not\n        mean “free.”\n\nMany Metatron decisions are not recoverable from the current tree because the\n        missing information is historical or counterfactual. The code tells you which\n        approach survived. It may not tell you which plausible alternatives were tried,\n        which production incident killed them, or which constraint the team expects every\n        future change to preserve.\n\n```\nDo not move JWT validation into every service. We tried that previously and\nit caused authorization semantics to diverge across services. Keep validation\nat the gateway.\n```\n\nAn AST or call graph can show that JWT validation happens at the gateway today. It\n        can show every service depending on the authenticated request. A generated summary\n        can accurately describe the current design. None of those artifacts necessarily\n        contains the causal statement: validation used to be distributed, the semantics\n        diverged, and the team deliberately centralized it. If that history never made it\n        into code, tests, a commit message, or documentation, model capability cannot infer\n        it reliably. It can only guess.\n\nThe boundary is not perfectly clean. Graft's node format includes a human Notes\n        region that survives regeneration, so teams can add knowledge beyond generated\n        summaries. Metatron's ingest also starts from recoverable structural and Git\n        signals. The distinction is about each project's center of gravity and default\n        lifecycle: Graft treats its generated `graft/` graph as a local,\n        gitignored, regenerable cache; Metatron treats reviewed decisions as shared,\n        Git-tracked source material.\n\n## What happens as coding models get stronger?\n\nA tempting argument says better models will make code maps unnecessary. I do not\n        think the evidence supports it. We need to separate two outcomes:\n\n- Correctness: does the extra context help the agent produce the\n          right answer or patch?\n- Efficiency: does it reduce tokens, tool calls, API requests,\n          wall-clock time, or cost?\nA strong model can reach a correctness ceiling while still wasting work getting\n        there. Graft's own published results are a useful example:\n\n               | Graft benchmark\n               | Correctness\n               | Efficiency reported by Graft\n            \n          \n          \n            \n               | [Controlled question benchmark\n162 runs, two repos, Claude Sonnet 5\n               | Cold 93%\nGraft push 93%\n               | 42% fewer tokens, 46% fewer tool calls, 60% less latency, 32% lower mean cost per task\n            \n            \n               | [SWE-bench Verified subset\n50 instances, Claude Sonnet 5, official grader\n               | Cold 27/50 (54%)\nGraft 33/50 (66%)\n               | 23% fewer tokens, 25% fewer tool calls, 24% fewer API requests, 32% less wall-clock time\n            \n          \n        \n      In the first benchmark, correctness was already 93%, yet the push configuration\n        removed a large share of the exploration cost. In the lower-baseline SWE-bench run,\n        Graft reported both a 12-point correctness gain and efficiency gains. That is\n        consistent with the idea that correctness gains shrink near a ceiling while\n        efficiency gains remain. It does not prove it: the task sets, graders, and\n        conditions differ, the controlled benchmark used a model judge, and the SWE-bench\n        run covered 50 instances. Graft's README also notes that its SWE-bench efficiency\n        totals are calculated on instances both arms resolved for a like-for-like comparison.\n\n        Metatron's own research reaches a similarly qualified boundary. Our\n        [Context Inheritance paper,\n        accepted at the forthcoming AgenticDev 2026 workshop, co-located with ASE,\n        reports that blind\n        frontier-authored context raised an 8B local model's gold-file localization from\n        21.1% to 47.2% on topic-matched tasks (+26.1 percentage points, p<0.0001).\n        The same executor's own lessons moved localization by 0.0 points. Resolve stayed\n        low and did not improve significantly at that tier, so this is a localization\n        result, not a general claim that the context doubled bug-fixing ability.\n\nOn constraint-sharing tasks, the frontier lifecycle moved resolve from 58.3% to\n        72.9%, but that result was directional rather than confirmed after the registered\n        multiple-comparison correction (raw p=0.041; Holm-adjusted\n        p=0.081). On a broader 88-instance sample, the same frontier executor\n        already resolved 90.9% without context and no delivery condition beat it. That is\n        evidence of a possible task-specific ceiling, not a universal law about models or\n        memory.\n\n        The paper evaluated the Repository Context Layer architecture in a minimal harness,\n        not Metatron's product end to end, and it did not test Metatron's automatic ingest.\n        The protocol, transcripts, raw outcomes, and verification scripts are available in\n        the [public replication artifact.\n\n## Why Git matters\n\nBoth projects use files and both understand Git, but Git plays a different role.\n        Graft uses Git to decide which working-tree files are visible and keeps its graph\n        fresh as those files change. By default, `graft build` adds the generated\n        `graft/` directory to `.gitignore`; teammates regenerate it\n        locally. That is sensible for derived structural state. A stale map should be\n        rebuilt from the source of truth.\n\nIn Metatron, the decision files themselves are source of truth. Git supplies\n        provenance, review, branching, blame, rollback, and temporal alignment with the\n        code. If a team changes the JWT rule, the context diff can land with the code diff.\n        If the change was wrong, both can be reverted. If two branches disagree, the\n        disagreement is visible rather than silently resolved by regeneration.\n\nThis is why the human gate matters. Persistent context is standing instruction to\n        future agents. A plausible but wrong code summary is inconvenient; a plausible but\n        wrong rule that every future agent obeys is institutional damage. Metatron makes\n        review part of the storage model because durable knowledge requires an owner and an\n        audit trail.\n\n## Why Graft and Metatron may belong together\n\nOnce the recoverable/non-recoverable distinction is explicit, the projects stop\n        looking like substitutes. An agent preparing to change authentication needs both a\n        current map of the gateway's callers and the historical warning against distributing\n        validation. One cannot replace the other.\n\nThe ideal context stack may therefore have three layers:\n\n- Source code: the executable ground truth.\n- Structural code intelligence: a Graft-like layer for symbols,\n          calls, dependencies, architecture, repository navigation, and blast radius.\n- Engineering memory: a Metatron-like layer for rationale,\n          constraints, rejected alternatives, edge cases, and reviewed lessons.\nThe structural layer reduces the price of understanding the present. The memory\n        layer makes parts of the past available at all.\n\n## The experiment I would like to see\n\nThe clean test is a factorial study, not two unrelated benchmark percentages placed\n        side by side. Run the same tasks, model, scaffold, tools, budget, and grading under\n        four conditions:\n\n```\n1. Cold baseline\n2. Graft\n3. Metatron\n4. Graft + Metatron\n```\n\nRepeat the matrix across several model capability levels. Measure correctness,\n        tokens, tool calls, API requests, latency, and cost. Include ordinary bug fixing,\n        but also tasks where success depends on a recurring team constraint or a previously\n        rejected approach — cases that test engineering memory rather than repository\n        localization alone.\n\nMy hypothesis is an interaction, not a winner. Structural context should help most\n        on large, unfamiliar, cross-file changes and may retain efficiency gains after its\n        correctness effect narrows. Reviewed engineering memory should help most when the\n        decisive fact is historical, conventional, or absent from the current code. The\n        combined arm should show whether faster navigation makes the right decision easier\n        to apply, or whether the two contexts sometimes compete for a finite attention\n        budget. Any of those outcomes would teach us more than “tool A scored X and tool B\n        scored Y” across incomparable experiments.\n\n## FAQ\n\nWhat is the main difference between Graft and Metatron?\n\n        Graft primarily derives structural intelligence from the current repository.\n        Metatron primarily preserves reviewed engineering decisions. In shorthand: Graft\n        answers WHAT / WHERE; Metatron answers WHY / WHY NOT.\n\nDo stronger coding models make Graft-like tools unnecessary?\n\n        No. Correctness gains may narrow on tasks a model already solves, while token,\n        tool-call, and latency gains remain valuable. Graft's controlled benchmark is a\n        concrete example: equal 93% correctness with substantially lower reported usage.\n\nCan the projects be used together?\n\n        Conceptually, yes. Their default artifacts and integration mechanisms differ, but a\n        coding agent can benefit from a current structural map and a reviewed decision\n        history in the same repository.\n\nGraft asks why an agent should rediscover the structure of a repository on every\n        task. Metatron asks the next question: why should it rediscover the team's engineering\n        knowledge?\n\nA capable agent can always spend more effort reading code. It cannot read an argument\n        nobody saved.\n\n      \n        \n          [Read Graft's source and benchmarks\n          [Explore Metatron\n          [Read the AgenticDev paper\n          [Inspect the research artifact\n        \n          \n          Written by Pavel Kerbel, creator of Metatron and co-author of the Context Inheritance study.","body_html":"<p>[Home/[Blog/Graft and Metatron</p>\n<pre><code>    Source note: I reviewed Graft 0.17.0 at\n    [`05760b0`,\n    Metatron 0.13.0 at\n    [`8cd80db`,\n    and both projects&#39; current READMEs, code, examples, and published benchmark material on September 6, 2026.\n    Graft moves quickly; follow the linked README for the latest feature and benchmark claims.</code></pre>\n<h2 id=\"two-projects-one-shared-observation\">Two projects, one shared observation</h2>\n<p>Every fresh coding-agent session pays an onboarding tax. The agent greps for a\n        name, opens a file, follows an import, finds a caller, backs out, and tries another\n        path. Some of that exploration is necessary. Much of it reconstructs facts that a\n        previous session — or a teammate — already reconstructed.</p>\n<p>Graft and Metatron both begin there: repository context should survive longer than\n        one agent run. But “context” is doing too much work in that sentence. A map of the\n        current implementation and a record of the team&#39;s engineering judgment are both\n        context, in the same way that a street map and a local&#39;s warning about a flood-prone\n        road are both navigation. One tells you what is there. The other tells you what\n        happened before.</p>\n<pre><code>    Graft is primarily WHAT / WHERE: symbols, calls, imports,\n    dependencies, affected files.\n\n    Metatron is primarily WHY / WHY NOT: rationale, constraints,\n    rejected approaches, lessons, conventions.</code></pre>\n<h2 id=\"what-graft-is-solving\">What Graft is solving</h2>\n<p>Graft is unusually serious about structural code intelligence. A plain\n        <code>graft build</code> uses tree-sitter to produce a per-symbol wiring graph and\n        per-file cards. The graph connects functions, classes, calls, imports, inheritance,\n        and implementations. Its query surface turns that structure into practical agent\n        operations: find relevant code, show a file&#39;s API without its bodies, trace callers\n        and dependencies, search by enclosing symbol, build a token-budgeted repository\n        map, and calculate the blast radius of a diff. An optional LSP layer adds\n        compiler-resolved edges where static syntax is not enough.</p>\n<p>The optional <code>--deep</code> build adds model-written file and symbol summaries,\n        concept nodes, and small “crux” excerpts containing the lines that carry the logic.\n        Sources are content-hashed; structural refreshes track the working tree, including\n        uncommitted changes. The generated graph is plain markdown and JSON, so an agent can\n        use its ordinary file-reading habits rather than depending on a proprietary search\n        service. Graft can also expose the same capabilities as MCP tools.</p>\n<p>That is a strong answer to a real problem. A call graph can tell an agent that an\n        authorization helper has twelve callers. A repository map can reveal that a small\n        file is a highly connected hub. Blast-radius analysis can keep a one-file patch from\n        ignoring four siblings. These are exactly the situations where naive grep-and-open\n        exploration is slow and brittle.</p>\n<pre><code>    Graft also publishes a\n    [repository-specific PocketBase evaluation:\n    across ten architecture questions and five implementation tasks, both arms touched\n    the same files as the maintainers on all five implementations, while the Graft arm\n    used 21% less cost and 14% less wall-clock time. “Same files” is a localization proxy\n    rather than a test-suite verdict, but it is well matched to Graft&#39;s navigation claim.</code></pre>\n<h2 id=\"what-metatron-is-solving\">What Metatron is solving</h2>\n<p>Metatron stores a different unit: a reviewed engineering decision. Each decision\n        has a pattern, scope, rationale, confidence, and source references. The default\n        files-first mode keeps those decisions as markdown under <code>context/</code>, next\n        to the code. An agent is instructed to consult them before planning and to record\n        durable lessons it discovers while working.</p>\n<p>The lifecycle is roughly:</p>\n<pre><code>consult → work → discover → propose → review → canonical decision</code></pre>\n<p>With the default pull-request review gate, an agent writes the proposed decision on\n        its working branch and human review of the PR is the curation act. Teams that want a\n        separate queue can stage it under <code>context/candidate/</code> before promotion.\n        Metatron&#39;s optional MCP mode adds relevance-ranked serving, feedback, and an explicit\n        candidate store, but the invariant is the same: nothing self-promotes into canonical\n        knowledge.</p>\n<p>Metatron can bootstrap candidates from structural signals and Git history, but that\n        is not the whole system. The important part is the write-back loop. An agent can\n        discover a constraint during a failed implementation, record it, and offer it for\n        review so that a future agent does not need to repeat the failure.</p>\n<h2 id=\"recoverable-context-and-non-recoverable-context\">Recoverable context and non-recoverable context</h2>\n<p>Here is the conceptual center of the comparison. Most of Graft&#39;s core facts are\n        recoverable from the current repository. Given enough time, tokens, and tool\n        calls, a sufficiently capable agent can parse the symbols, follow the imports, read\n        the bodies, inspect the tests, and reconstruct much of the architecture. Graft makes\n        that reconstruction faster, cheaper, and more dependable. “Recoverable” does not\n        mean “free.”</p>\n<p>Many Metatron decisions are not recoverable from the current tree because the\n        missing information is historical or counterfactual. The code tells you which\n        approach survived. It may not tell you which plausible alternatives were tried,\n        which production incident killed them, or which constraint the team expects every\n        future change to preserve.</p>\n<pre><code>Do not move JWT validation into every service. We tried that previously and\nit caused authorization semantics to diverge across services. Keep validation\nat the gateway.</code></pre>\n<p>An AST or call graph can show that JWT validation happens at the gateway today. It\n        can show every service depending on the authenticated request. A generated summary\n        can accurately describe the current design. None of those artifacts necessarily\n        contains the causal statement: validation used to be distributed, the semantics\n        diverged, and the team deliberately centralized it. If that history never made it\n        into code, tests, a commit message, or documentation, model capability cannot infer\n        it reliably. It can only guess.</p>\n<p>The boundary is not perfectly clean. Graft&#39;s node format includes a human Notes\n        region that survives regeneration, so teams can add knowledge beyond generated\n        summaries. Metatron&#39;s ingest also starts from recoverable structural and Git\n        signals. The distinction is about each project&#39;s center of gravity and default\n        lifecycle: Graft treats its generated <code>graft/</code> graph as a local,\n        gitignored, regenerable cache; Metatron treats reviewed decisions as shared,\n        Git-tracked source material.</p>\n<h2 id=\"what-happens-as-coding-models-get-stronger\">What happens as coding models get stronger?</h2>\n<p>A tempting argument says better models will make code maps unnecessary. I do not\n        think the evidence supports it. We need to separate two outcomes:</p>\n<ul><li><p>Correctness: does the extra context help the agent produce the</p><pre><code>    right answer or patch?</code></pre></li><li><p>Efficiency: does it reduce tokens, tool calls, API requests,</p><pre><code>    wall-clock time, or cost?</code></pre>\n<p>A strong model can reach a correctness ceiling while still wasting work getting\n      there. Graft&#39;s own published results are a useful example:</p>\n<pre><code>         | Graft benchmark\n         | Correctness\n         | Efficiency reported by Graft\n\n\n\n\n         | [Controlled question benchmark</code></pre>\n<p>162 runs, two repos, Claude Sonnet 5\n             | Cold 93%\nGraft push 93%\n             | 42% fewer tokens, 46% fewer tool calls, 60% less latency, 32% lower mean cost per task</p>\n<pre><code>         | [SWE-bench Verified subset</code></pre>\n<p>50 instances, Claude Sonnet 5, official grader\n             | Cold 27/50 (54%)\nGraft 33/50 (66%)\n             | 23% fewer tokens, 25% fewer tool calls, 24% fewer API requests, 32% less wall-clock time</p>\n<pre><code>In the first benchmark, correctness was already 93%, yet the push configuration\n  removed a large share of the exploration cost. In the lower-baseline SWE-bench run,\n  Graft reported both a 12-point correctness gain and efficiency gains. That is\n  consistent with the idea that correctness gains shrink near a ceiling while\n  efficiency gains remain. It does not prove it: the task sets, graders, and\n  conditions differ, the controlled benchmark used a model judge, and the SWE-bench\n  run covered 50 instances. Graft&#39;s README also notes that its SWE-bench efficiency\n  totals are calculated on instances both arms resolved for a like-for-like comparison.\n\n  Metatron&#39;s own research reaches a similarly qualified boundary. Our\n  [Context Inheritance paper,\n  accepted at the forthcoming AgenticDev 2026 workshop, co-located with ASE,\n  reports that blind\n  frontier-authored context raised an 8B local model&#39;s gold-file localization from\n  21.1% to 47.2% on topic-matched tasks (+26.1 percentage points, p&lt;0.0001).\n  The same executor&#39;s own lessons moved localization by 0.0 points. Resolve stayed\n  low and did not improve significantly at that tier, so this is a localization\n  result, not a general claim that the context doubled bug-fixing ability.</code></pre></li></ul>\n<p>On constraint-sharing tasks, the frontier lifecycle moved resolve from 58.3% to\n        72.9%, but that result was directional rather than confirmed after the registered\n        multiple-comparison correction (raw p=0.041; Holm-adjusted\n        p=0.081). On a broader 88-instance sample, the same frontier executor\n        already resolved 90.9% without context and no delivery condition beat it. That is\n        evidence of a possible task-specific ceiling, not a universal law about models or\n        memory.</p>\n<pre><code>    The paper evaluated the Repository Context Layer architecture in a minimal harness,\n    not Metatron&#39;s product end to end, and it did not test Metatron&#39;s automatic ingest.\n    The protocol, transcripts, raw outcomes, and verification scripts are available in\n    the [public replication artifact.</code></pre>\n<h2 id=\"why-git-matters\">Why Git matters</h2>\n<p>Both projects use files and both understand Git, but Git plays a different role.\n        Graft uses Git to decide which working-tree files are visible and keeps its graph\n        fresh as those files change. By default, <code>graft build</code> adds the generated\n        <code>graft/</code> directory to <code>.gitignore</code>; teammates regenerate it\n        locally. That is sensible for derived structural state. A stale map should be\n        rebuilt from the source of truth.</p>\n<p>In Metatron, the decision files themselves are source of truth. Git supplies\n        provenance, review, branching, blame, rollback, and temporal alignment with the\n        code. If a team changes the JWT rule, the context diff can land with the code diff.\n        If the change was wrong, both can be reverted. If two branches disagree, the\n        disagreement is visible rather than silently resolved by regeneration.</p>\n<p>This is why the human gate matters. Persistent context is standing instruction to\n        future agents. A plausible but wrong code summary is inconvenient; a plausible but\n        wrong rule that every future agent obeys is institutional damage. Metatron makes\n        review part of the storage model because durable knowledge requires an owner and an\n        audit trail.</p>\n<h2 id=\"why-graft-and-metatron-may-belong-together\">Why Graft and Metatron may belong together</h2>\n<p>Once the recoverable/non-recoverable distinction is explicit, the projects stop\n        looking like substitutes. An agent preparing to change authentication needs both a\n        current map of the gateway&#39;s callers and the historical warning against distributing\n        validation. One cannot replace the other.</p>\n<p>The ideal context stack may therefore have three layers:</p>\n<ul><li>Source code: the executable ground truth.</li><li><p>Structural code intelligence: a Graft-like layer for symbols,</p><pre><code>    calls, dependencies, architecture, repository navigation, and blast radius.</code></pre></li><li><p>Engineering memory: a Metatron-like layer for rationale,</p><pre><code>    constraints, rejected alternatives, edge cases, and reviewed lessons.</code></pre>\n<p>The structural layer reduces the price of understanding the present. The memory\n      layer makes parts of the past available at all.</p></li></ul>\n<h2 id=\"the-experiment-i-would-like-to-see\">The experiment I would like to see</h2>\n<p>The clean test is a factorial study, not two unrelated benchmark percentages placed\n        side by side. Run the same tasks, model, scaffold, tools, budget, and grading under\n        four conditions:</p>\n<pre><code>1. Cold baseline\n2. Graft\n3. Metatron\n4. Graft + Metatron</code></pre>\n<p>Repeat the matrix across several model capability levels. Measure correctness,\n        tokens, tool calls, API requests, latency, and cost. Include ordinary bug fixing,\n        but also tasks where success depends on a recurring team constraint or a previously\n        rejected approach — cases that test engineering memory rather than repository\n        localization alone.</p>\n<p>My hypothesis is an interaction, not a winner. Structural context should help most\n        on large, unfamiliar, cross-file changes and may retain efficiency gains after its\n        correctness effect narrows. Reviewed engineering memory should help most when the\n        decisive fact is historical, conventional, or absent from the current code. The\n        combined arm should show whether faster navigation makes the right decision easier\n        to apply, or whether the two contexts sometimes compete for a finite attention\n        budget. Any of those outcomes would teach us more than “tool A scored X and tool B\n        scored Y” across incomparable experiments.</p>\n<h2 id=\"faq\">FAQ</h2>\n<p>What is the main difference between Graft and Metatron?</p>\n<pre><code>    Graft primarily derives structural intelligence from the current repository.\n    Metatron primarily preserves reviewed engineering decisions. In shorthand: Graft\n    answers WHAT / WHERE; Metatron answers WHY / WHY NOT.</code></pre>\n<p>Do stronger coding models make Graft-like tools unnecessary?</p>\n<pre><code>    No. Correctness gains may narrow on tasks a model already solves, while token,\n    tool-call, and latency gains remain valuable. Graft&#39;s controlled benchmark is a\n    concrete example: equal 93% correctness with substantially lower reported usage.</code></pre>\n<p>Can the projects be used together?</p>\n<pre><code>    Conceptually, yes. Their default artifacts and integration mechanisms differ, but a\n    coding agent can benefit from a current structural map and a reviewed decision\n    history in the same repository.</code></pre>\n<p>Graft asks why an agent should rediscover the structure of a repository on every\n        task. Metatron asks the next question: why should it rediscover the team&#39;s engineering\n        knowledge?</p>\n<p>A capable agent can always spend more effort reading code. It cannot read an argument\n        nobody saved.</p>\n<pre><code>      [Read Graft&#39;s source and benchmarks\n      [Explore Metatron\n      [Read the AgenticDev paper\n      [Inspect the research artifact\n    \n      \n      Written by Pavel Kerbel, creator of Metatron and co-author of the Context Inheritance study.</code></pre>","headings":[{"level":2,"text":"Two projects, one shared observation","id":"two-projects-one-shared-observation"},{"level":2,"text":"What Graft is solving","id":"what-graft-is-solving"},{"level":2,"text":"What Metatron is solving","id":"what-metatron-is-solving"},{"level":2,"text":"Recoverable context and non-recoverable context","id":"recoverable-context-and-non-recoverable-context"},{"level":2,"text":"What happens as coding models get stronger?","id":"what-happens-as-coding-models-get-stronger"},{"level":2,"text":"Why Git matters","id":"why-git-matters"},{"level":2,"text":"Why Graft and Metatron may belong together","id":"why-graft-and-metatron-may-belong-together"},{"level":2,"text":"The experiment I would like to see","id":"the-experiment-i-would-like-to-see"},{"level":2,"text":"FAQ","id":"faq"}]}}