{"article":{"slug":"hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised","title":"hutch: local code reviews in emacs for the mildly disenfranchised","subtitle":null,"summary":"The author introduces Hutch, a Magit-based code review interface for Emacs that runs an LLM review agent locally on staged, unpushed or branch changes, favours verified SEARCH/REPLACE patches over prose comments, traces agent runs with Perfetto, and reports eval results across 40 PRs and several models.","content_type":"blog_post","language":"en","canonical_url":"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/","author":{"name":null,"url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"kitallis.in","url":"https://kitallis.in/","listing_slug":null,"listing":null},"topics":[{"name":"Emacs","slug":"emacs","url":"https://listedarticles.com/topics/emacs"},{"name":"Code Review","slug":"code-review","url":"https://listedarticles.com/topics/code-review"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Developer Tools","slug":"developer-tools","url":"https://listedarticles.com/topics/developer-tools"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1836,"reading_minutes":8,"published_at":"2026-10-05T00:00:00.000Z","added_at":"2026-10-06T08:12:16.259Z","updated_at":"2026-10-06T08:12:16.259Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised","markdown_url":"https://listedarticles.com/articles/hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised.md","example":false,"citation":"kitallis.in. \"hutch: local code reviews in emacs for the mildly disenfranchised.\" 5 Oct 2026. https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/"},"body_markdown":"# hutch: local code reviews in emacs for the mildly disenfranchised\n\nI haven't had a real job in four years. I closed down a [startup](https://tramline.app) I'd been building, just last month. During all these years, I spent most of that time at the back end of the frontier of AI agents. But I've finally caught up. It's been some [500 days](https://en.wikipedia.org/wiki/List_of_large_language_models#2025) since coding agents have really picked up, and they're genuinely more productive than, previously, [instructed](https://www.youtube.com/watch?v=U_cSLPv34xk).\n\nEven though I still prefer the pedagogical aspect of AI over the task-completing automaton aspects, the latter has driven all sorts of tooling around reviewing code, and not just writing and deploying it. The typical review agent party-line is: agents jump in, before your colleagues do, spray logorrhea across twenty pull requests before you have had a chance to wake up and look at your phone. This works, sometimes, for some people. But if you're like me, you still have humans reviewing code before it ships to users, and it's better to respect those people and their time. This is the case, regardless of where you sit on the balance of game-changer to curmudgeon.\n\nAll that is to say, no matter which direction agents take to get better with time, I hope we still *care* about things. Not in the way of formalizing care, with high-fidelity agent instructions and prompts or some superior upholding of taste sort of thing, but something as simple as announcing: *hey I'm still here, and I understand all this*.\n\nSo as a long-time emacs user, I present yet another attempt at wedging LLMs, agents and coding harnesses, now inside your text buffers (!) with [Hutch](https://github.com/adjaecent/magit-hutch). It's a small, Magit-induced code-review interface that fits a standard Magit commit-push workflow *locally* and hopefully helps reclaim some load created upstream.\n\n## quick tour[#](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#quick_tour)\n\nOpen up Magit, and hit the dispatcher binding (usually `d`) and you'll see a `Hutch code review` action put up next to the DWIM binding. Hutch operates on three different scopes: staged changes, un-pushed changes, and changes between current branch and working branch. By default, it's staged changes only, since that's most useful.\n\nOnce a review starts, you'll see a nice little progress bar in a new `*magit-hutch: code review*` buffer until the findings[<sup>1</sup>](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fn-1) are complete. This is a read-only buffer, but you can still perform the [pre-bound actions](https://github.com/adjaecent/magit-hutch#usage).\n\nEach finding has a type (suggestion, comment or LGTM), a file name, relevant line numbers, a title and a description. Suggestions additionally have a patch diff. Hutch piggybacks on [Magit](https://magit.vc) and [Transient](https://github.com/magit/transient) to render these menus so it behaves much like its own interface; keyboard-driven sub-menus, diff coloring, and highlighting. Suggestions are special since they can be applied. To mark a suggestion for application, you queue it with `m`.\n\nThen bulk-apply all queued suggestions with `A`. The application is scope-aware, so if you queue a finding for staged changes, it will apply the fix directly to the staged files.\n\nThat's it! Getting started should hopefully be pretty simple and intuitive for existing emacs users. There are of course a few interesting things going on behind the scenes, some of which I'll cover in the next few sections.\n\n## patches over comments[#](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#patches_over_comments)\n\nA big UX handicap of showing review comments and suggested patches in-buffer is that there is no existing connective tissue of a commenting system. With GitHub, though, the review UI collapses outdated comments on new commits and most review bots sit over the [suggestion](https://docs.github.com/en/pull-requests/how-tos/review-pull-requests/incorporating-feedback-in-your-pull-request#applying-suggested-changes) mechanic if they have changes to suggest.\n\nHutch is made with a bias towards patches, rather than just prosaic comments. According to the [Aider leaderboard](https://aider.chat/docs/leaderboards) (and through some of my own experiments), the `SEARCH/REPLACE` diffs are a lot more obedient across different models than just asking the model to author correct patches with precise line numbers.\n\nFor an [Aider-style diff](https://aider.chat/docs/more/edit-formats.html), you have to ensure there's enough surrounding context for the `SEARCH` to be unique, and ideally also preserve indentation. In Hutch's case, the tool's function schema naturally decomposes the *file, search, and replace* fields:\n\n```\nsrc/utils.clj\n<<<<<<< SEARCH\n(defn add [a b]\n  (+ a b))\n=======\n(defn add [a b c]\n  (+ a b c))\n>>>>>>> REPLACE\n```\nI've noticed that a lot of older (or cheaper) models tend to recall the `SEARCH` block from memory when asked for diffs, instead of copying it verbatim from `read_file`, `read_diff` or `surrounding_context` calls. This invariably botches them entirely. So we get them verified before submission. If `SEARCH` is missing or matches more than once, the finding is downgraded to a plain comment. On a unique hit, Hutch locally creates a unified diff:\n\nOnce a series of udiffs and comments are rendered, they can be marked and bulk applied. Hutch applies them per-file, lowest hunk first (bottom-up) so line positions are minimally disturbed. Each finding runs its own `git apply` and a bad application marks itself `invalid` so the rest can continue to land.\n\nAll this patching and commenting infrastructure pulls its weight, since with only a couple of keystrokes, you hopefully get less reading and parsing work and more actionable triaging. None of this guarantees patches-always of course, and it shouldn't.\n\nWith more powerful models, a simpler diffing method might generally work pretty well. But for a tool that's built to work across different and cheaper models, it's essential to be maximally supportive. In general, I feel like a key point of much of the agentic infrastructure we build is to have knobs for optimizing token:cost ratios. This could often mean thorny workarounds for good-enough models.\n\n## barely enough tooling[#](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#barely_enough_tooling)\n\nHutch has a fairly minimal toolset for pulling context:\n\n1. `read_diff`\n2. `read_file`\n3. `search_codebase`\n4. `surrounding_context`\n\nOut of these, `surrounding_context` is the more interesting one. It wraps over [Tree-sitter](https://batsov.com/articles/2026/02/27/building-emacs-major-modes-with-treesitter-lessons-learned/#why-tree-sitter) and uses grammars that are installed. It works by letting the model widen out to the enclosing definition of a relevant line and further out, as needed. In my tests, the overall read token consumption compared to simply blasting `read_file` was anecdotally lower with comparable levels of review quality[<sup>2</sup>](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fn-2).\n\nAll the findings from the model are submitted to the agent at once. On the write side of things, `verify_block` locally verifies diffs, and along with other comments and LGTM notices, submits them through a `submit_review` tool call. `submit_review` itself runs through some post-processing work, like gating hallucinations about files and line numbers, trimming the length of descriptions and downgrading patches to comments if they don't apply cleanly.\n\nOnce the submission lands, the output from all this work is persisted durably under `refs/hutch/id` and can be separately committed as a means of sharing (with `magit-post-commit-hook`) or for repainting later. If you squint hard enough, it might appear like a change identifier for a stacked-diff [review tool](https://blog.tangled.org/stacking), but its purpose is to keep reviews in the git tree, rather than identify changesets for human reviews. We don't really care about multi-party human reviews, it's all local.\n\n## evaluating[#](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#evaluating)\n\nThe one unfortunate part about benchmarking Hutch is how ungainly it is to pull comparison-ready output from text buffers. I initially ran the evals by invoking multiple headless emacsen and tee-ing the agent output before it was rendered, but eventually settled on emitting [Perfetto](https://perfetto.dev) traces and using them as the underlying medium for evals.\n\nI haven't seen agents traced through Perfetto elsewhere. This is likely for good reason. They aren't meant for this kind of thing really. They don't have a first-class notion of what a \"prompt\" or a \"tool call\" is. It's designed for kernels and browsers and not an abstract system with tons of prose.\n\nBut for a single-player, emacs-local agent, it sort of works. You can answer all kinds of structural questions like *why did this review take 40 rounds?*, *what tools were run in parallel?*, or *how much wall time was spent in reading diffs?*, and so on. But more importantly, it's free and infra-free. If you set `hutch-trace-dir`, it will emit Perfetto traces and you can just load them up on [ui.perfetto.dev](https://ui.perfetto.dev). Easy.\n\nHere's an example to fetch tool calls and their total times. This is the entire pipeline. No dashboards or SDKs required:\n\n```\nSELECT\n  name                       AS tool,\n  COUNT(*)                   AS calls,\n  ROUND(SUM(dur) / 1e6, 1)   AS total_ms\nFROM slice\nWHERE category = 'tool'\nGROUP BY name\nORDER BY total_ms DESC;\n-- tool                 calls  total_ms\n--------------------------------------\n-- search_codebase      18     4210.3\n-- read_diff            7      2103.1\n-- surrounding_context  12     880.5\n-- read_file            3      412.7\n-- submit_review        1      42.9\n```\nWith this set up, we take a mix of strategies from [Martian’s code review](https://github.com/withmartian/code-review-benchmark) benchmark and the [CR-Bench preprint](https://arxiv.org/html/2603.11078v1) and compute Precision, Recall, and Fβ scores. The evals are described in more detail[<sup>3</sup>](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fn-3) in the [eval/README.org](https://github.com/adjaecent/magit-hutch/blob/main/eval/README.org) section. But broadly, we run the bench against 40 PRs, 132 goldens, and use GPT 5.2 as a classifying judge. The eval pipeline goes off and runs queries directly on the traces. Looking at the numbers, I believe we land somewhere around the #16 mark on Martian’s Offline Benchmark [leaderboard](https://codereview.withmartian.com/?mode=offline), which is pretty competitive for a no-memory, single-shot agent.\n\nOutside of classified scoring, there are a few interesting things about the agent itself:\n\nDifferent models tend to catch different bugs. Out of 132 goldens, each model hits 40-50 goldens, with an overlap of 18 hits across all three models. Which means hypothetically, if all three ran combined, it would catch ~55% more bugs than one model alone.\n\nPretty lousy agreement across the models on what a bug is, I'd say.\n\nGPT 5.5 tends to hit my default round limit (80) a lot more than the other models for roughly the same hit rate. Opus 4.8 takes 3x fewer turns to complete.\n\nOn token efficiency, Opus is much cheaper on output tokens used per good finding by a respectable margin, but burns 3x more context on inputs, possibly due to the growing context Hutch resends each round.\n\n## dead on arrival[#](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#dead_on_arrival)\n\nThis is all probably too late, as I've been told. No one really writes or reviews code, uses editors or version control by hand any longer. I made this for myself and for workflows that I still practice. I don't want to purport any arguments about whether one should or shouldn't use LLMs with emacs. The tool has more to do with unlocking a certain kind of workflow than the overreach of agents in niche locations.\n\nIf this continues to be useful, I'd like to add a conversational mode for every finding (like CodeRabbit) and perhaps maintain a context tree learnt from and committable to the codebase to improve review quality and speed.\n\n1. In the example, I use GLM-5.2 as the underlying model, but this is configurable to whatever backend the excellent [gptel](https://github.com/karthink/gptel) project supports.[↩](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fnref1)\n2. The characterization tests and evals are covered under the [evaluating](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#evaluating) section, but I haven't yet gotten a chance to verify this claim empirically.[↩](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fnref2)\n3. There are some biases and nuances to consider before treating the hard metrics as truly objective. But I've elided them from the post since they are described in more detail in the [README](https://github.com/adjaecent/magit-hutch/blob/main/eval/README.org) .[↩](https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fnref3)\n","body_html":"<h1 id=\"hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised\">hutch: local code reviews in emacs for the mildly disenfranchised</h1>\n<p>I haven&#39;t had a real job in four years. I closed down a <a href=\"https://tramline.app\" rel=\"nofollow ugc noopener\">startup</a> I&#39;d been building, just last month. During all these years, I spent most of that time at the back end of the frontier of AI agents. But I&#39;ve finally caught up. It&#39;s been some <a href=\"https://en.wikipedia.org/wiki/List_of_large_language_models#2025\" rel=\"nofollow ugc noopener\">500 days</a> since coding agents have really picked up, and they&#39;re genuinely more productive than, previously, <a href=\"https://www.youtube.com/watch?v=U_cSLPv34xk\" rel=\"nofollow ugc noopener\">instructed</a>.</p>\n<p>Even though I still prefer the pedagogical aspect of AI over the task-completing automaton aspects, the latter has driven all sorts of tooling around reviewing code, and not just writing and deploying it. The typical review agent party-line is: agents jump in, before your colleagues do, spray logorrhea across twenty pull requests before you have had a chance to wake up and look at your phone. This works, sometimes, for some people. But if you&#39;re like me, you still have humans reviewing code before it ships to users, and it&#39;s better to respect those people and their time. This is the case, regardless of where you sit on the balance of game-changer to curmudgeon.</p>\n<p>All that is to say, no matter which direction agents take to get better with time, I hope we still <em>care</em> about things. Not in the way of formalizing care, with high-fidelity agent instructions and prompts or some superior upholding of taste sort of thing, but something as simple as announcing: <em>hey I&#39;m still here, and I understand all this</em>.</p>\n<p>So as a long-time emacs user, I present yet another attempt at wedging LLMs, agents and coding harnesses, now inside your text buffers (!) with <a href=\"https://github.com/adjaecent/magit-hutch\" rel=\"nofollow ugc noopener\">Hutch</a>. It&#39;s a small, Magit-induced code-review interface that fits a standard Magit commit-push workflow <em>locally</em> and hopefully helps reclaim some load created upstream.</p>\n<h2 id=\"quick-tour\">quick tour<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#quick_tour\" rel=\"nofollow ugc noopener\">#</a></h2>\n<p>Open up Magit, and hit the dispatcher binding (usually <code>d</code>) and you&#39;ll see a <code>Hutch code review</code> action put up next to the DWIM binding. Hutch operates on three different scopes: staged changes, un-pushed changes, and changes between current branch and working branch. By default, it&#39;s staged changes only, since that&#39;s most useful.</p>\n<p>Once a review starts, you&#39;ll see a nice little progress bar in a new <code>*magit-hutch: code review*</code> buffer until the findings<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fn-1\" rel=\"nofollow ugc noopener\">&lt;sup&gt;1&lt;/sup&gt;</a> are complete. This is a read-only buffer, but you can still perform the <a href=\"https://github.com/adjaecent/magit-hutch#usage\" rel=\"nofollow ugc noopener\">pre-bound actions</a>.</p>\n<p>Each finding has a type (suggestion, comment or LGTM), a file name, relevant line numbers, a title and a description. Suggestions additionally have a patch diff. Hutch piggybacks on <a href=\"https://magit.vc\" rel=\"nofollow ugc noopener\">Magit</a> and <a href=\"https://github.com/magit/transient\" rel=\"nofollow ugc noopener\">Transient</a> to render these menus so it behaves much like its own interface; keyboard-driven sub-menus, diff coloring, and highlighting. Suggestions are special since they can be applied. To mark a suggestion for application, you queue it with <code>m</code>.</p>\n<p>Then bulk-apply all queued suggestions with <code>A</code>. The application is scope-aware, so if you queue a finding for staged changes, it will apply the fix directly to the staged files.</p>\n<p>That&#39;s it! Getting started should hopefully be pretty simple and intuitive for existing emacs users. There are of course a few interesting things going on behind the scenes, some of which I&#39;ll cover in the next few sections.</p>\n<h2 id=\"patches-over-comments\">patches over comments<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#patches_over_comments\" rel=\"nofollow ugc noopener\">#</a></h2>\n<p>A big UX handicap of showing review comments and suggested patches in-buffer is that there is no existing connective tissue of a commenting system. With GitHub, though, the review UI collapses outdated comments on new commits and most review bots sit over the <a href=\"https://docs.github.com/en/pull-requests/how-tos/review-pull-requests/incorporating-feedback-in-your-pull-request#applying-suggested-changes\" rel=\"nofollow ugc noopener\">suggestion</a> mechanic if they have changes to suggest.</p>\n<p>Hutch is made with a bias towards patches, rather than just prosaic comments. According to the <a href=\"https://aider.chat/docs/leaderboards\" rel=\"nofollow ugc noopener\">Aider leaderboard</a> (and through some of my own experiments), the <code>SEARCH/REPLACE</code> diffs are a lot more obedient across different models than just asking the model to author correct patches with precise line numbers.</p>\n<p>For an <a href=\"https://aider.chat/docs/more/edit-formats.html\" rel=\"nofollow ugc noopener\">Aider-style diff</a>, you have to ensure there&#39;s enough surrounding context for the <code>SEARCH</code> to be unique, and ideally also preserve indentation. In Hutch&#39;s case, the tool&#39;s function schema naturally decomposes the <em>file, search, and replace</em> fields:</p>\n<pre><code>src/utils.clj\n&lt;&lt;&lt;&lt;&lt;&lt;&lt; SEARCH\n(defn add [a b]\n  (+ a b))\n=======\n(defn add [a b c]\n  (+ a b c))\n&gt;&gt;&gt;&gt;&gt;&gt;&gt; REPLACE</code></pre>\n<p>I&#39;ve noticed that a lot of older (or cheaper) models tend to recall the <code>SEARCH</code> block from memory when asked for diffs, instead of copying it verbatim from <code>read_file</code>, <code>read_diff</code> or <code>surrounding_context</code> calls. This invariably botches them entirely. So we get them verified before submission. If <code>SEARCH</code> is missing or matches more than once, the finding is downgraded to a plain comment. On a unique hit, Hutch locally creates a unified diff:</p>\n<p>Once a series of udiffs and comments are rendered, they can be marked and bulk applied. Hutch applies them per-file, lowest hunk first (bottom-up) so line positions are minimally disturbed. Each finding runs its own <code>git apply</code> and a bad application marks itself <code>invalid</code> so the rest can continue to land.</p>\n<p>All this patching and commenting infrastructure pulls its weight, since with only a couple of keystrokes, you hopefully get less reading and parsing work and more actionable triaging. None of this guarantees patches-always of course, and it shouldn&#39;t.</p>\n<p>With more powerful models, a simpler diffing method might generally work pretty well. But for a tool that&#39;s built to work across different and cheaper models, it&#39;s essential to be maximally supportive. In general, I feel like a key point of much of the agentic infrastructure we build is to have knobs for optimizing token:cost ratios. This could often mean thorny workarounds for good-enough models.</p>\n<h2 id=\"barely-enough-tooling\">barely enough tooling<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#barely_enough_tooling\" rel=\"nofollow ugc noopener\">#</a></h2>\n<p>Hutch has a fairly minimal toolset for pulling context:</p>\n<ol><li><code>read_diff</code></li><li><code>read_file</code></li><li><code>search_codebase</code></li><li><code>surrounding_context</code></li></ol>\n<p>Out of these, <code>surrounding_context</code> is the more interesting one. It wraps over <a href=\"https://batsov.com/articles/2026/02/27/building-emacs-major-modes-with-treesitter-lessons-learned/#why-tree-sitter\" rel=\"nofollow ugc noopener\">Tree-sitter</a> and uses grammars that are installed. It works by letting the model widen out to the enclosing definition of a relevant line and further out, as needed. In my tests, the overall read token consumption compared to simply blasting <code>read_file</code> was anecdotally lower with comparable levels of review quality<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fn-2\" rel=\"nofollow ugc noopener\">&lt;sup&gt;2&lt;/sup&gt;</a>.</p>\n<p>All the findings from the model are submitted to the agent at once. On the write side of things, <code>verify_block</code> locally verifies diffs, and along with other comments and LGTM notices, submits them through a <code>submit_review</code> tool call. <code>submit_review</code> itself runs through some post-processing work, like gating hallucinations about files and line numbers, trimming the length of descriptions and downgrading patches to comments if they don&#39;t apply cleanly.</p>\n<p>Once the submission lands, the output from all this work is persisted durably under <code>refs/hutch/id</code> and can be separately committed as a means of sharing (with <code>magit-post-commit-hook</code>) or for repainting later. If you squint hard enough, it might appear like a change identifier for a stacked-diff <a href=\"https://blog.tangled.org/stacking\" rel=\"nofollow ugc noopener\">review tool</a>, but its purpose is to keep reviews in the git tree, rather than identify changesets for human reviews. We don&#39;t really care about multi-party human reviews, it&#39;s all local.</p>\n<h2 id=\"evaluating\">evaluating<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#evaluating\" rel=\"nofollow ugc noopener\">#</a></h2>\n<p>The one unfortunate part about benchmarking Hutch is how ungainly it is to pull comparison-ready output from text buffers. I initially ran the evals by invoking multiple headless emacsen and tee-ing the agent output before it was rendered, but eventually settled on emitting <a href=\"https://perfetto.dev\" rel=\"nofollow ugc noopener\">Perfetto</a> traces and using them as the underlying medium for evals.</p>\n<p>I haven&#39;t seen agents traced through Perfetto elsewhere. This is likely for good reason. They aren&#39;t meant for this kind of thing really. They don&#39;t have a first-class notion of what a &quot;prompt&quot; or a &quot;tool call&quot; is. It&#39;s designed for kernels and browsers and not an abstract system with tons of prose.</p>\n<p>But for a single-player, emacs-local agent, it sort of works. You can answer all kinds of structural questions like <em>why did this review take 40 rounds?</em>, <em>what tools were run in parallel?</em>, or <em>how much wall time was spent in reading diffs?</em>, and so on. But more importantly, it&#39;s free and infra-free. If you set <code>hutch-trace-dir</code>, it will emit Perfetto traces and you can just load them up on <a href=\"https://ui.perfetto.dev\" rel=\"nofollow ugc noopener\">ui.perfetto.dev</a>. Easy.</p>\n<p>Here&#39;s an example to fetch tool calls and their total times. This is the entire pipeline. No dashboards or SDKs required:</p>\n<pre><code>SELECT\n  name                       AS tool,\n  COUNT(*)                   AS calls,\n  ROUND(SUM(dur) / 1e6, 1)   AS total_ms\nFROM slice\nWHERE category = &#39;tool&#39;\nGROUP BY name\nORDER BY total_ms DESC;\n-- tool                 calls  total_ms\n--------------------------------------\n-- search_codebase      18     4210.3\n-- read_diff            7      2103.1\n-- surrounding_context  12     880.5\n-- read_file            3      412.7\n-- submit_review        1      42.9</code></pre>\n<p>With this set up, we take a mix of strategies from <a href=\"https://github.com/withmartian/code-review-benchmark\" rel=\"nofollow ugc noopener\">Martian’s code review</a> benchmark and the <a href=\"https://arxiv.org/html/2603.11078v1\" rel=\"nofollow ugc noopener\">CR-Bench preprint</a> and compute Precision, Recall, and Fβ scores. The evals are described in more detail<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fn-3\" rel=\"nofollow ugc noopener\">&lt;sup&gt;3&lt;/sup&gt;</a> in the <a href=\"https://github.com/adjaecent/magit-hutch/blob/main/eval/README.org\" rel=\"nofollow ugc noopener\">eval/README.org</a> section. But broadly, we run the bench against 40 PRs, 132 goldens, and use GPT 5.2 as a classifying judge. The eval pipeline goes off and runs queries directly on the traces. Looking at the numbers, I believe we land somewhere around the #16 mark on Martian’s Offline Benchmark <a href=\"https://codereview.withmartian.com/?mode=offline\" rel=\"nofollow ugc noopener\">leaderboard</a>, which is pretty competitive for a no-memory, single-shot agent.</p>\n<p>Outside of classified scoring, there are a few interesting things about the agent itself:</p>\n<p>Different models tend to catch different bugs. Out of 132 goldens, each model hits 40-50 goldens, with an overlap of 18 hits across all three models. Which means hypothetically, if all three ran combined, it would catch ~55% more bugs than one model alone.</p>\n<p>Pretty lousy agreement across the models on what a bug is, I&#39;d say.</p>\n<p>GPT 5.5 tends to hit my default round limit (80) a lot more than the other models for roughly the same hit rate. Opus 4.8 takes 3x fewer turns to complete.</p>\n<p>On token efficiency, Opus is much cheaper on output tokens used per good finding by a respectable margin, but burns 3x more context on inputs, possibly due to the growing context Hutch resends each round.</p>\n<h2 id=\"dead-on-arrival\">dead on arrival<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#dead_on_arrival\" rel=\"nofollow ugc noopener\">#</a></h2>\n<p>This is all probably too late, as I&#39;ve been told. No one really writes or reviews code, uses editors or version control by hand any longer. I made this for myself and for workflows that I still practice. I don&#39;t want to purport any arguments about whether one should or shouldn&#39;t use LLMs with emacs. The tool has more to do with unlocking a certain kind of workflow than the overreach of agents in niche locations.</p>\n<p>If this continues to be useful, I&#39;d like to add a conversational mode for every finding (like CodeRabbit) and perhaps maintain a context tree learnt from and committable to the codebase to improve review quality and speed.</p>\n<ol><li>In the example, I use GLM-5.2 as the underlying model, but this is configurable to whatever backend the excellent <a href=\"https://github.com/karthink/gptel\" rel=\"nofollow ugc noopener\">gptel</a> project supports.<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fnref1\" rel=\"nofollow ugc noopener\">↩</a></li><li>The characterization tests and evals are covered under the <a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#evaluating\" rel=\"nofollow ugc noopener\">evaluating</a> section, but I haven&#39;t yet gotten a chance to verify this claim empirically.<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fnref2\" rel=\"nofollow ugc noopener\">↩</a></li><li>There are some biases and nuances to consider before treating the hard metrics as truly objective. But I&#39;ve elided them from the post since they are described in more detail in the <a href=\"https://github.com/adjaecent/magit-hutch/blob/main/eval/README.org\" rel=\"nofollow ugc noopener\">README</a> .<a href=\"https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/#fnref3\" rel=\"nofollow ugc noopener\">↩</a></li></ol>","headings":[{"level":1,"text":"hutch: local code reviews in emacs for the mildly disenfranchised","id":"hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised"},{"level":2,"text":"quick tour#","id":"quick-tour"},{"level":2,"text":"patches over comments#","id":"patches-over-comments"},{"level":2,"text":"barely enough tooling#","id":"barely-enough-tooling"},{"level":2,"text":"evaluating#","id":"evaluating"},{"level":2,"text":"dead on arrival#","id":"dead-on-arrival"}]}}