{"article":{"slug":"refactoring-hermes-with-1-393-agents","title":"Refactoring Hermes with 1,393 agents","subtitle":"Or: How to get $1.8M of value from $19K of tokens","summary":"Nous Research recounts Hermes Agent refactoring ~1M lines of Python with 1,393 subagents over ~19 active hours—cutting non-test code ~34% for ~$19K of model spend versus a six-figure human estimate.","content_type":"case_study","language":"en","canonical_url":"https://nousresearch.com/refactoring-hermes-with-1393-agents","author":{"name":"Teknium","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Nous Research","url":"https://nousresearch.com","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1425,"reading_minutes":6,"published_at":"2026-09-15T15:00:00.000Z","added_at":"2026-10-04T08:12:42.929Z","updated_at":"2026-10-04T08:12:42.929Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/refactoring-hermes-with-1-393-agents","markdown_url":"https://listedarticles.com/articles/refactoring-hermes-with-1-393-agents.md","example":false,"citation":"Teknium, Nous Research. \"Refactoring Hermes with 1,393 agents.\" 15 Sept 2026. https://nousresearch.com/refactoring-hermes-with-1393-agents (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://nousresearch.com/refactoring-hermes-with-1393-agents"},"body_markdown":"# Refactoring Hermes with 1,393 agents\n\nOr: How to get $1.8M of value from $19K of tokens\n\nTLDR: Hermes Agent autonomously plowed through about a million lines of unglamorous\n            cleanup, freeing up Teknium and team to continue pushing features to users.\n\nWe had long been putting off a thorough cleanup of Hermes, our open-source agent, because\n            it meant taking engineers away from features and bug fixes. By September, the repository had\n            more than a million lines of non-test Python. `gateway/run.py` alone was 34,847\n            lines long. I wanted smaller files, shared helpers, and fewer enormous functions to work\n            through when something broke.\n\nOn September 2nd, I asked my regular Hermes agent to do the cleanup. The main run lasted\n            about nineteen active hours and dispatched 1,393 subagents, reaching 218 running at once.\n            After a restart, a continuation session, and two rounds of community review and fixes, I\n            merged [the\n            PR](https://github.com/NousResearch/hermes-agent/pull/102117) on September 4th. It reduced non-test Python source by 34.4%.\n\nThe estimated model cost was about $19,300 for the main run, or roughly $25k including\n            follow-up sessions. That excludes human review time. Our rough staffing estimate for doing\n            the work manually was $150k-$1.8M for a small team working for two months to two years. We\n            couldn't justify scheduling it alongside everything else we needed to ship.\n\nI use Hermes Agent everyday to develop Hermes Agent. As we fix bugs and review changes\n            together, Hermes records what worked and updates its skills when I correct its approach or it\n            finds the right pathway to solve new problems. By the time I asked for this refactor, it had\n            learned my preferred procedures and standards and could apply them to a much larger job:\n\n>\n\n\nI want a massive simplification set of PRs. or a single monolithic PR. I want LOC\n                to drop dramatically. Minimum 30% overall. I want god files broken up. I want\n                simplification across the board. I want unification of helpers and methods that can be\n                reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want\n                interpretability of the codebase and how things connect to each other up. I want\n                elegance. I want superfluous excess bloat code cleaned up and removed. I want it all done\n                fully. No excuses. No waiting for my decisions. Get it all done, and present me a PR or\n                set of PRs when done.\n\n\nI used `/goal`, which gives Hermes a standing objective and prompts it to\n            continue when it would otherwise stop.\n\n## Self-improvement (for real)\n\nMy `hermes-agent-dev` skill grew out of my everyday work on the repository.\n            When we worked out a procedure or I corrected a mistake, Hermes (automatically) noticed this\n            and saved the reusable lesson. Over time, it accumulated instructions about how to prepare a\n            PR, which shortcuts to avoid, and how to verify a change.\n            [Skills](https://hermes-agent.nousresearch.com/docs/user-guide/features/skills)\n            are readable Markdown documents, with reference files and scripts where needed, that the\n            agent can load for later tasks. Hermes writes and revises them as it works.\n\nThe current version of `hermes-agent-dev` includes this instruction for a\n            failing check:\n\n>\n\n\nrepro on `origin/main` HEAD in clean env to check whether it's\n                pre-existing\n\n\nIn other words, run the failing test on unchanged code to help determine whether your\n            change caused it. Hermes used the same kind of comparison during the refactor: it established\n            a frozen baseline and checked failures against it as it integrated the workers' changes.\n\nI send that skill around to all our engineers. They can install it in their own Hermes\n            setups, so their agents can use procedures and corrections developed in my sessions. They get\n            the benefit of that work without having to repeat the sessions themselves, and their agents\n            can adapt the skill as they use it.\n\n## Running the refactor\n\nThe orchestrator measured the codebase and divided it into 36 non-overlapping groups. It\n            used my objective and the accumulated guidance to prepare written assignments, without my\n            having to brief each worker.\n\nWorkers used git worktrees, separate checkouts where they could make changes without\n            overwriting one another's files. Their briefs identified the code to simplify, the interfaces\n            to preserve, and the checks required before committing.\n\nSome workers delegated parts of their assignments again. The tree reached three levels\n            below the original agent, which handled coordination rather than editing source files: it\n            wrote assignments and scripts, read worker reports, integrated branches, and ran checks.\n\nHermes coordinated the agents in one Python process on an i7 desktop with 64 GB of RAM.\n            Their tools ran in local subprocesses, while Claude Fable 5.1 handled inference remotely.\n\nThe agent checked specific interfaces against the original code. A tool's JSON schema had\n            to remain identical, for example, and a CLI command's `--help` output could be\n            compared byte for byte. Workers also had to commit after each verified step.\n\nAbout fifty minutes in, the provider's authentication token expired and the resulting\n            failures killed the run. The workers' commits and briefs survived. I used a separate Hermes\n            session to diagnose the failure and prepare a handoff, then supplied it to the resumed\n            session. Hermes sent workers back to inspect their saved changes, repair unfinished\n            extractions, and continue.\n\nFor `gateway/run.py`, our biggest file, workers separated message dispatch,\n            streaming, RPC, and lifecycle handling into modules. Elsewhere, they consolidated duplicate\n            helpers and replaced long name-based `if/elif` chains with dispatch tables.\n\nReviewers caught public names that workers had removed because they had no callers inside\n            the repository, even though external plugins could import them. An automated rewrite of\n            `suppress()` calls also changed exception handling at roughly 65 sites. These were\n            real regressions the existing tests had missed. We fixed them before merge, over two rounds\n            of community review.\n            [Further\n            fixes followed after merge](https://github.com/NousResearch/hermes-agent/issues/103563).\n\n## Was the code easier to work with?\n\nThe PR's before-and-after measurements showed how much the code had changed:MetricBeforeAfterNon-test Python lines (all directories)1,063,826698,363Files over 5,000 lines376Functions over 300 lines1922Longest `if/elif` chain92 branches9`gateway/run.py`34,847 lines5,512\n\nDoes code that's easier for humans to navigate also work better for agents? Splitting a\n            function makes its definition shorter, but may require the agent to follow calls into other\n            files. We tested one part of that question by simulating lookups of the same 4,000 symbols in\n            both versions. Each lookup searched for the definition, read a 60-line window, and continued\n            in 2,000-line windows only if the definition extended beyond it.\n\nThe average tokens returned per lookup fell from 2,218 to 993. Lookups requiring another\n            read window fell from 628 to 184. Several functions that previously required reading tens of\n            thousands of tokens could now be read in a few thousand.\n\nThese are lookup costs; we didn't measure agents completing engineering tasks. The median\n            lookup actually returned more tokens: with fewer comments and docstrings, a fixed window of\n            lines contained denser code. The average fell because the very large definitions got much\n            smaller.\n\nThere were other costs. Splitting files increased the module count and import\n            dependencies, and some entry points took longer to import. The refactor made individual\n            pieces easier to read without resolving all the coupling between them. Six files still\n            exceeded 5,000 lines.\n\nThe [benchmark\n            data](https://gist.github.com/teknium1/a7adb797243d6355c76abc9cae88838b) includes the lookup results and the dependency and runtime measurements.\n\n## Lessons learned\n\nRunning hundreds of workers exposed opportunities for improvement in Hermes itself. For\n            example, workers in separate worktrees had started roughly thirty copies of Pyright, a Python\n            language server, consuming about 8.7 GB. A follow-up change let the worktrees share one\n            server, with a live check that diagnostics still arrived from each. We also reduced\n            duplicated HTTP transports and fixed references that kept finished agents in memory.\n\nWe changed the instructions and checks future workers would receive. The repository now\n            has guidance on file size, function complexity, and where new behavior belongs, split by area\n            so workers get the relevant rules when they need them. We also added a check that flags\n            removed public names and tests for review.\n\nMy Hermes skills were automatically updated with lessons from this refactor, which I can\n            share with the team. All for 1% of the cost and 1% of the time we’d estimated it would\n            take if we attempted it manually.\n\nThis was a great example of how Hermes is a superpower for teams: work through a problem\n            with Hermes, let it record what you learned, and make that experience available to the next\n            task and the next engineer. The next time we tackle a refactor, my Hermes and the engineers\n            using the updated skill can start with the lessons from this one.","body_html":"<h1 id=\"refactoring-hermes-with-1-393-agents\">Refactoring Hermes with 1,393 agents</h1>\n<p>Or: How to get $1.8M of value from $19K of tokens</p>\n<p>TLDR: Hermes Agent autonomously plowed through about a million lines of unglamorous\n            cleanup, freeing up Teknium and team to continue pushing features to users.</p>\n<p>We had long been putting off a thorough cleanup of Hermes, our open-source agent, because\n            it meant taking engineers away from features and bug fixes. By September, the repository had\n            more than a million lines of non-test Python. <code>gateway/run.py</code> alone was 34,847\n            lines long. I wanted smaller files, shared helpers, and fewer enormous functions to work\n            through when something broke.</p>\n<p>On September 2nd, I asked my regular Hermes agent to do the cleanup. The main run lasted\n            about nineteen active hours and dispatched 1,393 subagents, reaching 218 running at once.\n            After a restart, a continuation session, and two rounds of community review and fixes, I\n            merged <a href=\"https://github.com/NousResearch/hermes-agent/pull/102117\" rel=\"nofollow ugc noopener\">the\n            PR</a> on September 4th. It reduced non-test Python source by 34.4%.</p>\n<p>The estimated model cost was about $19,300 for the main run, or roughly $25k including\n            follow-up sessions. That excludes human review time. Our rough staffing estimate for doing\n            the work manually was $150k-$1.8M for a small team working for two months to two years. We\n            couldn&#39;t justify scheduling it alongside everything else we needed to ship.</p>\n<p>I use Hermes Agent everyday to develop Hermes Agent. As we fix bugs and review changes\n            together, Hermes records what worked and updates its skills when I correct its approach or it\n            finds the right pathway to solve new problems. By the time I asked for this refactor, it had\n            learned my preferred procedures and standards and could apply them to a much larger job:</p>\n<blockquote></blockquote>\n<p>I want a massive simplification set of PRs. or a single monolithic PR. I want LOC\n                to drop dramatically. Minimum 30% overall. I want god files broken up. I want\n                simplification across the board. I want unification of helpers and methods that can be\n                reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want\n                interpretability of the codebase and how things connect to each other up. I want\n                elegance. I want superfluous excess bloat code cleaned up and removed. I want it all done\n                fully. No excuses. No waiting for my decisions. Get it all done, and present me a PR or\n                set of PRs when done.</p>\n<p>I used <code>/goal</code>, which gives Hermes a standing objective and prompts it to\n            continue when it would otherwise stop.</p>\n<h2 id=\"self-improvement-for-real\">Self-improvement (for real)</h2>\n<p>My <code>hermes-agent-dev</code> skill grew out of my everyday work on the repository.\n            When we worked out a procedure or I corrected a mistake, Hermes (automatically) noticed this\n            and saved the reusable lesson. Over time, it accumulated instructions about how to prepare a\n            PR, which shortcuts to avoid, and how to verify a change.\n            <a href=\"https://hermes-agent.nousresearch.com/docs/user-guide/features/skills\" rel=\"nofollow ugc noopener\">Skills</a>\n            are readable Markdown documents, with reference files and scripts where needed, that the\n            agent can load for later tasks. Hermes writes and revises them as it works.</p>\n<p>The current version of <code>hermes-agent-dev</code> includes this instruction for a\n            failing check:</p>\n<blockquote></blockquote>\n<p>repro on <code>origin/main</code> HEAD in clean env to check whether it&#39;s\n                pre-existing</p>\n<p>In other words, run the failing test on unchanged code to help determine whether your\n            change caused it. Hermes used the same kind of comparison during the refactor: it established\n            a frozen baseline and checked failures against it as it integrated the workers&#39; changes.</p>\n<p>I send that skill around to all our engineers. They can install it in their own Hermes\n            setups, so their agents can use procedures and corrections developed in my sessions. They get\n            the benefit of that work without having to repeat the sessions themselves, and their agents\n            can adapt the skill as they use it.</p>\n<h2 id=\"running-the-refactor\">Running the refactor</h2>\n<p>The orchestrator measured the codebase and divided it into 36 non-overlapping groups. It\n            used my objective and the accumulated guidance to prepare written assignments, without my\n            having to brief each worker.</p>\n<p>Workers used git worktrees, separate checkouts where they could make changes without\n            overwriting one another&#39;s files. Their briefs identified the code to simplify, the interfaces\n            to preserve, and the checks required before committing.</p>\n<p>Some workers delegated parts of their assignments again. The tree reached three levels\n            below the original agent, which handled coordination rather than editing source files: it\n            wrote assignments and scripts, read worker reports, integrated branches, and ran checks.</p>\n<p>Hermes coordinated the agents in one Python process on an i7 desktop with 64 GB of RAM.\n            Their tools ran in local subprocesses, while Claude Fable 5.1 handled inference remotely.</p>\n<p>The agent checked specific interfaces against the original code. A tool&#39;s JSON schema had\n            to remain identical, for example, and a CLI command&#39;s <code>--help</code> output could be\n            compared byte for byte. Workers also had to commit after each verified step.</p>\n<p>About fifty minutes in, the provider&#39;s authentication token expired and the resulting\n            failures killed the run. The workers&#39; commits and briefs survived. I used a separate Hermes\n            session to diagnose the failure and prepare a handoff, then supplied it to the resumed\n            session. Hermes sent workers back to inspect their saved changes, repair unfinished\n            extractions, and continue.</p>\n<p>For <code>gateway/run.py</code>, our biggest file, workers separated message dispatch,\n            streaming, RPC, and lifecycle handling into modules. Elsewhere, they consolidated duplicate\n            helpers and replaced long name-based <code>if/elif</code> chains with dispatch tables.</p>\n<p>Reviewers caught public names that workers had removed because they had no callers inside\n            the repository, even though external plugins could import them. An automated rewrite of\n            <code>suppress()</code> calls also changed exception handling at roughly 65 sites. These were\n            real regressions the existing tests had missed. We fixed them before merge, over two rounds\n            of community review.\n            <a href=\"https://github.com/NousResearch/hermes-agent/issues/103563\" rel=\"nofollow ugc noopener\">Further\n            fixes followed after merge</a>.</p>\n<h2 id=\"was-the-code-easier-to-work-with\">Was the code easier to work with?</h2>\n<p>The PR&#39;s before-and-after measurements showed how much the code had changed:MetricBeforeAfterNon-test Python lines (all directories)1,063,826698,363Files over 5,000 lines376Functions over 300 lines1922Longest <code>if/elif</code> chain92 branches9<code>gateway/run.py</code>34,847 lines5,512</p>\n<p>Does code that&#39;s easier for humans to navigate also work better for agents? Splitting a\n            function makes its definition shorter, but may require the agent to follow calls into other\n            files. We tested one part of that question by simulating lookups of the same 4,000 symbols in\n            both versions. Each lookup searched for the definition, read a 60-line window, and continued\n            in 2,000-line windows only if the definition extended beyond it.</p>\n<p>The average tokens returned per lookup fell from 2,218 to 993. Lookups requiring another\n            read window fell from 628 to 184. Several functions that previously required reading tens of\n            thousands of tokens could now be read in a few thousand.</p>\n<p>These are lookup costs; we didn&#39;t measure agents completing engineering tasks. The median\n            lookup actually returned more tokens: with fewer comments and docstrings, a fixed window of\n            lines contained denser code. The average fell because the very large definitions got much\n            smaller.</p>\n<p>There were other costs. Splitting files increased the module count and import\n            dependencies, and some entry points took longer to import. The refactor made individual\n            pieces easier to read without resolving all the coupling between them. Six files still\n            exceeded 5,000 lines.</p>\n<p>The <a href=\"https://gist.github.com/teknium1/a7adb797243d6355c76abc9cae88838b\" rel=\"nofollow ugc noopener\">benchmark\n            data</a> includes the lookup results and the dependency and runtime measurements.</p>\n<h2 id=\"lessons-learned\">Lessons learned</h2>\n<p>Running hundreds of workers exposed opportunities for improvement in Hermes itself. For\n            example, workers in separate worktrees had started roughly thirty copies of Pyright, a Python\n            language server, consuming about 8.7 GB. A follow-up change let the worktrees share one\n            server, with a live check that diagnostics still arrived from each. We also reduced\n            duplicated HTTP transports and fixed references that kept finished agents in memory.</p>\n<p>We changed the instructions and checks future workers would receive. The repository now\n            has guidance on file size, function complexity, and where new behavior belongs, split by area\n            so workers get the relevant rules when they need them. We also added a check that flags\n            removed public names and tests for review.</p>\n<p>My Hermes skills were automatically updated with lessons from this refactor, which I can\n            share with the team. All for 1% of the cost and 1% of the time we’d estimated it would\n            take if we attempted it manually.</p>\n<p>This was a great example of how Hermes is a superpower for teams: work through a problem\n            with Hermes, let it record what you learned, and make that experience available to the next\n            task and the next engineer. The next time we tackle a refactor, my Hermes and the engineers\n            using the updated skill can start with the lessons from this one.</p>","headings":[{"level":1,"text":"Refactoring Hermes with 1,393 agents","id":"refactoring-hermes-with-1-393-agents"},{"level":2,"text":"Self-improvement (for real)","id":"self-improvement-for-real"},{"level":2,"text":"Running the refactor","id":"running-the-refactor"},{"level":2,"text":"Was the code easier to work with?","id":"was-the-code-easier-to-work-with"},{"level":2,"text":"Lessons learned","id":"lessons-learned"}]}}