{"article":{"slug":"a-deep-dive-into-ponytail","title":"A deep dive into Ponytail","subtitle":null,"summary":"Flavio Copes explains Ponytail, a plugin that stops coding agents from over-engineering: how it works, how to install it in ChatGPT, Codex, and Claude Code, intensity settings, and a practical workflow for keeping agent-built features small.","content_type":"tutorial","language":"en","canonical_url":"https://flaviocopes.com/ponytail/","author":{"name":"Flavio Copes","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"flaviocopes.com","url":"https://flaviocopes.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Productivity","slug":"productivity","url":"https://listedarticles.com/topics/productivity"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2212,"reading_minutes":10,"published_at":"2026-09-23T15:56:01.772Z","added_at":"2026-09-23T15:56:01.772Z","updated_at":"2026-09-23T15:56:01.772Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/a-deep-dive-into-ponytail","markdown_url":"https://listedarticles.com/articles/a-deep-dive-into-ponytail.md","example":false,"citation":"Flavio Copes, flaviocopes.com. \"A deep dive into Ponytail.\" 23 Sept 2026. https://flaviocopes.com/ponytail/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://flaviocopes.com/ponytail/"},"body_markdown":"[Home](</>) / [AI](</tags/ai/>)\n\n# A deep dive into Ponytail\n\nBy [Flavio Copes](</about/>)\n\nUpdated Sep 17, 2026\n\nLearn how Ponytail stops AI coding agents from over-engineering: install it in ChatGPT, Codex, or Claude Code, choose a mode, and review diffs.\n\n~~~\n\nAI coding agents like writing code.\n\nSometimes they like it too much.\n\nYou ask for a date field. The agent installs a date picker, adds a wrapper component, creates a theme file, and starts discussing time zones.\n\nThe browser already has this:\n    \n    \n    <input type=\"date\" />\n\n[Ponytail](<https://ponytail.dev/>) is a ruleset that pushes coding agents toward that kind of answer.\n\nIt was created by Dietrich Gebert. The project describes its approach as the lazy senior developer: understand the problem, find the smallest correct solution, and stop.\n\nLazy does not mean careless here. Ponytail tells the agent to preserve validation, error handling, security, accessibility, and anything you explicitly requested. It aims to reduce the code you have to own, not play code golf.\n\n## How Ponytail works\n\nPonytail gives the agent a seven-step ladder.\n\nThe agent stops at the first step that solves the real problem:\n\n  1. Does this need to exist?\n  2. Does the codebase already have it?\n  3. Does the standard library provide it?\n  4. Does the native platform provide it?\n  5. Does an installed dependency already solve it?\n  6. Can it be one line?\n  7. Only then, write the minimum new code.\n\nThe order makes the agent reuse an existing date-field component before reaching for the native HTML input.\n\nSuppose the field needs a calendar system the browser input cannot represent. The native feature no longer solves the task, so the agent keeps moving down the ladder.\n\nPonytail also tells the agent to inspect the relevant code first. It should trace the real flow, search for existing helpers, and fix a bug at its shared cause instead of patching one visible symptom.\n\nIt is lazy about the solution, not the investigation.\n\n## Install Ponytail in the ChatGPT app\n\nOpen **Settings** , then select **Plugins**.\n\nOpen the **Add** menu and choose **Add plugin marketplace**. Enter `DietrichGebert/ponytail` as the source, then click **Add marketplace** :\n\n![Adding the Ponytail plugin marketplace in the ChatGPT app](/images/ponytail/chatgpt-add-marketplace.webp)\n\nReturn to the Plugins page and select **Personal**. Find Ponytail and click **Install** :\n\n![Installing Ponytail from the personal plugin marketplace in the ChatGPT app](/images/ponytail/chatgpt-install-plugin.webp)\n\nOpen a new Codex task and type:\n    \n    \n    /ponytail\n\nThe composer shows the installed Ponytail skills. Select the main Ponytail skill and send the prompt:\n\n![Activating the Ponytail skill from a Codex task in the ChatGPT app](/images/ponytail/chatgpt-use-ponytail.webp)\n\nPonytail confirms the active intensity. The default is `full`.\n\n## Install Ponytail in Codex CLI\n\nPonytail is available as a Codex plugin.\n\nAdd its marketplace:\n    \n    \n    codex plugin marketplace add DietrichGebert/ponytail\n\nThen install the plugin:\n    \n    \n    codex plugin add ponytail@ponytail\n\nStart Codex and open the hooks screen:\n    \n    \n    /hooks\n\nReview and trust Ponytail’s lifecycle hooks, then start a new task.\n\nThe hooks activate Ponytail when a session starts and pass the active mode to subagents.\n\nIf you install from the terminal and use the ChatGPT app too, restart the app. It will load the same plugin.\n\nPonytail starts in `full` mode by default. You can check that the bundled skills are available with:\n    \n    \n    @ponytail-help\n\nCodex CLI invokes skills with `@`. The ChatGPT app and Claude Code use slash commands.\n\n## Install Ponytail in Claude Code\n\nIn Claude Code, send these as two separate prompts.\n\nFirst, add the marketplace:\n    \n    \n    /plugin marketplace add DietrichGebert/ponytail\n\nWait for that command to finish. Then install Ponytail:\n    \n    \n    /plugin install ponytail@ponytail\n\nStart a new session after installation. Ponytail activates in `full` mode by default.\n\nYou can check the current mode with:\n    \n    \n    /ponytail\n\nThe same two installation commands work in the Code tab of the Claude Code desktop app.\n\nComing soon · waiting lists open\n\nAI Workshop and Solo Lab: new cohorts\n\nSolo Lab starts on 26 October: three weeks on finding customers and making a living from your own software. AI Workshop starts on 30 November: three weeks on directing coding agents and reviewing what they build. Join the waiting lists and I'll email you when enrollment opens.\n\n[See the cohorts →](</cohorts/>)\n\n## Install Ponytail in other CLIs\n\nGitHub Copilot CLI, Gemini CLI and Pi have their own install commands. Copilot CLI uses the same marketplace flow as Codex:\n    \n    \n    copilot plugin marketplace add DietrichGebert/ponytail\n    copilot plugin install ponytail@ponytail\n\nGemini CLI installs it as an extension:\n    \n    \n    gemini extensions install https://github.com/DietrichGebert/ponytail\n\nAnd Pi:\n    \n    \n    pi install git:github.com/DietrichGebert/ponytail\n\nThe Claude Code and Codex plugins use small Node.js lifecycle hooks. If `node` is missing from the non-interactive shell’s `PATH`, the skills still work, but automatic activation does not.\n\nThe command examples below use Codex CLI’s `@` syntax. In the ChatGPT app or Claude Code, replace `@` with `/`.\n\n## Try it on a small feature\n\nLet’s give the agent a task that often attracts too much code:\n    \n    \n    Add a birthday field to the profile form.\n    Use the existing form styles and keep the current validation behavior.\n\nIn `full` mode, Ponytail first inspects the form and its existing components.\n\nIf plain HTML covers the requirement, the result might be close to this:\n    \n    \n    <label for=\"birthday\">Birthday</label>\n    <input id=\"birthday\" name=\"birthday\" type=\"date\" />\n\nThe result needs no new package, calendar wrapper, or custom date parser.\n\nNotice that the prompt did not say “write the fewest lines possible.” Ponytail applies the ladder as a standing coding rule.\n\nIf the application already uses a form component with error messages and accessible descriptions, the agent should use that component instead. Reusing the codebase is higher on the ladder than using the browser directly.\n\n## Choose an intensity\n\nPonytail has three active modes:\n\n  * `lite` builds what you requested and mentions a smaller alternative\n  * `full` enforces the complete ladder and is the default\n  * `ultra` treats speculative requirements aggressively and prefers deletion\n\nIn Codex CLI, switch modes by invoking the main skill:\n    \n    \n    @ponytail lite\n    \n    \n    @ponytail full\n    \n    \n    @ponytail ultra\n\nUse `lite` when you want the agent to follow the requested design but still point out a cheaper option.\n\nMy default would be `full`. It makes the smaller choice without turning every task into an argument.\n\nI would reserve `ultra` for experiments, cleanup work, and features whose requirements are still negotiable. It may challenge the feature itself, which is useful only when the feature is open to challenge.\n\nTo turn Ponytail off, say `normal mode` or run:\n    \n    \n    @ponytail off\n\nThe active mode lasts for the session.\n\n## Set the default mode\n\nYou can change the default with a config file at:\n    \n    \n    ~/.config/ponytail/config.json\n\nFor example, this starts new sessions in `lite` mode:\n    \n    \n    {\n      \"defaultMode\": \"lite\"\n    }\n\nValid values are `lite`, `full`, `ultra`, and `off`.\n\nYou can also set `PONYTAIL_DEFAULT_MODE`. The environment variable takes priority over the config file.\n\n## Review the current diff\n\nPonytail includes a separate review skill for code that already exists in the current diff.\n\nRun:\n    \n    \n    @ponytail-review Review the current diff\n\nThis is a read-only review. It looks for a small set of problems:\n\n  * dead code and unused flexibility\n  * standard-library features rebuilt by hand\n  * dependencies that duplicate native platform features\n  * abstractions with one implementation\n  * logic that can keep its behavior with fewer lines\n\nThe result is a delete list. It does not apply the changes.\n\nThis narrow scope is useful. A normal code review still needs to inspect correctness, security, performance, and product behavior.\n\n`@ponytail-review` asks a different question: what can we remove?\n\n## Audit the whole repository\n\nThe review skill only checks the current diff.\n\nTo inspect the complete codebase, run:\n    \n    \n    @ponytail-audit Audit this repository for over-engineering\n\nThe audit ranks opportunities to delete, reuse, or replace code. It also reports the number of lines and dependencies it thinks could disappear.\n\nTreat that total as a review estimate. Inspect every finding before changing working code.\n\nAn abstraction with one implementation may look unnecessary today. It may also represent a real boundary required by a plugin API, a test seam, or a second implementation living outside the repository.\n\nThe agent sees code. You still own the context.\n\n## Track deliberate shortcuts\n\nSometimes the smallest correct implementation has a known limit.\n\nPonytail uses a `ponytail:` comment to record that limit and the reason to revisit it:\n    \n    \n    // ponytail: linear scan, add an index when lists exceed 10,000 items\n\nThe comment should name both the ceiling and the upgrade trigger.\n\nYou can collect those comments with:\n    \n    \n    @ponytail-debt\n\nThis gives you a ledger of deliberate shortcuts. It also flags comments that do not say when the shortcut should be replaced.\n\nThe command does not change the code. If you want a permanent ledger, ask the agent to save the result after reviewing it.\n\n## What Ponytail will not remove\n\nThe shortest program is an empty file. That does not make it correct.\n\nPonytail has explicit boundaries. It should not simplify away:\n\n  * validation at a trust boundary\n  * error handling that prevents data loss\n  * security controls\n  * accessibility basics\n  * behavior you explicitly requested\n\nIt also asks for one runnable check around non-trivial logic. A branch, parser, loop, money flow, or security path should leave behind a small test or assertion that fails when the logic breaks.\n\nThese rules are the difference between **minimal code** and **careless code**.\n\nPonytail does not replace a security review. It does not prove that a small implementation is safe. It keeps safety inside its definition of “works.”\n\n## What the benchmark shows\n\nThe project includes a [published agentic benchmark](<https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/results/2026-06-18-agentic.md>).\n\nIt ran real Claude Code sessions against a FastAPI and React repository. The same agent completed 12 feature tasks with and without Ponytail, with four runs per task.\n\nAcross those tasks, the reported result was:\n\n  * 54% fewer added lines\n  * 22% fewer tokens\n  * 20% lower cost\n  * 27% less time\n\nThe largest reductions happened when a native browser control replaced a custom component. Date and color pickers dropped by more than 90%.\n\nThe backend tasks were different. When the existing implementation was already small, the versions converged. Ponytail found little to remove because little bloat existed.\n\nThe benchmark also ran 20 adversarial safety checks. The Ponytail runs kept every tested guard. A shorter “YAGNI and one-liners” prompt failed one path-traversal check.\n\nThese numbers are interesting, but they are not a promise for every project. The benchmark used one model, one repository, a small task set, and four runs per task.\n\nThe benchmark suggests Ponytail helps most when the task gives an agent room to overbuild.\n\n## How I would use Ponytail\n\nI would combine Ponytail with [fstack](</fstack/>), not replace fstack with it.\n\nfstack helps me clarify the task, write a plan, build in small steps, check the result, and ship. Ponytail controls the implementation choices inside that workflow.\n\nFor a project such as [Port Pilot](</i-launched-port-pilot/>), I would use this sequence:\n\n  1. Record the product behavior and safety boundaries in `AGENTS.md`.\n  2. Use fstack to clarify and plan the change.\n  3. Build in Ponytail’s `full` mode.\n  4. Run the real tests and a normal correctness review.\n  5. Run `@ponytail-review` against the final diff.\n\nPort Pilot can stop local processes. Confirmation rules and process-safety checks are part of the product, not optional complexity. Ponytail may shrink the implementation around those rules, but it should never delete the rules.\n\nI would not use Ponytail as the product manager. It cannot decide whether a planned abstraction is part of a public API, whether tomorrow’s integration is already contracted, or whether a hardware calibration setting looks unused because the test machine happens to be accurate.\n\nI would also keep it away from prose. Ponytail’s own scope is coding. Writing needs a different review skill with different rules.\n\n## Use Ponytail with other agents\n\nPonytail also supports OpenCode, Cursor, Windsurf, Cline, Kiro, Zed, and several other coding agents.\n\nSome install it as a plugin. Others load a copied rules file or the repository’s `AGENTS.md`.\n\nThe [Ponytail repository](<https://github.com/DietrichGebert/ponytail>) keeps the current instructions for each host.\n\nIf you want to understand how the bundled `SKILL.md` files, activation descriptions, hooks, and supporting resources fit together, my free [AI Agent Skills course](</courses/ai-agent-skills/>) builds that model from the beginning.\n\n## Remove Ponytail\n\nTo remove the Codex plugin, run:\n    \n    \n    codex plugin remove ponytail\n\nThis removes the plugin files. Ponytail may leave its mode config under `~/.config/ponytail/`.\n\nIf you want to remove that state too, follow the cleanup instructions in the repository before uninstalling the plugin. The cleanup script is part of the plugin, so removing the plugin first also removes the script.\n\nPonytail packages one small idea as a persistent rule: understand everything you need, then build only what you need. I think that is a good default for a coding agent.\n\nTagged: [AI](</tags/ai/>) · [All topics](</blog/#topics>)\n\n[Follow @flaviocopes](<https://twitter.com/flaviocopes>)\n\nWant me to talk about your product? You can [sponsor this site](</sponsor/>).\n\nFree ebooks\n\nWant to go deeper? Subscribe to my newsletter and open my free download library for books, courses, and software.\n\n[Open the download library →](</access>)\n\n~~~\n\nRelated posts about ai:\n\n  * [A deep dive into Pi](</pi/>)\n\n  * [How AI models are measured and compared](</ai-benchmarks/>)\n\n  * [A deep dive into fx](</fx/>)\n\n  * [Slop grenades](</slop-grenades/>)\n\n  * [Grok Bot vs Cursor Projects: when to use which](</grok-bot-vs-cursor-projects/>)\n\n  * [Solving an issue with sending emails on my server](</sendy-slow-emails-ai/>)\n\n  * [Updating self-hosted Plausible Analytics using AI](</updating-plausible-with-ai/>)\n\n  * [A deep dive into Jev, TypeSafe's System One model](</jev/>)","body_html":"<p><a href=\"/\">Home</a> / <a href=\"/tags/ai/\">AI</a></p>\n<h1 id=\"a-deep-dive-into-ponytail\">A deep dive into Ponytail</h1>\n<p>By <a href=\"/about/\">Flavio Copes</a></p>\n<p>Updated Sep 17, 2026</p>\n<p>Learn how Ponytail stops AI coding agents from over-engineering: install it in ChatGPT, Codex, or Claude Code, choose a mode, and review diffs.</p>\n<pre><code>\nAI coding agents like writing code.\n\nSometimes they like it too much.\n\nYou ask for a date field. The agent installs a date picker, adds a wrapper component, creates a theme file, and starts discussing time zones.\n\nThe browser already has this:\n    \n    \n    &lt;input type=&quot;date&quot; /&gt;\n\n[Ponytail](&lt;https://ponytail.dev/&gt;) is a ruleset that pushes coding agents toward that kind of answer.\n\nIt was created by Dietrich Gebert. The project describes its approach as the lazy senior developer: understand the problem, find the smallest correct solution, and stop.\n\nLazy does not mean careless here. Ponytail tells the agent to preserve validation, error handling, security, accessibility, and anything you explicitly requested. It aims to reduce the code you have to own, not play code golf.\n\n## How Ponytail works\n\nPonytail gives the agent a seven-step ladder.\n\nThe agent stops at the first step that solves the real problem:\n\n  1. Does this need to exist?\n  2. Does the codebase already have it?\n  3. Does the standard library provide it?\n  4. Does the native platform provide it?\n  5. Does an installed dependency already solve it?\n  6. Can it be one line?\n  7. Only then, write the minimum new code.\n\nThe order makes the agent reuse an existing date-field component before reaching for the native HTML input.\n\nSuppose the field needs a calendar system the browser input cannot represent. The native feature no longer solves the task, so the agent keeps moving down the ladder.\n\nPonytail also tells the agent to inspect the relevant code first. It should trace the real flow, search for existing helpers, and fix a bug at its shared cause instead of patching one visible symptom.\n\nIt is lazy about the solution, not the investigation.\n\n## Install Ponytail in the ChatGPT app\n\nOpen **Settings** , then select **Plugins**.\n\nOpen the **Add** menu and choose **Add plugin marketplace**. Enter `DietrichGebert/ponytail` as the source, then click **Add marketplace** :\n\n![Adding the Ponytail plugin marketplace in the ChatGPT app](/images/ponytail/chatgpt-add-marketplace.webp)\n\nReturn to the Plugins page and select **Personal**. Find Ponytail and click **Install** :\n\n![Installing Ponytail from the personal plugin marketplace in the ChatGPT app](/images/ponytail/chatgpt-install-plugin.webp)\n\nOpen a new Codex task and type:\n    \n    \n    /ponytail\n\nThe composer shows the installed Ponytail skills. Select the main Ponytail skill and send the prompt:\n\n![Activating the Ponytail skill from a Codex task in the ChatGPT app](/images/ponytail/chatgpt-use-ponytail.webp)\n\nPonytail confirms the active intensity. The default is `full`.\n\n## Install Ponytail in Codex CLI\n\nPonytail is available as a Codex plugin.\n\nAdd its marketplace:\n    \n    \n    codex plugin marketplace add DietrichGebert/ponytail\n\nThen install the plugin:\n    \n    \n    codex plugin add ponytail@ponytail\n\nStart Codex and open the hooks screen:\n    \n    \n    /hooks\n\nReview and trust Ponytail’s lifecycle hooks, then start a new task.\n\nThe hooks activate Ponytail when a session starts and pass the active mode to subagents.\n\nIf you install from the terminal and use the ChatGPT app too, restart the app. It will load the same plugin.\n\nPonytail starts in `full` mode by default. You can check that the bundled skills are available with:\n    \n    \n    @ponytail-help\n\nCodex CLI invokes skills with `@`. The ChatGPT app and Claude Code use slash commands.\n\n## Install Ponytail in Claude Code\n\nIn Claude Code, send these as two separate prompts.\n\nFirst, add the marketplace:\n    \n    \n    /plugin marketplace add DietrichGebert/ponytail\n\nWait for that command to finish. Then install Ponytail:\n    \n    \n    /plugin install ponytail@ponytail\n\nStart a new session after installation. Ponytail activates in `full` mode by default.\n\nYou can check the current mode with:\n    \n    \n    /ponytail\n\nThe same two installation commands work in the Code tab of the Claude Code desktop app.\n\nComing soon · waiting lists open\n\nAI Workshop and Solo Lab: new cohorts\n\nSolo Lab starts on 26 October: three weeks on finding customers and making a living from your own software. AI Workshop starts on 30 November: three weeks on directing coding agents and reviewing what they build. Join the waiting lists and I&#39;ll email you when enrollment opens.\n\n[See the cohorts →](&lt;/cohorts/&gt;)\n\n## Install Ponytail in other CLIs\n\nGitHub Copilot CLI, Gemini CLI and Pi have their own install commands. Copilot CLI uses the same marketplace flow as Codex:\n    \n    \n    copilot plugin marketplace add DietrichGebert/ponytail\n    copilot plugin install ponytail@ponytail\n\nGemini CLI installs it as an extension:\n    \n    \n    gemini extensions install https://github.com/DietrichGebert/ponytail\n\nAnd Pi:\n    \n    \n    pi install git:github.com/DietrichGebert/ponytail\n\nThe Claude Code and Codex plugins use small Node.js lifecycle hooks. If `node` is missing from the non-interactive shell’s `PATH`, the skills still work, but automatic activation does not.\n\nThe command examples below use Codex CLI’s `@` syntax. In the ChatGPT app or Claude Code, replace `@` with `/`.\n\n## Try it on a small feature\n\nLet’s give the agent a task that often attracts too much code:\n    \n    \n    Add a birthday field to the profile form.\n    Use the existing form styles and keep the current validation behavior.\n\nIn `full` mode, Ponytail first inspects the form and its existing components.\n\nIf plain HTML covers the requirement, the result might be close to this:\n    \n    \n    &lt;label for=&quot;birthday&quot;&gt;Birthday&lt;/label&gt;\n    &lt;input id=&quot;birthday&quot; name=&quot;birthday&quot; type=&quot;date&quot; /&gt;\n\nThe result needs no new package, calendar wrapper, or custom date parser.\n\nNotice that the prompt did not say “write the fewest lines possible.” Ponytail applies the ladder as a standing coding rule.\n\nIf the application already uses a form component with error messages and accessible descriptions, the agent should use that component instead. Reusing the codebase is higher on the ladder than using the browser directly.\n\n## Choose an intensity\n\nPonytail has three active modes:\n\n  * `lite` builds what you requested and mentions a smaller alternative\n  * `full` enforces the complete ladder and is the default\n  * `ultra` treats speculative requirements aggressively and prefers deletion\n\nIn Codex CLI, switch modes by invoking the main skill:\n    \n    \n    @ponytail lite\n    \n    \n    @ponytail full\n    \n    \n    @ponytail ultra\n\nUse `lite` when you want the agent to follow the requested design but still point out a cheaper option.\n\nMy default would be `full`. It makes the smaller choice without turning every task into an argument.\n\nI would reserve `ultra` for experiments, cleanup work, and features whose requirements are still negotiable. It may challenge the feature itself, which is useful only when the feature is open to challenge.\n\nTo turn Ponytail off, say `normal mode` or run:\n    \n    \n    @ponytail off\n\nThe active mode lasts for the session.\n\n## Set the default mode\n\nYou can change the default with a config file at:\n    \n    \n    ~/.config/ponytail/config.json\n\nFor example, this starts new sessions in `lite` mode:\n    \n    \n    {\n      &quot;defaultMode&quot;: &quot;lite&quot;\n    }\n\nValid values are `lite`, `full`, `ultra`, and `off`.\n\nYou can also set `PONYTAIL_DEFAULT_MODE`. The environment variable takes priority over the config file.\n\n## Review the current diff\n\nPonytail includes a separate review skill for code that already exists in the current diff.\n\nRun:\n    \n    \n    @ponytail-review Review the current diff\n\nThis is a read-only review. It looks for a small set of problems:\n\n  * dead code and unused flexibility\n  * standard-library features rebuilt by hand\n  * dependencies that duplicate native platform features\n  * abstractions with one implementation\n  * logic that can keep its behavior with fewer lines\n\nThe result is a delete list. It does not apply the changes.\n\nThis narrow scope is useful. A normal code review still needs to inspect correctness, security, performance, and product behavior.\n\n`@ponytail-review` asks a different question: what can we remove?\n\n## Audit the whole repository\n\nThe review skill only checks the current diff.\n\nTo inspect the complete codebase, run:\n    \n    \n    @ponytail-audit Audit this repository for over-engineering\n\nThe audit ranks opportunities to delete, reuse, or replace code. It also reports the number of lines and dependencies it thinks could disappear.\n\nTreat that total as a review estimate. Inspect every finding before changing working code.\n\nAn abstraction with one implementation may look unnecessary today. It may also represent a real boundary required by a plugin API, a test seam, or a second implementation living outside the repository.\n\nThe agent sees code. You still own the context.\n\n## Track deliberate shortcuts\n\nSometimes the smallest correct implementation has a known limit.\n\nPonytail uses a `ponytail:` comment to record that limit and the reason to revisit it:\n    \n    \n    // ponytail: linear scan, add an index when lists exceed 10,000 items\n\nThe comment should name both the ceiling and the upgrade trigger.\n\nYou can collect those comments with:\n    \n    \n    @ponytail-debt\n\nThis gives you a ledger of deliberate shortcuts. It also flags comments that do not say when the shortcut should be replaced.\n\nThe command does not change the code. If you want a permanent ledger, ask the agent to save the result after reviewing it.\n\n## What Ponytail will not remove\n\nThe shortest program is an empty file. That does not make it correct.\n\nPonytail has explicit boundaries. It should not simplify away:\n\n  * validation at a trust boundary\n  * error handling that prevents data loss\n  * security controls\n  * accessibility basics\n  * behavior you explicitly requested\n\nIt also asks for one runnable check around non-trivial logic. A branch, parser, loop, money flow, or security path should leave behind a small test or assertion that fails when the logic breaks.\n\nThese rules are the difference between **minimal code** and **careless code**.\n\nPonytail does not replace a security review. It does not prove that a small implementation is safe. It keeps safety inside its definition of “works.”\n\n## What the benchmark shows\n\nThe project includes a [published agentic benchmark](&lt;https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/results/2026-06-18-agentic.md&gt;).\n\nIt ran real Claude Code sessions against a FastAPI and React repository. The same agent completed 12 feature tasks with and without Ponytail, with four runs per task.\n\nAcross those tasks, the reported result was:\n\n  * 54% fewer added lines\n  * 22% fewer tokens\n  * 20% lower cost\n  * 27% less time\n\nThe largest reductions happened when a native browser control replaced a custom component. Date and color pickers dropped by more than 90%.\n\nThe backend tasks were different. When the existing implementation was already small, the versions converged. Ponytail found little to remove because little bloat existed.\n\nThe benchmark also ran 20 adversarial safety checks. The Ponytail runs kept every tested guard. A shorter “YAGNI and one-liners” prompt failed one path-traversal check.\n\nThese numbers are interesting, but they are not a promise for every project. The benchmark used one model, one repository, a small task set, and four runs per task.\n\nThe benchmark suggests Ponytail helps most when the task gives an agent room to overbuild.\n\n## How I would use Ponytail\n\nI would combine Ponytail with [fstack](&lt;/fstack/&gt;), not replace fstack with it.\n\nfstack helps me clarify the task, write a plan, build in small steps, check the result, and ship. Ponytail controls the implementation choices inside that workflow.\n\nFor a project such as [Port Pilot](&lt;/i-launched-port-pilot/&gt;), I would use this sequence:\n\n  1. Record the product behavior and safety boundaries in `AGENTS.md`.\n  2. Use fstack to clarify and plan the change.\n  3. Build in Ponytail’s `full` mode.\n  4. Run the real tests and a normal correctness review.\n  5. Run `@ponytail-review` against the final diff.\n\nPort Pilot can stop local processes. Confirmation rules and process-safety checks are part of the product, not optional complexity. Ponytail may shrink the implementation around those rules, but it should never delete the rules.\n\nI would not use Ponytail as the product manager. It cannot decide whether a planned abstraction is part of a public API, whether tomorrow’s integration is already contracted, or whether a hardware calibration setting looks unused because the test machine happens to be accurate.\n\nI would also keep it away from prose. Ponytail’s own scope is coding. Writing needs a different review skill with different rules.\n\n## Use Ponytail with other agents\n\nPonytail also supports OpenCode, Cursor, Windsurf, Cline, Kiro, Zed, and several other coding agents.\n\nSome install it as a plugin. Others load a copied rules file or the repository’s `AGENTS.md`.\n\nThe [Ponytail repository](&lt;https://github.com/DietrichGebert/ponytail&gt;) keeps the current instructions for each host.\n\nIf you want to understand how the bundled `SKILL.md` files, activation descriptions, hooks, and supporting resources fit together, my free [AI Agent Skills course](&lt;/courses/ai-agent-skills/&gt;) builds that model from the beginning.\n\n## Remove Ponytail\n\nTo remove the Codex plugin, run:\n    \n    \n    codex plugin remove ponytail\n\nThis removes the plugin files. Ponytail may leave its mode config under `~/.config/ponytail/`.\n\nIf you want to remove that state too, follow the cleanup instructions in the repository before uninstalling the plugin. The cleanup script is part of the plugin, so removing the plugin first also removes the script.\n\nPonytail packages one small idea as a persistent rule: understand everything you need, then build only what you need. I think that is a good default for a coding agent.\n\nTagged: [AI](&lt;/tags/ai/&gt;) · [All topics](&lt;/blog/#topics&gt;)\n\n[Follow @flaviocopes](&lt;https://twitter.com/flaviocopes&gt;)\n\nWant me to talk about your product? You can [sponsor this site](&lt;/sponsor/&gt;).\n\nFree ebooks\n\nWant to go deeper? Subscribe to my newsletter and open my free download library for books, courses, and software.\n\n[Open the download library →](&lt;/access&gt;)\n</code></pre>\n<p>Related posts about ai:</p>\n<ul><li><a href=\"/pi/\">A deep dive into Pi</a></li><li><a href=\"/ai-benchmarks/\">How AI models are measured and compared</a></li><li><a href=\"/fx/\">A deep dive into fx</a></li><li><a href=\"/slop-grenades/\">Slop grenades</a></li><li><a href=\"/grok-bot-vs-cursor-projects/\">Grok Bot vs Cursor Projects: when to use which</a></li><li><a href=\"/sendy-slow-emails-ai/\">Solving an issue with sending emails on my server</a></li><li><a href=\"/updating-plausible-with-ai/\">Updating self-hosted Plausible Analytics using AI</a></li><li><a href=\"/jev/\">A deep dive into Jev, TypeSafe&#39;s System One model</a></li></ul>","headings":[{"level":1,"text":"A deep dive into Ponytail","id":"a-deep-dive-into-ponytail"}]}}