{"article":{"slug":"agentic-sdd-giving-claude-code-agents-a-real-engineering-process","title":"Agentic-SDD: Giving Claude Code Agents a Real Engineering Process","subtitle":null,"summary":"Agentic-SDD is a Claude Code plugin that gates coding agents through a six-stage require→plan→analyze→implement→verify→fix pipeline with on-disk status, specialist agents, and a bundled knowledge-search engine.","content_type":"tutorial","language":"en","canonical_url":"https://dev.to/mallikarjunht/agentic-sdd-giving-claude-code-agents-a-real-engineering-process-5249","author":{"name":"Mallikarjun H T","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"DEV Community","url":"https://dev.to/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"},{"name":"Developer Tools","slug":"developer-tools","url":"https://listedarticles.com/topics/developer-tools"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"}],"about_listings":[],"cover_image_url":"https://media2.dev.to/dynamic/image/width=1200,height=627,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbcq95kksitu6xwi8wueq.png","license":"all-rights-reserved","word_count":1303,"reading_minutes":6,"published_at":"2026-10-01T18:10:08.900Z","added_at":"2026-10-01T18:10:08.900Z","updated_at":"2026-10-01T18:10:08.900Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/agentic-sdd-giving-claude-code-agents-a-real-engineering-process","markdown_url":"https://listedarticles.com/articles/agentic-sdd-giving-claude-code-agents-a-real-engineering-process.md","example":false,"citation":"Mallikarjun H T, DEV Community. \"Agentic-SDD: Giving Claude Code Agents a Real Engineering Process.\" 1 Oct 2026. https://dev.to/mallikarjunht/agentic-sdd-giving-claude-code-agents-a-real-engineering-process-5249 (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://dev.to/mallikarjunht/agentic-sdd-giving-claude-code-agents-a-real-engineering-process-5249"},"body_markdown":"# Agentic-SDD: Giving Claude Code Agents a Real Engineering Process (and a Knowledge Base That Doesn't Depend on Anyone Else)\n\nAsk an AI coding agent to \"fix the bug\" and it will fix *a* bug — usually the\n\none nearest the surface, in whatever file it opens first. It rarely stops to\n\nask whether the fix addresses the real requirement, whether a sibling caller\n\nhas the same problem, or whether a reviewer would actually sign off on it.\n\nThat's not a model failure so much as a missing process: humans don't skip\n\ndesign and review because we're smarter in the moment, we skip it because\n\nnobody built the discipline into the loop.\n\n**Agentic-SDD** is a Claude Code plugin that tries to build that discipline\n\nin — a six-stage pipeline (`require → plan → analyze → implement → verify →`) gated by a persistent on-disk status file instead of conversation\n\nfix\n\nmemory, backed by an 18-agent specialist roster, with its own bundled\n\nknowledge-search engine so it doesn't depend on any other plugin to ground\n\nits decisions in your actual codebase.\n\nRepo: github.com/MallikarjunHt/agentic-sdd\n\n## The core idea: six gated stages, not a chat loop\n\nA `/sdd.require <feature-id>` call doesn't just write code. It walks a real\n\nfeature through:\n\n1. \n**Require.** A business-analyst agent turns a raw ticket/request into`1-spec.md` — problem, proposed solution, explicit non-goals, risks, and\nGiven/When/Then acceptance criteria. You approve it before anything else\nhappens.\n2. \n**Plan.** An architect agent turns the approved spec into`2-plan.md` and`3-tasks.md` — a concrete Definition of Done, a file map, and a numbered\ntask list, each task tagged with a suggested owner from a 12-role\nspecialist bench (Java, Angular, React, Python, UI, DevOps, QA, BA, DBA,\nSecurity, Architect, Senior Dev), scored against trigger keywords. You\napprove this too.\n3. \n**Analyze.** A dedicated gap-analysis agent checks the plan against the\nspec's acceptance criteria*and* the real current codebase, before a\nsingle line of code is written. CRITICAL findings block progress until\nresolved or explicitly descoped.\n4. \n**Implement.** Each task is implemented by its routed specialist, at a\nmodel tier (cheap/mid/expensive) scored from that task's own complexity —\na one-line config change doesn't need the same model as a cross-cutting\nrefactor.\n5. \n**Verify.** Three separate agents check constitution conformance,\nquality/security, and test adequacy — plus a fourth, a dedicated devil's\nadvocate, whose only job is to look for what a checklist-shaped review\nstructurally can't catch: concurrency, migration/rollback risk, backward\ncompatibility. A PASS proposes a living-documentation diff, reviewed in\nthe normal PR, never committed silently.\n6. \n**Fix.** If verify fails, a fixer agent gets exactly three attempts before\nit has to stop and recommend a plan change instead of looping forever —\na real circuit breaker, not an infinite retry.\n\nEvery stage after `require` reads and writes a `status.json` per feature, so\n\na crashed session resumes instead of restarting, and stages genuinely cannot\n\nrun out of order. That one design choice — a file as the gate, not the\n\nmodel's own sense of \"did I already do this?\" — is doing most of the\n\nreliability work here.\n\n## The profile system: one plugin, any repo\n\nRather than hardcoding one tech stack's conventions, `/sdd.init` writes a\n\n`specs/constitution.md` grounded in *that specific target repo's* real\n\nconventions — build tool, test framework, lint command, forbidden/required\n\npatterns — not a generic textbook standard. A repo's own detected config\n\nalways wins over the constitution file when they disagree. One plugin\n\ninstall, many target repos, each honestly represented.\n\n## The part I almost shipped wrong: knowledge grounding\n\nThe original design had every stage's \"gather context\" step call out to an\n\nMCP tool contract — any server exposing a `knowledge_search`-shaped tool\n\nname could plug in. Clean, decoupled, and *wrong* for the actual goal: if\n\nsomeone uninstalls whatever server was providing that tool, the plugin loses\n\na capability it should always have.\n\nSo I vendored the actual search engine — a small Lucene-based Java CLI, BM25\n\nfull-text search with an optional ONNX vector-embedding upgrade — directly\n\ninto the plugin, under its own renamed package so it carries no trace of\n\nwhere it came from. A few lessons from doing that for real, not\n\nhypothetically:\n\n- \n**I almost committed a 150MB jar.** GitHub rejects pushes over 100MB per\nfile. The fix was obvious in retrospect: vendor the*source* (a few\nhundred KB), build the jar locally with Maven on first use, and gitignore\nthe build output. The plugin repo stays small; the binary gets built once\nper machine.\n- \n**The embedding model is optional by design, not by accident.** Full\nsemantic search needs a ~90MB ONNX model I also didn't want to ship by\ndefault. The engine already had a clean fallback — no model vendored means\nBM25-only, reported plainly by its own`doctor` command — so the default\ninstall just... works, lexical search only, with vector search as an\nopt-in upgrade documented for later.\n- \n**A relative default path is a trap.** The CLI's`--index-dir` and`--models-dir` flags default to cwd-relative paths. I'd already seen this\nexact class of bug once before (a tool silently reporting \"no embedding\nmodel\" because it was invoked from the wrong working directory, not\nbecause the model was actually missing) — so every call from inside the\nplugin passes both as absolute paths. Boring, but it's the difference\nbetween \"works on my machine\" and \"works.\"\n\n## Proof, not a pitch\n\nI ran the full pipeline end-to-end against a real Dependabot CVE alert in a\n\nproduction Java/npm monorepo — a transitive `yaml` dependency vulnerable to\n\nstack-overflow on deeply nested input. All six stages ran for real: the spec\n\ncaptured the actual CVE and affected manifest, the plan scoped it to a\n\none-line `npm` override (no functional changes, matching the kind of\n\ndiscipline a dependency-only fix needs), analyze confirmed `yaml` was\n\ntransitive rather than direct (so an override, not a direct bump, was the\n\nright mechanism), implement regenerated the lockfile properly instead of\n\nhand-editing it, and verify confirmed the resolved version actually cleared\n\nthe vulnerable range — straight from the regenerated lockfile, since the\n\nenvironment's own vulnerability-feed proxy wasn't reachable that day. Two\n\nreal commits, nothing fabricated, nothing silently skipped.\n\n## What's still a known gap\n\nStated plainly, because a tool that hides its own limitations is worse than\n\none that states them:\n\n- No automated validation of the plugin's own prompts/templates — there's no compiler for a markdown instruction file. A change is proven by running it against a real feature, not by a test suite.\n- No cross-repo coordination — each install targets one repo via its own config. A feature spanning two repos needs two separate, manually coordinated runs.\n- The wiki-mirror feature (a page-per-stage sync to Confluence or similar) is the one remaining capability that depends on an external MCP server being connected, and degrades to \"local artifact only\" when it isn't — deliberately, since a wiki is a genuinely external system, not something a plugin should vendor a copy of.\n\nIf you're building something similar — or just tired of watching an agent\n\nconfidently fix the wrong thing — the repo's up, MIT licensed, and the\n\npipeline is designed to be read, not just run:\n\ngithub.com/MallikarjunHt/agentic-sdd.\n\n- No automated validation of the plugin's own prompts/templates — there's no compiler for a markdown instruction file. A change is proven by running it against a real feature, not by a test suite.\n- No cross-repo coordination — each install targets one repo via its own config. A feature spanning two repos needs two separate, manually coordinated runs.\n- The wiki-mirror feature (a page-per-stage sync to Confluence or similar) is the one remaining capability that depends on an external MCP server being connected, and degrades to \"local artifact only\" when it isn't — deliberately, since a wiki is a genuinely external system, not something a plugin should vendor a copy of.\n\nIf you're building something similar — or just tired of watching an agent\n\nconfidently fix the wrong thing — the repo's up, MIT licensed, and the\n\npipeline is designed to be read, not just run:\n\ngithub.com/MallikarjunHt/agentic-sdd.","body_html":"<h1 id=\"agentic-sdd-giving-claude-code-agents-a-real-engineering-process\">Agentic-SDD: Giving Claude Code Agents a Real Engineering Process (and a Knowledge Base That Doesn&#39;t Depend on Anyone Else)</h1>\n<p>Ask an AI coding agent to &quot;fix the bug&quot; and it will fix *a* bug — usually the</p>\n<p>one nearest the surface, in whatever file it opens first. It rarely stops to</p>\n<p>ask whether the fix addresses the real requirement, whether a sibling caller</p>\n<p>has the same problem, or whether a reviewer would actually sign off on it.</p>\n<p>That&#39;s not a model failure so much as a missing process: humans don&#39;t skip</p>\n<p>design and review because we&#39;re smarter in the moment, we skip it because</p>\n<p>nobody built the discipline into the loop.</p>\n<p><strong>Agentic-SDD</strong> is a Claude Code plugin that tries to build that discipline</p>\n<p>in — a six-stage pipeline (<code>require → plan → analyze → implement → verify →</code>) gated by a persistent on-disk status file instead of conversation</p>\n<p>fix</p>\n<p>memory, backed by an 18-agent specialist roster, with its own bundled</p>\n<p>knowledge-search engine so it doesn&#39;t depend on any other plugin to ground</p>\n<p>its decisions in your actual codebase.</p>\n<p>Repo: github.com/MallikarjunHt/agentic-sdd</p>\n<h2 id=\"the-core-idea-six-gated-stages-not-a-chat-loop\">The core idea: six gated stages, not a chat loop</h2>\n<p>A <code>/sdd.require &lt;feature-id&gt;</code> call doesn&#39;t just write code. It walks a real</p>\n<p>feature through:</p>\n<ol><li></li></ol>\n<p><strong>Require.</strong> A business-analyst agent turns a raw ticket/request into<code>1-spec.md</code> — problem, proposed solution, explicit non-goals, risks, and\nGiven/When/Then acceptance criteria. You approve it before anything else\nhappens.</p>\n<ol start=\"2\"><li></li></ol>\n<p><strong>Plan.</strong> An architect agent turns the approved spec into<code>2-plan.md</code> and<code>3-tasks.md</code> — a concrete Definition of Done, a file map, and a numbered\ntask list, each task tagged with a suggested owner from a 12-role\nspecialist bench (Java, Angular, React, Python, UI, DevOps, QA, BA, DBA,\nSecurity, Architect, Senior Dev), scored against trigger keywords. You\napprove this too.</p>\n<ol start=\"3\"><li></li></ol>\n<p><strong>Analyze.</strong> A dedicated gap-analysis agent checks the plan against the\nspec&#39;s acceptance criteria<em>and</em> the real current codebase, before a\nsingle line of code is written. CRITICAL findings block progress until\nresolved or explicitly descoped.</p>\n<ol start=\"4\"><li></li></ol>\n<p><strong>Implement.</strong> Each task is implemented by its routed specialist, at a\nmodel tier (cheap/mid/expensive) scored from that task&#39;s own complexity —\na one-line config change doesn&#39;t need the same model as a cross-cutting\nrefactor.</p>\n<ol start=\"5\"><li></li></ol>\n<p><strong>Verify.</strong> Three separate agents check constitution conformance,\nquality/security, and test adequacy — plus a fourth, a dedicated devil&#39;s\nadvocate, whose only job is to look for what a checklist-shaped review\nstructurally can&#39;t catch: concurrency, migration/rollback risk, backward\ncompatibility. A PASS proposes a living-documentation diff, reviewed in\nthe normal PR, never committed silently.</p>\n<ol start=\"6\"><li></li></ol>\n<p><strong>Fix.</strong> If verify fails, a fixer agent gets exactly three attempts before\nit has to stop and recommend a plan change instead of looping forever —\na real circuit breaker, not an infinite retry.</p>\n<p>Every stage after <code>require</code> reads and writes a <code>status.json</code> per feature, so</p>\n<p>a crashed session resumes instead of restarting, and stages genuinely cannot</p>\n<p>run out of order. That one design choice — a file as the gate, not the</p>\n<p>model&#39;s own sense of &quot;did I already do this?&quot; — is doing most of the</p>\n<p>reliability work here.</p>\n<h2 id=\"the-profile-system-one-plugin-any-repo\">The profile system: one plugin, any repo</h2>\n<p>Rather than hardcoding one tech stack&#39;s conventions, <code>/sdd.init</code> writes a</p>\n<p><code>specs/constitution.md</code> grounded in <em>that specific target repo&#39;s</em> real</p>\n<p>conventions — build tool, test framework, lint command, forbidden/required</p>\n<p>patterns — not a generic textbook standard. A repo&#39;s own detected config</p>\n<p>always wins over the constitution file when they disagree. One plugin</p>\n<p>install, many target repos, each honestly represented.</p>\n<h2 id=\"the-part-i-almost-shipped-wrong-knowledge-grounding\">The part I almost shipped wrong: knowledge grounding</h2>\n<p>The original design had every stage&#39;s &quot;gather context&quot; step call out to an</p>\n<p>MCP tool contract — any server exposing a <code>knowledge_search</code>-shaped tool</p>\n<p>name could plug in. Clean, decoupled, and <em>wrong</em> for the actual goal: if</p>\n<p>someone uninstalls whatever server was providing that tool, the plugin loses</p>\n<p>a capability it should always have.</p>\n<p>So I vendored the actual search engine — a small Lucene-based Java CLI, BM25</p>\n<p>full-text search with an optional ONNX vector-embedding upgrade — directly</p>\n<p>into the plugin, under its own renamed package so it carries no trace of</p>\n<p>where it came from. A few lessons from doing that for real, not</p>\n<p>hypothetically:</p>\n<ul><li></li></ul>\n<p><strong>I almost committed a 150MB jar.</strong> GitHub rejects pushes over 100MB per\nfile. The fix was obvious in retrospect: vendor the<em>source</em> (a few\nhundred KB), build the jar locally with Maven on first use, and gitignore\nthe build output. The plugin repo stays small; the binary gets built once\nper machine.</p>\n<ul><li></li></ul>\n<p><strong>The embedding model is optional by design, not by accident.</strong> Full\nsemantic search needs a ~90MB ONNX model I also didn&#39;t want to ship by\ndefault. The engine already had a clean fallback — no model vendored means\nBM25-only, reported plainly by its own<code>doctor</code> command — so the default\ninstall just... works, lexical search only, with vector search as an\nopt-in upgrade documented for later.</p>\n<ul><li></li></ul>\n<p><strong>A relative default path is a trap.</strong> The CLI&#39;s<code>--index-dir</code> and<code>--models-dir</code> flags default to cwd-relative paths. I&#39;d already seen this\nexact class of bug once before (a tool silently reporting &quot;no embedding\nmodel&quot; because it was invoked from the wrong working directory, not\nbecause the model was actually missing) — so every call from inside the\nplugin passes both as absolute paths. Boring, but it&#39;s the difference\nbetween &quot;works on my machine&quot; and &quot;works.&quot;</p>\n<h2 id=\"proof-not-a-pitch\">Proof, not a pitch</h2>\n<p>I ran the full pipeline end-to-end against a real Dependabot CVE alert in a</p>\n<p>production Java/npm monorepo — a transitive <code>yaml</code> dependency vulnerable to</p>\n<p>stack-overflow on deeply nested input. All six stages ran for real: the spec</p>\n<p>captured the actual CVE and affected manifest, the plan scoped it to a</p>\n<p>one-line <code>npm</code> override (no functional changes, matching the kind of</p>\n<p>discipline a dependency-only fix needs), analyze confirmed <code>yaml</code> was</p>\n<p>transitive rather than direct (so an override, not a direct bump, was the</p>\n<p>right mechanism), implement regenerated the lockfile properly instead of</p>\n<p>hand-editing it, and verify confirmed the resolved version actually cleared</p>\n<p>the vulnerable range — straight from the regenerated lockfile, since the</p>\n<p>environment&#39;s own vulnerability-feed proxy wasn&#39;t reachable that day. Two</p>\n<p>real commits, nothing fabricated, nothing silently skipped.</p>\n<h2 id=\"what-s-still-a-known-gap\">What&#39;s still a known gap</h2>\n<p>Stated plainly, because a tool that hides its own limitations is worse than</p>\n<p>one that states them:</p>\n<ul><li>No automated validation of the plugin&#39;s own prompts/templates — there&#39;s no compiler for a markdown instruction file. A change is proven by running it against a real feature, not by a test suite.</li><li>No cross-repo coordination — each install targets one repo via its own config. A feature spanning two repos needs two separate, manually coordinated runs.</li><li>The wiki-mirror feature (a page-per-stage sync to Confluence or similar) is the one remaining capability that depends on an external MCP server being connected, and degrades to &quot;local artifact only&quot; when it isn&#39;t — deliberately, since a wiki is a genuinely external system, not something a plugin should vendor a copy of.</li></ul>\n<p>If you&#39;re building something similar — or just tired of watching an agent</p>\n<p>confidently fix the wrong thing — the repo&#39;s up, MIT licensed, and the</p>\n<p>pipeline is designed to be read, not just run:</p>\n<p>github.com/MallikarjunHt/agentic-sdd.</p>\n<ul><li>No automated validation of the plugin&#39;s own prompts/templates — there&#39;s no compiler for a markdown instruction file. A change is proven by running it against a real feature, not by a test suite.</li><li>No cross-repo coordination — each install targets one repo via its own config. A feature spanning two repos needs two separate, manually coordinated runs.</li><li>The wiki-mirror feature (a page-per-stage sync to Confluence or similar) is the one remaining capability that depends on an external MCP server being connected, and degrades to &quot;local artifact only&quot; when it isn&#39;t — deliberately, since a wiki is a genuinely external system, not something a plugin should vendor a copy of.</li></ul>\n<p>If you&#39;re building something similar — or just tired of watching an agent</p>\n<p>confidently fix the wrong thing — the repo&#39;s up, MIT licensed, and the</p>\n<p>pipeline is designed to be read, not just run:</p>\n<p>github.com/MallikarjunHt/agentic-sdd.</p>","headings":[{"level":1,"text":"Agentic-SDD: Giving Claude Code Agents a Real Engineering Process (and a Knowledge Base That Doesn't Depend on Anyone Else)","id":"agentic-sdd-giving-claude-code-agents-a-real-engineering-process"},{"level":2,"text":"The core idea: six gated stages, not a chat loop","id":"the-core-idea-six-gated-stages-not-a-chat-loop"},{"level":2,"text":"The profile system: one plugin, any repo","id":"the-profile-system-one-plugin-any-repo"},{"level":2,"text":"The part I almost shipped wrong: knowledge grounding","id":"the-part-i-almost-shipped-wrong-knowledge-grounding"},{"level":2,"text":"Proof, not a pitch","id":"proof-not-a-pitch"},{"level":2,"text":"What's still a known gap","id":"what-s-still-a-known-gap"}]}}