Agentic-SDD: Giving Claude Code Agents a Real Engineering Process (and a Knowledge Base That Doesn't Depend on Anyone Else)
Ask an AI coding agent to "fix the bug" and it will fix *a* bug — usually the
one nearest the surface, in whatever file it opens first. It rarely stops to
ask whether the fix addresses the real requirement, whether a sibling caller
has the same problem, or whether a reviewer would actually sign off on it.
That's not a model failure so much as a missing process: humans don't skip
design and review because we're smarter in the moment, we skip it because
nobody built the discipline into the loop.
Agentic-SDD is a Claude Code plugin that tries to build that discipline
in — a six-stage pipeline (require → plan → analyze → implement → verify →) gated by a persistent on-disk status file instead of conversation
fix
memory, backed by an 18-agent specialist roster, with its own bundled
knowledge-search engine so it doesn't depend on any other plugin to ground
its decisions in your actual codebase.
Repo: github.com/MallikarjunHt/agentic-sdd
The core idea: six gated stages, not a chat loop
A /sdd.require <feature-id> call doesn't just write code. It walks a real
feature through:
Require. A business-analyst agent turns a raw ticket/request into1-spec.md — problem, proposed solution, explicit non-goals, risks, and
Given/When/Then acceptance criteria. You approve it before anything else
happens.
Plan. An architect agent turns the approved spec into2-plan.md and3-tasks.md — a concrete Definition of Done, a file map, and a numbered
task list, each task tagged with a suggested owner from a 12-role
specialist bench (Java, Angular, React, Python, UI, DevOps, QA, BA, DBA,
Security, Architect, Senior Dev), scored against trigger keywords. You
approve this too.
Analyze. A dedicated gap-analysis agent checks the plan against the spec's acceptance criteriaand the real current codebase, before a single line of code is written. CRITICAL findings block progress until resolved or explicitly descoped.
Implement. Each task is implemented by its routed specialist, at a model tier (cheap/mid/expensive) scored from that task's own complexity — a one-line config change doesn't need the same model as a cross-cutting refactor.
Verify. Three separate agents check constitution conformance, quality/security, and test adequacy — plus a fourth, a dedicated devil's advocate, whose only job is to look for what a checklist-shaped review structurally can't catch: concurrency, migration/rollback risk, backward compatibility. A PASS proposes a living-documentation diff, reviewed in the normal PR, never committed silently.
Fix. If verify fails, a fixer agent gets exactly three attempts before it has to stop and recommend a plan change instead of looping forever — a real circuit breaker, not an infinite retry.
Every stage after require reads and writes a status.json per feature, so
a crashed session resumes instead of restarting, and stages genuinely cannot
run out of order. That one design choice — a file as the gate, not the
model's own sense of "did I already do this?" — is doing most of the
reliability work here.
The profile system: one plugin, any repo
Rather than hardcoding one tech stack's conventions, /sdd.init writes a
specs/constitution.md grounded in that specific target repo's real
conventions — build tool, test framework, lint command, forbidden/required
patterns — not a generic textbook standard. A repo's own detected config
always wins over the constitution file when they disagree. One plugin
install, many target repos, each honestly represented.
The part I almost shipped wrong: knowledge grounding
The original design had every stage's "gather context" step call out to an
MCP tool contract — any server exposing a knowledge_search-shaped tool
name could plug in. Clean, decoupled, and wrong for the actual goal: if
someone uninstalls whatever server was providing that tool, the plugin loses
a capability it should always have.
So I vendored the actual search engine — a small Lucene-based Java CLI, BM25
full-text search with an optional ONNX vector-embedding upgrade — directly
into the plugin, under its own renamed package so it carries no trace of
where it came from. A few lessons from doing that for real, not
hypothetically:
I almost committed a 150MB jar. GitHub rejects pushes over 100MB per file. The fix was obvious in retrospect: vendor thesource (a few hundred KB), build the jar locally with Maven on first use, and gitignore the build output. The plugin repo stays small; the binary gets built once per machine.
The embedding model is optional by design, not by accident. Full
semantic search needs a ~90MB ONNX model I also didn't want to ship by
default. The engine already had a clean fallback — no model vendored means
BM25-only, reported plainly by its owndoctor command — so the default
install just... works, lexical search only, with vector search as an
opt-in upgrade documented for later.
A relative default path is a trap. The CLI's--index-dir and--models-dir flags default to cwd-relative paths. I'd already seen this
exact class of bug once before (a tool silently reporting "no embedding
model" because it was invoked from the wrong working directory, not
because the model was actually missing) — so every call from inside the
plugin passes both as absolute paths. Boring, but it's the difference
between "works on my machine" and "works."
Proof, not a pitch
I ran the full pipeline end-to-end against a real Dependabot CVE alert in a
production Java/npm monorepo — a transitive yaml dependency vulnerable to
stack-overflow on deeply nested input. All six stages ran for real: the spec
captured the actual CVE and affected manifest, the plan scoped it to a
one-line npm override (no functional changes, matching the kind of
discipline a dependency-only fix needs), analyze confirmed yaml was
transitive rather than direct (so an override, not a direct bump, was the
right mechanism), implement regenerated the lockfile properly instead of
hand-editing it, and verify confirmed the resolved version actually cleared
the vulnerable range — straight from the regenerated lockfile, since the
environment's own vulnerability-feed proxy wasn't reachable that day. Two
real commits, nothing fabricated, nothing silently skipped.
What's still a known gap
Stated plainly, because a tool that hides its own limitations is worse than
one that states them:
- No automated validation of the plugin's own prompts/templates — there's no compiler for a markdown instruction file. A change is proven by running it against a real feature, not by a test suite.
- No cross-repo coordination — each install targets one repo via its own config. A feature spanning two repos needs two separate, manually coordinated runs.
- The wiki-mirror feature (a page-per-stage sync to Confluence or similar) is the one remaining capability that depends on an external MCP server being connected, and degrades to "local artifact only" when it isn't — deliberately, since a wiki is a genuinely external system, not something a plugin should vendor a copy of.
If you're building something similar — or just tired of watching an agent
confidently fix the wrong thing — the repo's up, MIT licensed, and the
pipeline is designed to be read, not just run:
github.com/MallikarjunHt/agentic-sdd.
- No automated validation of the plugin's own prompts/templates — there's no compiler for a markdown instruction file. A change is proven by running it against a real feature, not by a test suite.
- No cross-repo coordination — each install targets one repo via its own config. A feature spanning two repos needs two separate, manually coordinated runs.
- The wiki-mirror feature (a page-per-stage sync to Confluence or similar) is the one remaining capability that depends on an external MCP server being connected, and degrades to "local artifact only" when it isn't — deliberately, since a wiki is a genuinely external system, not something a plugin should vendor a copy of.
If you're building something similar — or just tired of watching an agent
confidently fix the wrong thing — the repo's up, MIT licensed, and the
pipeline is designed to be read, not just run:
github.com/MallikarjunHt/agentic-sdd.
