---
title: "I Got Tired of Vibe Coding, So I Rebuilt the SDLC as Twelve Agent Skills"
slug: i-got-tired-of-vibe-coding-so-i-rebuilt-the-sdlc-as-twelve-agent-skills
url: https://listedarticles.com/articles/i-got-tired-of-vibe-coding-so-i-rebuilt-the-sdlc-as-twelve-agent-skills
canonical_url: https://medium.com/@zeeshan_hanif/i-got-tired-of-vibe-coding-so-i-rebuilt-the-sdlc-as-twelve-agent-skills-229773eafb07
content_type: tutorial
language: en
published_at: 2026-09-23T16:06:47.000Z
updated_at: 2026-09-23T18:32:19.201Z
author: "Zeeshan Hanif"
author_url: https://medium.com/@zeeshan_hanif
authored_by: human
publisher: "Medium"
publisher_url: https://medium.com/
topics: ["AI Agents", "Programming", "Software Engineering", "Open Source", "Developer Tools"]
license: all-rights-reserved
word_count: 4397
reading_minutes: 19
citation: "Zeeshan Hanif, Medium. \"I Got Tired of Vibe Coding, So I Rebuilt the SDLC as Twelve Agent Skills.\" 23 Sept 2026. https://medium.com/@zeeshan_hanif/i-got-tired-of-vibe-coding-so-i-rebuilt-the-sdlc-as-twelve-agent-skills-229773eafb07 (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# I Got Tired of Vibe Coding, So I Rebuilt the SDLC as Twelve Agent Skills

> Zeeshan Hanif open-sources a twelve-skill agent kit that turns vibe coding into a disciplined pipeline—from requirements interview to deployed, verified software with decisions traced on disk.

# I Got Tired of Vibe Coding, So I Rebuilt the SDLC as Twelve Agent Skills

## The problem nobody wants to admit

Coding agents are extraordinarily good at writing code. They are bad at knowing what to write, in what order, and whether it’s actually done.

Left to itself, an agent improvises. It redesigns mid-build. It weakens a failing test until it passes. It loses its place between sessions. And it ships something that satisfies the last message in the chat rather than the requirement you started with.

Here’s how most AI-assisted projects actually get built today. You open a chat, describe an app in three sentences, and the agent starts generating. It makes a hundred small decisions — database shape, auth approach, folder structure, error handling — and tells you about none of them. Twenty minutes later you have something that runs. Two weeks later you have something nobody understands, including the agent, because the context that produced those decisions evaporated the moment the session ended.

For a weekend prototype, that’s genuinely fine. But the moment a project has to outlive the chat window that created it, the cracks show:

There are no requirements — anywhere. Ask “does this handle password reset?” and the honest answer is “let me grep the code and find out.” The requirements exist only as a fossil record of prompts scattered across chat histories.

Decisions have no memory. Why Postgres and not Mongo? The agent had a reason at 2:14 PM on a Tuesday, and it’s gone. A later session either rediscovers the reasoning or silently contradicts it.

Context rot compounds. Long agent sessions degrade. So you start a fresh one — which knows nothing, because everything it needed lived in the previous session’s context window instead of on disk.

Nothing verifies anything. The agent says “done, all tests pass.” Did it write tests that check the acceptance criteria, or tests that describe whatever the code happens to do? Without a requirement to trace back to, “the tests pass” is a statement about the tests, not the software.

None of this is new. Our industry spent fifty years building answers to exactly these problems: requirements specifications, architecture decision records, traceability matrices, independent acceptance verification, staged delivery. Most of that discipline was already eroding before agents arrived — the paperwork was expensive, and teams under deadline pressure dropped it first. Agents accelerated that retreat. When code arrives in seconds, the artifacts around it — the spec, the decision record, the traceability — become the slow part, and the part that gets skipped.

That’s the insight I couldn’t shake:

>

The traditional SDLC was never wrong — it was just too expensive for humans to run rigorously. Agents make it cheap.

An agent doesn’t groan at maintaining a traceability matrix. It doesn’t skip the architecture doc because of sprint pressure. The clerical burden that made heavyweight processes collapse under their own weight is precisely what LLMs are best at.

So I built it: Agentic SDLC Kit — twelve Agent Skills that carry a project from “I have an idea” to a deployed, verified, maintainable system. They follow the open SKILL.md standard, so they run in Claude Code, Codex, and any other agent that loads skills.Press enter or click to view image in full size

## One requirement, end to end

Before the machinery, the thing the machinery exists for. Say the requirements interview captures a sentence like this:

>

Here is what that one sentence becomes as it moves through the pipeline, and the file each piece lands in:Press enter or click to view image in full size

The trace shows eight of the twelve skills, because those are the ones that produce something for this requirement. Scaffolding and deployment work on the project as a whole, and the orchestrator and seam checker sit above every requirement rather than inside one.

Nothing in that chain is remembered. Every step is a document read from disk by a skill that may be running in a session opened ten minutes or ten days later. FR-AUTH-001 means the same thing at every step, and the traceability matrix says so:Press enter or click to view image in full size

Each column is written by a different skill — more on that shortly.

## Running the pipeline

The README walks through a complete run, with the exact prompt that triggers each skill, in order. The first one is all you need to start. Open your agent on a new project and type:

>

I’m kicking off a new project — help me gather and write up the requirements.

Then see how different it feels when the agent interviews you before writing anything.

Every step after that is the same kind of sentence. Once the plan exists, the whole construction loop is four of them, repeated for each feature:

```
Design the next feature from the plan.
Design the next feature's screens.
Implement the next feature from the plan.
Verify the next feature.
```

Or hand the loop to the orchestrator with one: “Run the loop until it hits something that needs me.”

## What makes this a kit and not twelve prompts

Six contracts make the skills behave like one system. If you take nothing else from this article, take these — they apply to any multi-step agent system you build.

1. Stable IDs, minted once, never recycled.FR-AUTH-007 in the SRS is the same ID the architecture cites, the plan schedules into a feature, the design turns into acceptance criteria, and the auditor signs off on. Six namespaces — requirements, use cases, decisions, screens, features, defects — all immutable. Removals are tombstoned, never deleted or renumbered, so every downstream reference stays resolvable forever.

2. Source-gated traceability. A shared vocabulary is also a way to manufacture authority. A fabricated FR-AUTH-019 looks exactly like a real one — it reads as grounded, it survives review, and three documents later something depends on a requirement that was never written. A missing citation is visible; an invented one isn't, which makes it the more corrosive of the two failures by a wide margin. So the rule is blunt: a skill cites an ID only when the document that defines it is present. Where there's nothing to cite, a skill states its reasoning in prose and says so. That's prevention at write time; skill twelve catches at audit time whatever still slips through.

3. The traceability matrix is a shared ledger with exclusive column ownership. Multiple skills write to docs/rtm.md, but each owns exactly one column and appends, never overwrites. Requirements owns the rows, planning owns Plan ref, architecture opens Design ref for the two design skills to extend, and acceptance-verification alone owns Test ref. A requirement's lifecycle reads left to right, and "fully verified" is computed from the intersection — never written into a cell by anyone.

4. Position is computed from artifacts, never remembered. No skill keeps a pointer to “where we are.” A feature folder with a technical design is designed; add a UI design and it’s ui-designed; a fully-checked task list is developer-done; an accepted acceptance report is verified. That’s why a brand-new session lands in exactly the right place. Context rot stops mattering when context is reconstructed from disk.

5. Escalate, don’t fork. When reality contradicts the design, a skill either implements the design’s intent or files an amendment against the owning document. It never quietly changes the spec. A feature needing a new entity is an architecture amendment. A screen needing an off-system color is a design-system amendment. No silent redesigns.

6. Green is demonstrated, not asserted. Done-whens are executed. Suites are re-run cold by an independent auditor. Tests and coverage gates are never loosened to pass.

Two chains run the length of the pipeline, because they’re where agentic projects usually rot. Testing is decided once, at the top — the architecture picks the strategy, the frameworks, and the coverage stance, and everything downstream realizes that decision instead of re-deciding it. Configuration is a first-class artifact — scaffolding mints a config template per deployable unit, implementation extends it whenever a task needs a new variable, and deployment reconciles the templates against the real environment before deploying. Green pipelines stop dying on a missing environment variable.

All six contracts rest on one fact: everything the pipeline knows is in the repo. Here’s everything it writes, and which skill writes each file.Press enter or click to view image in full size

## The linear phase — five skills that run once per project

### 1. requirements-engineering — the interview that becomes an SRS

The pipeline’s entry point. Instead of transcribing what you happen to mention, it proactively enumerates the standard sub-requirements for each capability area and has you confirm, extend, or trim them. “User authentication” is never one line — it expands into sign-up, sign-in, email verification, forgot/reset password, logout, session expiry, lockout, rate limiting. A requirement catalog prompts the commonly-forgotten areas too, and it walks the standard quality categories — performance, security, reliability and the rest — so non-functional requirements come out measurable.

It produces a structured SRS, a use-case document, and the traceability matrix. Functional requirements use either EARS — five constrained sentence patterns — or classic “shall” statements; you choose once, before the first requirement is minted, and that choice binds every later amendment. Every run includes an unwanted-behavior pass, so error and failure counterparts become their own requirements instead of afterthoughts in prose. Amendments later add, modify, or tombstone requirements without disturbing the IDs already in flight.

The benefit: “what did we agree to build?” has a canonical, citable answer before a line of code exists — and it can change later without destroying history.

### 2. software-architecture — decisions with receipts

Its governing principle: architecture is driven by quality attributes and constraints, not by technology. It elicits what actually forces decisions — scale, availability, consistency, security, team, timeline, lock-in tolerance — and derives the design, rather than reaching for a familiar stack first. With an SRS present it reads the requirements as the primary source and interviews only for the gaps; with use cases present it mines them for the runtime view, since exception flows reveal failure handling and actors reveal trust boundaries.

Output: an arc42-based document right-sized to the system, C4 diagrams as embedded Mermaid, and Architecture Decision Records with a “Requirements addressed” field citing the exact IDs that drove each decision. Plus the testing decisions above, and a deployment view recording the environment set and promotion order — the contract the deployment skill later executes.

The benefit: “why Postgres?” now has a written answer with a requirement ID attached, and ADR supersession preserves the decision history instead of overwriting it.

### 3. ux-foundations — the architecture of the UI

The design-phase sibling to software-architecture. It derives the structure and design language of the UI rather than jumping to pixel-perfect screens, on the principle of one shared core plus a per-surface layer — so an admin portal, marketing site, and mobile app feel like one product without flattening their real differences. Personas come straight from the SRS’s user classes, and the SRS’s accessibility requirements are a non-negotiable bar: an imported palette that fails the contrast requirement gets adjusted, and the change is recorded.

The visual direction comes from one of four source modes:
- Research — derive a new direction from scratch.
- Extract — read it out of reference images you drop in a folder.
- Ingest — take it from an existing design file or brand book.
- Connect — pull it over MCP from a design tool such as Claude Design or Figma.

It always asks; detection tells it what’s possible, only you say what’s wanted. And every mode ends by playing back a proposed direction for confirmation, because extraction and research are approximations, never facts.

The fourth mode deserves a note, because it comes with an ordering constraint. By the time you reach this step you hold an SRS and an architecture document — a written account of who the users are, which surfaces exist, and what the accessibility bar is. Hand those documents, whole or in the relevant parts, to a design tool like Claude Design or Figma, generate a real visual design from them, and connect it over MCP. What ux-foundations pulls across is not screens but the language underneath them — tokens, styles, components.

The mode you choose also fixes the foundation’s provenance, and that is what the screen-level skill inherits two steps later: whether screens are registered from the design tool or generated fresh, and from where. So design-tool work belongs before ux-foundations, not after it. Generate a design once the foundation has already been set from a different source, and you own two design systems that don’t agree.

Three outputs with a strict authority split: ux-foundations.md (personas, IA, flows, screen inventory with stable SCR IDs), design.md (the render-time system a coding agent loads when building UI), and tokens.json (canonical W3C DTCG tokens — the single source from which design.md's CSS is derived).

The benefit: downstream agents don’t improvise UI. Screens already have IDs, tokens already have values, components already have specified states.

### 4. implementation-planning — from documents to build order

The bridge from design to construction. Two principles: slice vertically, never horizontally — every unit of work cuts through UI, API, domain, and data to deliver something demonstrable, never “build all the tables” — and make the architecture executable before complete: build the walking skeleton first, not the easiest feature.

It produces epics and thin vertical slices, each with a stable FEAT-NNN ID and each tracing the requirements it implements, the use cases it realizes, and the screens it touches. Plus a risk- and dependency-ordered sequence, the first slice specified in full, and a Must-requirement coverage check — every Must-priority requirement lands in a slice or is consciously deferred. No silent gaps. It plans the whole app's breadth but details only what's next: depth-on-demand, not the waterfall trap.

The benefit: “what do we build first, and why?” gets a defensible answer, and the plan doubles as the work queue every loop skill reads.

### 5. project-scaffolding — the first skill whose output is a running system

Scaffolding usually means “generate some folders.” This one runs the ecosystem’s official generators — discovering them and verifying their names and flags against live docs before first use, never from memory, which is what keeps the skill self-updating as frameworks change. The stack itself is an input, never a decision: the ADRs already chose it.

You get a real repo with module boundaries enforced by lint and import rules, a wired walking skeleton (UI shell → API → domain stub → database → back), a compose-first local stack with pinned images so a teammate or a fresh agent session starts the same system you did, CI, config templates, and a test harness running the frameworks the architecture named. Agent instructions live in AGENTS.md at the root, with CLAUDE.md holding nothing but a pointer to it — one source of truth instead of two that drift.

Then it verifies itself empirically: clean install, every unit builds, the skeleton test passes from a cold start, boundary rules hold, and every path in AGENTS.md resolves. Anything unfixable is flagged, never silently shipped.

The benefit: construction starts from a repo that provably works, with the project’s conventions physically installed where fresh sessions will find them.

## 6. initial-deployment — runs once, whenever you choose

Scaffolding stops at deploy-ready by contract. This skill executes exactly those artifacts and gets the system running in the cloud. It runs once, and you choose when: on the bare skeleton right here, mid-loop after a few features, or once the whole plan is built. It gates nothing downstream.

Doing it early is what I’d recommend — deploying the walking skeleton proves the system deploys before features pile on. Deploying later gives its measurement phase more to measure, at the cost of a first push that ships N features at once, with many more candidate causes when something breaks. Either way, once it has run, CD carries every push after it.

Two gates define its character:
- Money. The deployment plan — resources, environments, rough cost class — is played back for one explicit confirmation. Nothing billable exists before that nod.
- Credentials. The skill preflights that your provider CLI is authenticated and blocks with instructions if it isn’t. It never asks for, stores, or writes a secret value — it wires the provider’s own secret mechanism and proves it with a non-secret canary.

It provisions the infrastructure in dependency order, extends CI into CD, and runs the skeleton’s own test against the deployed environment to confirm it actually works.

The benefit: the gap between “works on my machine” and “runs in production, verifiably” becomes a skill instead of a weekend of improvisation.

## The Construction Loop — four skills, once per feature, forever

Here the pipeline changes species. Everything so far runs once per project; these four run once per feature, in a fresh context each time, reading everything they need from disk.Press enter or click to view image in full size

### 7. detailed-design — one feature, fully specified

Turns planned intent into a buildable technical design. Three principles: the what is fixed upstream and this skill designs the how; the live codebase is a mandatory input, not an obstacle — feature twelve is designed against code that has evolved past the skeleton, and where code and documents diverge, reality wins (with the divergence noted); and depth here and only here.

Per feature it writes technical-design.md — API contracts, schema migrations within the entities architecture already owns, acceptance criteria derived from the requirement statements — and tasks.md, an ordered, individually verifiable program. Each task's done-when is classified at design time: behavioral tasks get test-artifact done-whens, structural tasks get demonstration done-whens, so implementation executes the classification instead of improvising ceremony tests. New entities or boundary changes escalate to an architecture amendment, never invented locally.

The benefit: the implementation agent receives a program, not a vibe.

### 8. ui-design — the adapter between design tools and code

The presentation half of each feature’s design, and structurally an adapter. How screens actually get designed varies wildly, so it routes across strategies per screen — following the source the foundation established — while emitting one uniform contract: docs/design-manifest.json, the screen registry that everything downstream reads regardless of what produced each screen. All design-tool variance is absorbed here; no downstream skill ever learns what a Figma node is.

It has three modes:
- Anchor — right after ux-foundations, design two or three compositionally demanding screens to prove the design system actually composes before features build on it.
- Per-feature — the loop step.
- Re-verification — after a design-system amendment, re-check the affected screens.

Screens conform to the design system or escalate. An off-system color is either corrected or the system is amended upstream — never forked locally.

The benefit: the pipeline works whether you have a designer, a Figma file, or nothing at all, and implementation never has to guess which design is authoritative.

### 9. feature-implementation — the disciplined builder

tasks.md is the program; this skill is the interpreter. It executes the feature's tasks autonomously, one at a time in order, against the live codebase, ending developer-done. It carries seven disciplines, each targeting a named failure mode of agentic implementation:
- The program is fixed. Improvisation is escalation, not initiative.
- Disk is the memory. Checkbox state is the position, and every iteration must be executable in a brand-new session.
- Fix-loops are bounded. Three attempts per task, each needing a changed hypothesis, then an honest stop — unchecked box, failure note, an explicit WIP commit a fresh session can pick up cold.
- Scope is this feature. No drive-by refactors.
- Conventions hold by construction. Frameworks inherited from the harness; UI built through tokens rather than raw values.
- Git is the checkpoint. One commit per completed task, so the log reads as the execution log of the task list.Press enter or click to view image in full size

And the seventh, my favorite — the anti-fake-green rule. Tests are never weakened, skipped, deleted, or edited to pass, and where coverage is enforced, the threshold, its scope, and its exclusion lists are equally off-limits. A red gate means missing tests, not a config to lower. Fixing a genuinely buggy test toward the design is legitimate and recorded; fixing it toward the code is the forbidden move in costume.

Repeated blocking on one task is treated as a meta-signal that the design has a systematic problem — it recommends going back to detailed-design rather than burning more attempts.

The benefit: the speed of autonomous implementation, with guardrails against the specific ways autonomous implementation cheats.

### 10. acceptance-verification — the independent auditor

The skill that makes the whole traceability story pay off. It runs after implementation and answers, without the implementer’s investment: does this feature actually satisfy its requirements?

Its central rule is derivation, not trust. Every check is re-derived from the authoritative documents — acceptance criteria cross-checked against the verbatim SRS and use-case statements — never from task checkboxes, delivery summaries, or any session’s claimed greens. Those are the artifacts under audit, not evidence. Claims are exhibits; observations are evidence.

So it audits each criterion for coverage and faithfulness, checks that the test actually exercises the behavior, runs an independent anti-fake-green review over the feature’s commit range, and re-executes every suite cold — feature tests, whole-repo, E2E, the coverage gate exactly as CI runs it. It verifies side-effect requirements by observation, inspecting the audit log’s entries rather than trusting a 200.Press enter or click to view image in full size

Auditor and mechanic stay separate: findings route, and the skill never fixes production code. The one thing it may change is the measurement — a weak test corrected toward the criterion, never toward the code. If the faithful test then fails, that’s a rework finding: the system working as designed. Three verdicts come out of it:
- Accepted — every criterion traced to a passing, faithful test.
- Rework — the implementation doesn’t satisfy a criterion; routes back to the builder.
- Design defect — the criterion itself is wrong or unbuildable; routes back upstream.

A feature is verified whole or not yet. Partial acceptance isn’t a verdict.

The benefit: “done” stops being the builder’s opinion and becomes an auditable fact.

## Above and across — the two meta skills

### 11. sdlc-orchestrator — the thin loop driver

With eleven skills in play, something has to answer “what runs next?” This one does — and it’s deliberately thin. Every loop skill already resolves its own position; the orchestrator adds only the global position, the invocation, and the routing. If logic here starts to look like design or build logic, it belongs in a stage skill instead.

It computes rather than stores — there is no orchestrator state file — and drives the loop at whatever scope you ask for: one stage, one feature cycle, or run-until-blocked. It announces its resolution before acting (“FEAT-006 is at ui-designed; invoking feature-implementation”) rather than presenting a menu, which would invite off-sequence violations of the plan’s dependency order.

It routes outcomes: rework loops back through implementation and re-verification, bounded at two rework cycles before surfacing — an auditor and a builder disagreeing repeatedly is a design-quality signal, not something to grind through. Change requests walk the amendment chain top-down so a feature is never minted without its requirements. Bugs become a defect-ledger entry, a failing test first, a scoped fix, then re-verification.

And the design constraint I’m proudest of — its absence changes nothing. Every skill stays independently invocable. The orchestrator is composition on top, not wiring inside.

### 12. pipeline-verify — the seam checker

Tier one of verification is inside every stage skill; each verifies its own contract at delivery. This is tier two: the seams between documents that no single stage can see — a citation written correctly against a document another skill later amended, a requirement every stage individually handled but no stage ever covered, a manifest locator that rotted when a heading moved.

It’s mechanical only, read-only, and derived-output-only. It checks that every cited ID resolves; that no active work cites a tombstoned item; that there are no orphans, meaning requirements nothing touches and screens nothing renders; that the RTM’s rows, column ownership, and computed verification hold; and that the architecture’s named frameworks and coverage stance are actually realized in CI.

Findings come out classified as error, warning, or info, grouped by owning skill and specific enough to fix without re-deriving. Missing documents scope their checks out and are reported as skipped: absence of evidence is never reported as cleanliness.

The benefit: document rot — the thing that kills every documentation-heavy process — gets caught mechanically instead of discovered six weeks later.

## Do you need all twelve?

No. The pipeline is built for a system you intend to keep, but every stage right-sizes its own output — a CRUD tool gets a tight architecture document, not the full arc42 treatment. A reasonable minimum: requirements → architecture → plan → scaffolding, then the loop. Add the two design skills when there’s meaningful UI, deployment when it needs to be reachable, the orchestrator when you’re tired of driving by hand, and the seam checker once there’s enough traceability to be worth auditing. Run only requirements-engineering and stop, if that's the part that helps.

## What actually changes

Put together, the kit inverts the default AI-development experience on every axis that was broken.Press enter or click to view image in full size

The deeper point is the one I started with: this isn’t nostalgia for heavyweight process. It’s arbitrage. The traditional SDLC encoded decades of hard-won lessons about how software projects fail, and we priced it out of reach by making humans do the clerical work. Agents changed the price. The discipline is affordable again — and it turns out rigorous process is exactly what makes autonomous agents auditable, because every claim they make becomes checkable against a document they didn’t write.

## Get the kit

The kit is open source under Apache 2.0 at github.com/zeeshanhanif/agentic-sdlc-kit. The skills CLI installs it in any agent that loads skills — Claude Code, Codex, Cursor and others:

```
npx skills add https://github.com/zeeshanhanif/agentic-sdlc-kit --skill '*'
```

Name one skill instead of '*' to try it alone — for example, just the requirements interview:

```
npx skills add https://github.com/zeeshanhanif/agentic-sdlc-kit --skill requirements-engineering
```

In Claude Code, you can also install it as a single plugin:

```
/plugin marketplace add zeeshanhanif/agentic-sdlc-kit
/plugin install agentic-sdlc-kit@zeeshanhanif
```

The plugin bundles all twelve as one unit, because they share an ID vocabulary and a traceability matrix with per-column ownership.

I’d genuinely like to hear where it breaks for you. That’s how the next twelve improvements get found.

Some of the improvements are already underway: I’m now extending the kit with hooks and agents alongside the skills.

If you’re building agentic systems and this resonated, follow along — I’m writing more about loop engineering, document-driven agent pipelines, and turning SDLC discipline into services-as-software.
