Table of contents
Spec-Driven Development is the hottest acronym of 2026. Conferences, LinkedIn threads, opinion pieces: everyone seems to have a take. Strip away the marketing and you find three radically different practices wearing the same label, with promises, costs and risks that have almost nothing in common.
Birgitta Bockeler at Thoughtworks recently mapped the field across three distinct flavors. This article summarizes what she describes, adds six months of our own experience shipping client code with these approaches, and lands on what genuinely pays off, what is over-engineering, and how to add just enough into Claude Code without falling into spec-mania.
If you’re evaluating Spec-Kit, Kiro or a plain SPEC.md for your team, here’s what to keep in mind before deciding.
Spec-first, spec-anchored, spec-as-source: three flavors explained
Bockeler’s distinction is useful because it points at what genuinely changes from one tool to another: not the format of the spec, but its role across the code lifecycle.
Spec-first: the spec is the input. You write a detailed intent (a PRD, a fleshed-out user story, a lightweight requirements doc), the AI agent uses it to plan, then to generate code. Once the code ships, the spec becomes a reference artifact, often abandoned in some corner of the repo. That’s GitHub Spec-Kit’s bet. Upside: framed kickoff. Downside: the spec drifts from the shipped code fast.
Spec-anchored: the spec is a living contract. It’s updated as the code evolves, and the agent reads it on every session to stay aligned. That’s Amazon Kiro’s pitch with the steering folder + spec structure. Upside: less drift, the agent retains context. Cost: someone has to maintain the spec actively, which adds non-trivial cognitive load.
Spec-as-source: the spec is the source of truth, code becomes a compiled artifact regenerated from it. Tessl Framework bets the farm on this one. Theoretical upside: zero drift, perfect traceability. Reality: you inherit every pathology of 2000s Model-Driven Development, plus the uncertainty layer of LLMs.
The arXiv paper Spec-Driven Development: From Code to Contract in the Age of AI (February 2026) frames these three levels as a spectrum, not as boxes. Useful read for understanding where each team can sit without locking into a single dogma.

GitHub Spec-Kit: orchestrated spec-first
GitHub Spec-Kit is the most visible tool right now because it ships behind GitHub, hence by default on millions of repos via Copilot. Its model: a three-step workflow (specify, plan, tasks), a constitution that defines immutable principles the agent must respect, and configurable prompts at each step.
In practice, you trigger the command, the agent walks you through clarifying the intent, then generates a technical plan, then breaks it into tasks. You validate at each transition. The result: a feature that starts with a clear frame, and a documented decision history.
Where Spec-Kit shines: kicking off a feature from zero with a team that doesn’t know the domain, or an agent that needs to run unattended for hours. Where it falls short: maintenance. Once the feature ships, who updates the spec? Honest answer: almost nobody. The spec files become fossils in the repo, and new features start fresh, ignoring the history.

Amazon Kiro: the spec as a living contract
Kiro (Amazon) takes the opposite angle. The spec is not a kickoff deliverable, it’s a permanent reference. The steering folder centralizes project conventions (code style, architectural patterns, domain vocabulary), and each feature carries its own spec that stays in sync with the code as things change.
Kiro’s bet: if the spec is always current, the AI agent doesn’t have to rediscover context every session. It reads the steering folder, opens the feature spec, and works with as fine-grained an understanding as a human who’s been on the project from day one.
In practice, this works when the team buys into the update discipline. The moment a developer ships a quick fix without touching the spec, the contract breaks, and the next agent that opens the feature gets stale information. Kiro is powerful for regulated codebases (healthcare, finance, defense) where traceability is mandatory anyway. For an SMB pivoting every quarter, the maintenance can become a bottleneck.

From Assess to harness engineering
On the Thoughtworks side, the Technology Radar Vol. 33 from November 2025 placed SDD in Assess: worth watching, not ready for general adoption. The verdict was cautious: interesting, but too many unknowns around maintenance and cognitive cost.
Six months later, the April 2026 Radar (Vol. 34) folded SDD into a broader concept: harness engineering. The idea: stop seeing SDD as an isolated practice and start seeing it as one lever (alongside Agent Skills, mutation testing, context engineering) for building scaffolding around the agent that forces it to ship reliable code without constant supervision.
The shift matters: we’re moving from how to write the right spec to how to wire the agent so it fails visibly and self-corrects. It’s the same logic as AI agents calling APIs instead of clicking buttons: the value isn’t in the interface or the spec itself, it’s in the harness that makes them reliable. 🔧
Honest verdict: where SDD actually pays off
After six months using these approaches on client code (SaaS migrations, e-commerce integrations, internal tooling), our read is clear:
Real wins:
- Multi-file features that touch several layers (model, controller, view, tests). The spec forces you to think the contract before diving into code.
- Long autonomous agent runs (overnight runs, migration batches). Without a spec, the agent drifts.
- Audited codebases where every decision must be traceable. The spec becomes proof of intent.
- Onboarding a new developer or a new agent. The spec is the summary documentation never delivers.
Pure ceremony:
- Ten-line bug fix: writing a spec for that doubles the fix time.
- Throwaway spike to test a hypothesis: the spec isn’t the artifact, the code itself is.
- Domain the team already knows cold: the test, review, merge loop is faster than spec, review, code, merge.
- Short feedback loop: if your tests run in two minutes, the spec slows you down.
Bockeler’s matrix nails it: SDD adds value when the problem is complex and clear. When it’s small or fuzzy, it’s friction.

Practical advice for Claude Code
For Claude Code specifically, skip the seventeen-file frameworks. One lightweight SPEC.md at the root of the feature folder is enough, with four sections:
- Goal: one sentence describing what the feature does for the user.
- Constraints: what you can’t break (compatibility, performance, security).
- Acceptance criteria: 3-5 concrete, observable, testable cases.
- Out-of-scope: what’s not in this iteration, and why.
The agent reads it on every run, you edit it together as the scope evolves, it lives with the code in the same PR. No external tooling, no constitution, no steering folder. Just one file that forces clarity of intent before the first line of code.
Frequently asked questions
Should we adopt Spec-Kit, Kiro or stick with a lightweight SPEC.md ?
It depends on context. Spec-Kit for kicking off a feature from zero with a team or agent that doesn’t know the domain. Kiro for codebases where traceability is mandatory (regulated, audited). A simple SPEC.md for most SMBs and agile projects. Start with the lightest, scale up only if you genuinely feel the need.
Does Spec-Driven Development slow down delivery ?
On small features, yes. On complex or multi-file features, it speeds delivery up because it removes the back-and-forth with the agent. Measure across three sprints before judging.
Is it compatible with Claude Code ?
Perfectly. Claude Code reads any markdown file in the repo. A SPEC.md at the root of the feature folder is all it needs to stay aligned on intent.
Spec-Driven Development done right is documenting just enough to give the agent a target, without turning every feature into a spec project. The discipline isn’t in the format, it’s in the daily call to stop writing the moment the spec has done its job.
So: are you running full Spec-Kit, anchoring with Kiro, or living happily with a single SPEC.md ?
Related reading: our previous analysis on the end of the interface and the rise of AI agents, the broader context where SDD fits. And our AI and automation services if you want to integrate these approaches into your stack.
Sources: Birgitta Bockeler, SDD: 3 Tools Compared · Martin Fowler’s take · arXiv paper Spec-Driven Development: From Code to Contract in the Age of AI (Feb. 2026) · Thoughtworks Technology Radar Vol. 33 and April 2026.
Exploring AI in your dev stack ?
We help SMB teams adopt Claude Code, GitHub Spec-Kit or Kiro without falling into the tooling-for-tooling trap. 30 minutes to clarify what’s actually worth it in your context.
