Spec-driven development vs vibe coding is usually pitched as a rivalry, as if a team has to pick a side. That framing is wrong. The two optimize for different things, they fail in different ways, and most teams that build with AI coding agents end up needing both — on different pieces of work, often in the same week.
The real question is not which one is better. It is which one fits the thing in front of you right now, and what happens when the prototype you vibe-coded on Tuesday is quietly running in production by Friday.
Key Takeaways
- Spec-driven development vs vibe coding is a difference in what you optimize for: vibe coding optimizes for discovery speed, spec-driven development for code that survives
- Vibe coding — a term Andrej Karpathy coined in 2025 — means describing intent in natural language and accepting the agent's output with minimal review
- Spec-driven development writes a reviewed specification first and treats it as the source of truth the code is generated from
- The measurable cost of unreviewed AI code is real: GitClear found refactoring fell from 25% of changed lines in 2021 to under 10% in 2024, while copy-pasted code rose
- Teams don't choose one — they vibe-code to discover, then convert what survives into a spec before it ships. The hard part is the handoff between the two
Spec-driven development vs vibe coding: the core difference
Spec-driven development vs vibe coding comes down to what you optimize for. Vibe coding describes intent in natural language and lets the agent generate code with minimal review — it optimizes for speed of discovery. Spec-driven development writes a reviewed specification first and treats it as the source of truth the implementation is derived from — it optimizes for code that other people, and other agents, can safely build on.
Neither is a best practice on its own. Vibe coding on a throwaway prototype is exactly right; spec-driven development there would be ceremony. Spec-driven development on a payments change three services deep is exactly right; vibe coding there is how you ship a contradiction with passing tests. The table below is the honest version — strengths on both sides.
| Dimension | Vibe coding | Spec-driven development |
|---|---|---|
| Optimizes for | Discovery speed | Durability and maintainability |
| Starting artifact | A prompt | A reviewed specification |
| Source of truth | The running output | The spec |
| Human review | After the fact, if at all | Before any code is written |
| Best for | Prototypes, spikes, disposable code | Production systems, shared codebases |
| Failure mode | Technical debt compounds silently | Overhead when the work is throwaway |
| Scales with team size | Poorly — intent lives in one head | Well — intent is written down and reviewed |
| Agent handoff | A prompt the next person has to reverse-engineer | A plan the next agent can read directly |
The row that matters most for a team is the second-to-last one. Vibe coding keeps the intent in the head of whoever typed the prompt. That is fine when that person is the only one who will ever touch the code. It stops being fine the moment a second person — or a second agent — has to understand why the code does what it does.
What vibe coding actually is
Vibe coding is an AI-assisted style where you describe what you want in plain language, accept the generated code largely on faith, and keep prompting until the software appears to work. Karpathy's original framing was that you "fully give in to the vibes" and "forget that the code even exists." The term caught on fast enough that Collins named it a Word of the Year for 2025.
The honest case for vibe coding is strong. For a prototype you intend to throw away, reading every line is wasted effort — the output is the spec, and if it runs, it has done its job. For solo exploration, a spike to learn whether an approach is even viable, or a one-off script, vibe coding is often the correct and fastest choice. It lowers the cost of trying something to almost nothing, and that is genuinely valuable.
The problem is not vibe coding. The problem is vibe-coded software that was supposed to be disposable and quietly wasn't.
What spec-driven development actually is
Spec-driven development inverts the order: you write the specification first, review it, and treat it as the source of truth the code is generated from. The spec — the problem, the user stories, the architecture, the interfaces — is the artifact people reason about and approve. The implementation follows from it rather than the other way around. GitHub's open-source Spec Kit frames the goal as making specifications executable: the spec generates the implementation rather than just loosely guiding it.
That sounds heavier than it is in practice. A spec does not have to be a fifty-page document; it has to be the set of decisions that would be expensive to get wrong, written down where someone other than the author can check them. The point is not paperwork. The point is that the decisions exist somewhere reviewable before an agent turns them into code — which is the whole argument of what spec-driven development for teams actually requires.
Spec-driven development's cost is real too: for genuinely disposable work, writing a spec first is overhead you will never recoup. The skill is knowing which bucket the work is in.
Where vibe coding wins, and where it quietly breaks
Vibe coding wins whenever the code is disposable and loses whenever it isn't. The trouble is that "disposable" is a decision made at the moment of writing, and production is a decision made later — often by someone else, often under deadline, often by promoting the prototype that already works rather than rebuilding it. The code outlives the assumption that justified skipping the spec.
This is where the cost stops being hypothetical. GitClear's analysis of 211 million lines of code found that as AI coding tools were adopted, refactoring sank from 25% of changed lines in 2021 to under 10% in 2024, and 2024 was the first year on record where developers introduced more duplicated code than they consolidated. Copy-pasted lines rose from 8.3% to 12.3% over the same period. That is the signature of shipping faster than anyone reviews — exactly what unbounded vibe coding produces at scale.
The failure mode is the diagonal path: code that took the "yes, disposable" branch and then survived anyway, without ever passing through a spec. On one person's side project that is a cleanup chore. On a team's shared codebase it is debt that compounds in the dark.
Why teams need the difference, not a side
For a solo developer, "spec-driven development vs vibe coding" is a personal workflow preference. For a team it is a coordination problem, and that changes the math. The reason has nothing to do with discipline and everything to do with where knowledge lives.
When you vibe-code, the intent — why this edge case matters, which service this must not touch, what the agent was actually asked for — lives in your head and in a prompt history nobody else reads. A teammate who picks up the code has to reverse-engineer the intent from the implementation. So does the next AI agent, which will confidently extend code it does not understand the constraints on. A spec is simply that intent, written down once, where the next human and the next agent can both read it. The cost of writing it is paid once; the cost of not writing it is paid every time someone else touches the code.
This gets sharper across repositories. A vibe-coded change that looks self-contained in one service can quietly assume something about a second service in a different repo — and nothing in a prompt-and-iterate loop surfaces that. A reviewed spec with real impact analysis across every repo a change touches catches the cross-boundary assumption before an agent builds on it. If your work rarely leaves one repository, this is academic. If you run microservices, it is the difference between catching a dependency in review and finding it in production — and you can put a number on what those misses cost with the rework ROI calculator.
So the team version is not "always spec." It is: vibe-code the disposable, spec the durable, and make the handoff between them deliberate rather than accidental.
How to choose, per piece of work
Choose per piece of work, not per team and not per quarter. The deciding question is the one in the diagram above: is this code disposable? If you will delete it regardless of whether it works, vibe-code it and move fast. If it will live in a shared codebase, be maintained by someone else, or be extended by an autonomous agent, it needs a spec before it ships — even if you vibe-coded your way to understanding the problem first.
The practical synthesis most mature teams land on is structured exploration with living specs: vibe-code to discover what you actually need, then formalize the parts that survived into a reviewed specification before they reach production. The exploration stays fast and cheap. The durable code still gets the review it needs. What makes or breaks this pattern is the handoff — the moment a prototype is promoted to "real." Teams that leave that moment implicit ship their prototypes by accident. Teams that make it an explicit gate get the speed of vibe coding and the durability of spec-driven development without pretending one of them is always the answer.
That gate — a reviewed, approved spec standing between exploration and production — is the thing worth building deliberately. It is where the discovery speed of vibe coding gets converted into something a team, and an agent, can safely build on.
Build the gate between exploration and production
The enterprise AI-SDLC playbook lays out how teams keep vibe coding's speed while giving durable work a reviewed, approved spec before it ships.
Download the playbook →