SDD — QA, Testing & Best Practices¶
Why SDD Matters for QA¶
In spec-driven development the spec already contains what QA needs most: acceptance criteria with IDs. That gives QA three jobs:
- Test the spec — review requirements before any code exists (cheapest bugs to fix).
- Turn acceptance criteria into tests — so the spec is enforced, not just documented.
- Verify the code against the spec — every requirement implemented and tested, nothing outside scope changed.
"The spec document is the blueprint. The safety net is the test suite." — quoted in Martin Fowler's fragments
From Acceptance Criteria to Tests¶
Given/When/Then scenarios map almost directly to BDD scenarios or test names:
**FR-002**: System MUST reject duplicate album names
1. **Given** an album "Trip" exists, **When** the user creates another album "Trip",
**Then** they see "An album with this name already exists"
import pytest
@pytest.mark.requirement("FR-002")
def test_duplicate_album_name_is_rejected(api, album_factory):
album_factory(name="Trip")
response = api.post("/albums", json={"name": "Trip", "photo_ids": [1, 2]})
assert response.status_code == 409
assert response.json()["message"] == "An album with this name already exists"
Register the marker once so pytest doesn't warn:
[tool.pytest.ini_options]
markers = ["requirement(id): links a test to a spec requirement"]
EARS requirements map the same way — WHEN <event> THE SYSTEM SHALL <behaviour> becomes
arrange the condition → trigger the event → assert the behaviour. Kiro goes one step further and generates
property-based tests from EARS requirements (useful evidence, but "evidence rather than proof", and a poor fit
for external services or non-deterministic behaviour).
Traceability¶
Keep one chain from requirement to result:
FR-002 (spec.md) ──► T014 [US1] (tasks.md) ──► test_duplicate_album_name_is_rejected ──► CI result
- Requirements have IDs (
FR-001,SC-001), tasks have IDs and story tags (T014 [US1]) - Tests reference requirement IDs (marker, test name or docstring)
- A small script or report can list requirements without tests — that is your coverage gap
# scripts/spec_coverage.py — requirements from spec.md vs requirement markers in tests
import pathlib
import re
spec = pathlib.Path("specs/001-photo-albums/spec.md").read_text()
required = set(re.findall(r"\*\*(FR-\d{3})\*\*", spec))
tests = "\n".join(p.read_text() for p in pathlib.Path("tests").rglob("test_*.py"))
covered = set(re.findall(r'requirement\("(FR-\d{3})"\)', tests))
missing = sorted(required - covered)
print("Requirements without tests:", missing or "none")
raise SystemExit(1 if missing else 0)
Verifying Code Against the Spec¶
| Check | How |
|---|---|
| Artifacts are consistent | Spec Kit analyze — read-only check across spec, plan and tasks |
| Code matches the spec | Spec Kit converge (adds tasks for gaps); OpenSpec /opsx:verify |
| Independent review | A separate reviewer agent/sub-agent compares the diff with the spec — avoids the implementing agent grading its own work |
| Scope | Nothing outside the listed files/interfaces changed |
| Acceptance | All tests linked to requirements pass; success criteria measured |
Spec Kit's tasks template makes test tasks (contract, integration) come before implementation — but only
if tests are requested. Ask for them explicitly in the constitution: "every functional requirement has an automated test".
Spec Review as a QA Activity¶
Treat the spec like a test object:
- Ambiguity — could two people implement this differently? → ask, mark
[NEEDS CLARIFICATION] - Testability — can you write a pass/fail check? "Fast" is not testable; "under 2 seconds at p95" is
- Completeness — errors, empty states, permissions, limits, concurrency, negative cases
- Consistency — no contradictions between requirements, stories and success criteria
- Scope — non-goals are written down
Spec Kit's checklist command produces requirements-quality checklists (CHK001…). A checked box means the
requirement is good — not that the feature is implemented.
Best Practices¶
| Practice | Why |
|---|---|
| Scale the process to the change — skip the spec if the diff fits in one sentence | Avoids ceremony for small fixes |
| Many small feature specs, stored in the repo next to the code | One mega-spec is unreadable and goes stale |
| Separate what/why from how | Specs stay stable when technology changes |
| Review one step at a time — spec, then plan, then tasks | Mistakes are cheaper early |
| Implement in a fresh context | The agent works from the spec, not from a long chat |
| Give the agent a runnable check — tests, build, linters | "If you can't verify it, don't ship it" |
| Keep specs alive — update the spec when requirements change (Kiro Sync Files, OpenSpec archive) | Prevents spec drift |
Keep rule files short — CLAUDE.md, AGENTS.md, steering |
Long files get ignored |
Anti-Patterns¶
| Anti-pattern | Symptom | Fix |
|---|---|---|
| Waterfall spec | Everything specified upfront, nothing learned during implementation | Iterate: spec → build a slice → update the spec |
| Markdown nobody reads | Long, repetitive specs approved without review | Shorter specs, IDs, review checklist |
| Spec as a guarantee | "It's in the spec, so the code does it" | Tests linked to requirements, independent review |
| Spec without tests | Acceptance criteria never become checks | Requirement → test traceability, coverage report |
| Spec drift | Code changed, spec didn't | Update spec in the same PR; living specs |
| Tech details in the spec | Spec breaks when the stack changes | Move the how into the plan |
| Chasing every review finding | Over-engineering | Fix what matters for the requirements |
When Not to Use SDD¶
- Typos, one-line fixes, small refactors — describe the change and do it
- Exploration and prototypes — vibe code first, write a spec when you know what you want
- Behaviour you can't specify well yet — spike first, then specify
- When nobody will review the spec — the process adds cost without the benefit
See also¶
- Spec-Driven Development
- SDD — Writing Good Specs
- QA & Testing Methodology
- Testing Pyramid
- Test Design Techniques