Module 2: How To Write A Good Spec
In this module, you study what a good spec is and will learn to write one yourself.
Start With A Brief And Let The Agent Draft The Spec
Anthropic engineer Addy Osmani’s first principle is to begin with a clear goal statement and a few requirements, then let the agent expand that brief into a detailed spec. Review the draft before any code. Agents excel at elaboration when they have a clear mission, and they drift off course without one.
Osmani asks the agent to write the result to a file such as spec.md. That file anchors the agent when a session restarts or the history grows too long.
The draft usually comes back structured, with an overview, a feature list, tech stack suggestions, and a data model. Review it, correct any hallucination or off-target detail, and only then move to code.
Osmani gives one exception. This works well unless you already have very specific technical requirements that must hold from the start.
Keep The Brief About What And Why
A high-level brief focuses on what and why, not on how, at least at first. Ask three questions. Who is the user, what do they need, and what does success look like?
Osmani’s example brief for a to-do app reads, “Build a web app where users can track tasks (to-do list), with user accounts, a database, and a simple UI.”
His success criteria follow the same style. “User can add, edit, complete tasks; data is saved persistently; the app is responsive and secure.”
Spec Kit’s docs say the same. Describe what you build and why, and let the agent write the spec around user experience and success criteria. That is the what-and-how split covered earlier.
Make The Agent Restate The Task, List Assumptions, And Flag Ambiguity
Larridin CTO Ameya Kanitkar recommends that when an agent helps write the spec, you force it to restate the task, list its assumptions, and flag any ambiguity.
The check catches the details the model would otherwise pick silently. Anything still open after it gets its own investigation before the spec is final. The lesson on the spike-first workflow shows how Larridin does that.
Osmani runs the same check in a read-only planning mode, such as Claude Code’s Plan Mode. The agent analyzes the codebase and drafts the spec, but writes no code.
Ask it to question you about the plan, then to review the plan for architecture, best practices, security risks, and testing strategy. Refine until there is no room for misinterpretation, and only then let the agent execute.
Goal: Turn a one-paragraph brief into a reviewed spec.md before the agent writes any code.
Steps:
- Switch your coding agent to its read-only planning mode, if it has one. Osmani’s example is Claude Code’s Plan Mode.
- Paste Osmani’s kickoff prompt, with your high-level brief in place of the placeholder:
You are an AI software engineer. Draft a detailed specification for [project X]
covering objectives, features, constraints, and a step-by-step plan.
- Before you read the draft, make the agent run Larridin’s check. Neither source gives a prompt for it, so this one is ours:
Before you finalize the spec: restate in your own words what needs to be built,
list every assumption you made, and flag anything in the brief that is ambiguous.
Ask me about each open point. Do not choose an answer yourself.
- Answer its questions.
- Ask it to review the plan for architecture, best practices, security risks, and testing strategy.
- Correct any hallucination or detail that misses your vision, until there is no room for misinterpretation.
- Leave planning mode.
- Save the result as spec.md.
Expected result: A spec.md with an overview, a feature list, tech stack suggestions, and a data model. The agent stated every assumption, you resolved every ambiguity, and no code exists yet.
The Six-Section Spec Template
Larridin CTO Ameya Kanitkar says to write the smallest spec that unambiguously specifies the system. If a section feels like fluff, cut it. If a decision feels obvious, write it down anyway, because “obvious to you” is not the same as “obvious to the model.”
The bar for detail is that the spec should be enough to hand to a junior engineer or a small model.
The Six Sections
Larridin’s reference implementation is a promoted spike; the lesson on the spike-first workflow covers how they build one. The lesson on test plans covers the sixth section in full.
Blend A PRD With An SRS
Osmani’s second principle is to treat the spec as a structured document with clear sections, not a loose pile of notes. A formal document gives a “literal-minded” agent a blueprint.
Osmani blends two document types. The Product Requirements Document (PRD) side keeps the user-centric “why” behind each feature. The Software Requirements Specification (SRS) side nails down specifics such as which database or API to use.
Use a consistent format. Many developers use Markdown headings or XML-like tags to mark sections, because models handle well-structured text better than free-form prose.
Be specific about your stack. Say “React 18 with TypeScript, Vite, and Tailwind CSS,” not “React project,” and include versions and key dependencies. Vague specs produce vague code.
Osmani warns that “minimal does not necessarily mean short.” Do not shy away from detail where it matters, but keep the spec focused.
Six Things Every Spec Must Cover
GitHub’s analysis of more than 2,500 agent configuration files found that the most effective specs cover six areas. Osmani turns them into a completeness checklist:
- Commands. Full commands with flags, early in the file, such as
npm test,pytest -v,npm run build. The agent references these constantly. - Testing. How to run tests, which framework, where test files live, and what coverage you expect.
- Project structure. Where source code, tests, and docs live, stated explicitly.
- Code style. One real code snippet that shows your style beats three paragraphs that describe it. Include naming conventions and formatting rules.
- Git workflow. Branch naming, commit message format, PR requirements.
- Boundaries. What the agent should never touch, such as secrets, vendor directories, production configs, and specific folders.
These six areas belong to the codebase-scoped layer, and Osmani blends them into one Project Spec document. The lesson on boundaries covers them in full.
Two Skeletons To Copy
Larridin’s six sections as a blank template:
# Spec: [feature or system]
## Problem statement
## Non-goals
## Assumptions
## Reference implementation
## Architecture
## Test plan
Osmani’s skeleton for the config-side areas, from his guide:
# Project Spec: My team's tasks app
## Objective
- Build a web app for small teams to manage tasks...
## Tech Stack
- React 18+, TypeScript, Vite, Tailwind CSS
- Node.js/Express backend, PostgreSQL, Prisma ORM
## Commands
- Build: `npm run build` (compiles TypeScript, outputs to dist/)
- Test: `npm test` (runs Jest, must pass before commits)
- Lint: `npm run lint --fix` (auto-fixes ESLint errors)
## Project Structure
- `src/` – Application source code
- `tests/` – Unit and integration tests
- `docs/` – Documentation
## Boundaries
- ✅ Always: Run tests before commits, follow naming conventions
- ⚠️ Ask first: Database schema changes, adding dependencies
- 🚫 Never: Commit secrets, edit node_modules/, modify CI config
How To Set Boundaries The Agent Will Respect
Osmani’s fourth principle is that the spec is both coach and referee. A good spec anticipates where the agent might go wrong and sets up guardrails. It also uses what you know, such as domain knowledge, edge cases, and gotchas, so the agent does not operate in a vacuum.
The Three Tiers: Always, Ask First, Never
GitHub’s study of 2,500 agent files found that the most effective specs use a three-tier boundary system rather than a flat list of do’s and don’ts.
The tiers tell the agent when to proceed, when to pause, and when to stop:
- ✅ Always do
- ⚠️ Ask first
- 🚫 Never do
“Never commit secrets” was the single most common helpful constraint in the study.
Put Your Domain Knowledge In The Spec
Your spec should reflect insights that only an experienced developer, or someone with context, would know. Osmani’s phrase for it is to pour your mentorship into the spec.
If you build an e-commerce agent and you know that products and categories have a many-to-many relationship, state that clearly. Do not assume the agent will infer it, because it might not.
If a library is notoriously tricky, mention the pitfalls to avoid. The spec can hold advice such as “when using library X, watch out for the memory leak issue in version Y and apply this workaround.”
Encode your style preferences too, such as “use functional components, not class components in React,” and the agent will emulate your style.
A small example anchors the agent to the exact format you want. Many engineers include one in the spec, such as “All API responses should be JSON,” followed by the error shape {"error": "message"}.
A Boundaries Block To Copy
Osmani’s three tiers with his example lines, as one block for your rules file or constitution, the codebase-scoped layer:
## Boundaries
✅ Always do
- Always run tests before commits.
- Always follow the naming conventions in the style guide.
- Always log errors to the monitoring service.
⚠️ Ask first
- Ask before modifying database schemas.
- Ask before adding new dependencies.
- Ask before changing CI/CD configuration.
🚫 Never do
- Never commit secrets or API keys.
- Never edit node_modules/ or vendor/.
- Never remove a failing test without explicit approval.
The Code: Your daily unfair advantage in software engineering.
Join 350,000+ software engineers, tech leads, and CTOs who start their morning with The Code.
How To Write Test Plans The Agent Cannot Misread
Larridin CTO Ameya Kanitkar makes the test plan part of the spec. Once you know what the system does, write the test plan before the implementation plan, not after. Test-driven development and spec-driven development fit together. Go over it and adjust it if something is wrong or missing.
There is no need to implement the tests yet. List them, name each one, and specify inputs and expected outputs. That list is a clear definition of done, and you measure the implementation plan against it.
Watch out for “ambiguous tests.” The two pairs at the end of this lesson show Ojstersek’s fix.
Keep What BDD Taught: Plain Language And Given/When/Then
Thoughtworks Technology Director Liu Shangqi argues that experience from behavior-driven development (BDD) is still valid, and this new technology should not change much in this area.
Specifications should still:
- Use domain-oriented ubiquitous language to describe business intent, not tech-bound implementations.
- Have a clear structure, with a common style to define scenarios with Given/When/Then.
- Strive for completeness yet conciseness, and cover the critical path without an enumeration of all cases. With AI, that also saves tokens.
- Aim for clarity and determinism. LLMs do not generate deterministic code, but clear specifications still reduce hallucinations.
Liu adds that the spec-by-example we use in BDD is essentially the few-shot prompt technique.
Put Tests In The Success Criteria, Or In A Conformance Suite
Anthropic engineer Addy Osmani advises that you incorporate a test plan, or even actual tests, into your spec and prompt flow. In the spec’s success criteria, say “these sample inputs should produce these outputs” or “the following unit tests should pass.”
He also relays Simon Willison’s case for conformance suites, which are language-independent tests, often YAML-based, that any implementation must pass. A conformance suite is more rigorous than ad-hoc unit tests, because it derives directly from the spec, and you can reuse it across implementations.
Willison sums it up. A robust test suite is “like giving the agents superpowers,” because they can validate and iterate quickly when tests fail.
Precise Tests, And A Scenario Skeleton To Copy
Ojstersek’s two pairs, side by side:
Handles large inputs gracefully
Processes 10k rows in under 2 seconds with memory under 500MB
Fails safely on bad input
Returns a 400 with a specific error code when the payload is missing the customer_id field
The standard Given/When/Then form Thoughtworks recommends for each scenario, as a blank skeleton:
Scenario: [what the user does, in the domain's own words]
Given [the starting state]
When [the action]
Then [the observable result]
Osmani’s one-line conformance reference, for the spec’s success criteria:
Must pass all cases in conformance/api-tests.yaml
Reverse Engineering A Real Spec
This lesson walks one real spec in its own order. The spec is on the left. Step through it, and the note beside each section says what it does for the agent and which rule it applies.
The spec is the zero-dependency brainstorm server design by Jesse Vincent, author of the superpowers skills library for coding agents. It is about 860 words and changes an existing system, not a greenfield one.
Zero-Dependency Brainstorm Server
Replace the brainstorm companion server’s vendored node_modules (express, ws, chokidar — 714 tracked files) with a single zero-dependency server.js using only Node.js built-ins.
Motivation
Vendoring node_modules into the git repo creates a supply chain risk: frozen dependencies don’t get security patches, 714 files of third-party code are committed without audit, and modifications to vendored code look like normal commits. While the actual risk is low (localhost-only dev server), eliminating it is straightforward.
Architecture
A single server.js file (~250-300 lines) using http, crypto, fs, and path. The file serves two roles:
- When run directly (
node server.js): starts the HTTP/WebSocket server - When required (
require('./server.js')): exports WebSocket protocol functions for unit testing
WebSocket Protocol
Implements RFC 6455 for text frames only:
Handshake: Compute Sec-WebSocket-Accept from client’s Sec-WebSocket-Key using SHA-1 + the RFC 6455 magic GUID. Return 101 Switching Protocols.
Frame decoding (client to server): Handle three masked length encodings:
- Small: payload < 126 bytes
- Medium: 126-65535 bytes (16-bit extended)
- Large: > 65535 bytes (64-bit extended)
XOR-unmask payload using 4-byte mask key. Return { opcode, payload, bytesConsumed } or null for incomplete buffers. Reject unmasked frames.
Frame encoding (server to client): Unmasked frames with the same three length encodings.
Opcodes handled: TEXT (0x01), CLOSE (0x08), PING (0x09), PONG (0x0A). Unrecognized opcodes get a close frame with status 1003 (Unsupported Data).
Deliberately skipped: Binary frames, fragmented messages, extensions (permessage-deflate), subprotocols. These are unnecessary for small JSON text messages between localhost clients. Extensions and subprotocols are negotiated in the handshake — by not advertising them, they are never active.
Buffer accumulation: Each connection maintains a buffer. On data, append and loop decodeFrame until it returns null or buffer is empty.
HTTP Server
Three routes:
GET /— Serve newest.htmlfrom screen directory by mtime. Detect full documents vs fragments, wrap fragments in frame template, inject helper.js. Returntext/html. When no.htmlfiles exist, serve a hardcoded waiting page (“Waiting for Claude to push a screen...”) with helper.js injected.GET /files/*— Serve static files from screen directory with MIME type lookup from a hardcoded extension map (html, css, js, png, jpg, gif, svg, json). Return 404 if not found.- Everything else — 404.
WebSocket upgrade handled via the 'upgrade' event on the HTTP server, separate from the request handler.
Configuration
Environment variables (all optional):
BRAINSTORM_PORT— port to bind (default: random high port 49152-65535)BRAINSTORM_HOST— interface to bind (default:127.0.0.1)BRAINSTORM_URL_HOST— hostname for the URL in startup JSON (default:localhostwhen host is127.0.0.1, otherwise same as host)BRAINSTORM_DIR— screen directory path (default:/tmp/brainstorm)
Startup Sequence
- Create
SCREEN_DIRif it doesn’t exist (mkdirSyncrecursive) - Load frame template and helper.js from
__dirname - Start HTTP server on configured host/port
- Start
fs.watchonSCREEN_DIR - On successful listen, log
server-startedJSON to stdout:{ type, port, host, url_host, url, screen_dir } - Write the same JSON to
SCREEN_DIR/.server-infoso agents can find connection details when stdout is hidden (background execution)
Application-Level WebSocket Messages
When a TEXT frame arrives from a client:
- Parse as JSON. If parsing fails, log to stderr and continue.
- Log to stdout as
{ source: 'user-event', ...event }. - If the event contains a
choiceproperty, append the JSON toSCREEN_DIR/.events(one line per event).
File Watching
fs.watch(SCREEN_DIR) replaces chokidar. On HTML file events:
- On new file (
renameevent for a file that exists): delete.eventsfile if present (unlinkSync), logscreen-addedto stdout as JSON - On file change (
changeevent): logscreen-updatedto stdout as JSON (do NOT clear.events) - Both events: send
{ type: 'reload' }to all connected WebSocket clients
Debounce per-filename with ~100ms timeout to prevent duplicate events (common on macOS and Linux).
Error Handling
- Malformed JSON from WebSocket clients: log to stderr, continue
- Unhandled opcodes: close with status 1003
- Client disconnects: remove from broadcast set
fs.watcherrors: log to stderr, continue- No graceful shutdown logic — shell scripts handle process lifecycle via SIGTERM
What Changes
| Before | After |
|---|---|
index.js + package.json + package-lock.json + 714 node_modules files | server.js (single file) |
| express, ws, chokidar dependencies | none |
| No static file serving | /files/* serves from screen directory |
What Stays the Same
helper.js— no changesframe-template.html— no changesstart-server.sh— one-line update:index.jstoserver.jsstop-server.sh— no changesvisual-companion.md— no changes- All existing server behavior and external contract
Platform Compatibility
server.jsuses only cross-platform Node built-insfs.watchis reliable for single flat directories on macOS, Linux, and Windows- Shell scripts require bash (Git Bash on Windows, which is required for Claude Code)
Testing
Unit tests (ws-protocol.test.js): Test WebSocket frame encoding/decoding, handshake computation, and protocol edge cases directly by requiring server.js exports.
Integration tests (server.test.js): Test full server behavior — HTTP serving, WebSocket communication, file watching, brainstorming workflow. Uses ws npm package as a test-only client dependency (not shipped to end users).
Zero-Dependency Brainstorm Server, by Jesse Vincent, MIT license. Shown in full.
Title
One sentence states what changes and into what, as 714 vendored files become one file with no dependencies. The agent knows the problem and the shape of done before it reads anything else. This is Larridin’s problem statement, in one paragraph.
Motivation
The section names the risk, unpatched third-party code, and then sizes it as low, on a localhost-only server. That size tells the agent how much to build. Without it, an agent treats every risk as critical and over-engineers the fix.
Architecture
The spec fixes one file, a line budget, four built-ins, the RFC, three length encodings, and four opcodes. Each is a decision the agent would otherwise make on its own, and silently. Written down, they cannot drift between sessions.
Deliberately skipped
Four features the server will not have, with the reason and the proof that the omission is safe. This is where the spec stops the agent from building a full WebSocket library. The non-goals sit next to the section they apply to.
Configuration and Startup Sequence
Every variable carries its default, and the one non-obvious startup step carries its reason. Obvious to you is not obvious to the model. A missing default is a guess the agent makes for you.
Error Handling
Each line pairs a condition with one response, so the agent can test each case. The last line is a decision not to build shutdown logic. A precise no is as useful as a precise yes.
What Changes and What Stays the Same
The table lists what goes and the one file that replaces it. The list names six files the agent must not touch, and the one exception. These boundaries make a brownfield change safe.
Platform Compatibility
Only someone with context knows these three facts, which built-ins are cross-platform, where the file watcher is reliable, and that the scripts need bash. Without them, the agent finds out in production.
Testing
Two test files, and what each covers. The test-only dependency is named as test-only, so the agent does not ship it. This list is the definition of done.
Each Section, And The Rule It Applies
The Heading Outline To Copy
The outline of the original, to reuse as a skeleton for a brownfield spec:
By this point you should have:
- A kickoff prompt that turns a one-paragraph brief into a reviewed spec.
- A check that makes the agent restate the task, list assumptions, and flag ambiguity.
- Larridin’s six-section template and Osmani’s Project Spec skeleton, ready to copy.
- A three-tier boundaries block for your rules file or constitution.
- Two ambiguous-versus-precise test pairs, a Given/When/Then skeleton, and a one-line conformance reference.
- A real, annotated brownfield spec to calibrate your own against.
Module 3: How To Run SDD End To End
This module helps you get started with GitHub Spec Kit, shares Microsoft’s adoption story, and shows how you can avoid common pitfalls with SDD.