Subscribe to Newsletter

Module 2: How To Write A Good Spec

In this module, you study what a good spec is and will learn to write one yourself.

Start module Module 2 of 3 · 5 lessons

2.1

Start With A Brief And Let The Agent Draft The Spec

Anthropic engineer Addy Osmani’s first principle is to begin with a clear goal statement and a few requirements, then let the agent expand that brief into a detailed spec. Review the draft before any code. Agents excel at elaboration when they have a clear mission, and they drift off course without one.

Osmani asks the agent to write the result to a file such as spec.md. That file anchors the agent when a session restarts or the history grows too long.

The draft usually comes back structured, with an overview, a feature list, tech stack suggestions, and a data model. Review it, correct any hallucination or off-target detail, and only then move to code.

Osmani gives one exception. This works well unless you already have very specific technical requirements that must hold from the start.

Keep The Brief About What And Why

A high-level brief focuses on what and why, not on how, at least at first. Ask three questions. Who is the user, what do they need, and what does success look like?

Osmani’s example brief for a to-do app reads, “Build a web app where users can track tasks (to-do list), with user accounts, a database, and a simple UI.”

His success criteria follow the same style. “User can add, edit, complete tasks; data is saved persistently; the app is responsive and secure.”

Spec Kit’s docs say the same. Describe what you build and why, and let the agent write the spec around user experience and success criteria. That is the what-and-how split covered earlier.

Make The Agent Restate The Task, List Assumptions, And Flag Ambiguity

Larridin CTO Ameya Kanitkar recommends that when an agent helps write the spec, you force it to restate the task, list its assumptions, and flag any ambiguity.

The check catches the details the model would otherwise pick silently. Anything still open after it gets its own investigation before the spec is final. The lesson on the spike-first workflow shows how Larridin does that.

Osmani runs the same check in a read-only planning mode, such as Claude Code’s Plan Mode. The agent analyzes the codebase and drafts the spec, but writes no code.

Ask it to question you about the plan, then to review the plan for architecture, best practices, security risks, and testing strategy. Refine until there is no room for misinterpretation, and only then let the agent execute.

HANDS-ON · The Kickoff Prompt, Run In Planning Mode

Goal: Turn a one-paragraph brief into a reviewed spec.md before the agent writes any code.

Steps:

  1. Switch your coding agent to its read-only planning mode, if it has one. Osmani’s example is Claude Code’s Plan Mode.
  2. Paste Osmani’s kickoff prompt, with your high-level brief in place of the placeholder:
prompt
You are an AI software engineer. Draft a detailed specification for [project X]
covering objectives, features, constraints, and a step-by-step plan.
  1. Before you read the draft, make the agent run Larridin’s check. Neither source gives a prompt for it, so this one is ours:
prompt
Before you finalize the spec: restate in your own words what needs to be built,
list every assumption you made, and flag anything in the brief that is ambiguous.
Ask me about each open point. Do not choose an answer yourself.
  1. Answer its questions.
  2. Ask it to review the plan for architecture, best practices, security risks, and testing strategy.
  3. Correct any hallucination or detail that misses your vision, until there is no room for misinterpretation.
  4. Leave planning mode.
  5. Save the result as spec.md.

Expected result: A spec.md with an overview, a feature list, tech stack suggestions, and a data model. The agent stated every assumption, you resolved every ambiguity, and no code exists yet.

2.2

The Six-Section Spec Template

Larridin CTO Ameya Kanitkar says to write the smallest spec that unambiguously specifies the system. If a section feels like fluff, cut it. If a decision feels obvious, write it down anyway, because “obvious to you” is not the same as “obvious to the model.”

The bar for detail is that the spec should be enough to hand to a junior engineer or a small model.

The Six Sections

SPEC.MD1Problem statementWhat you solve and for whom. One paragraph.2Non-goalsWhat the system will not do.3AssumptionsDecisions made, so the model does not guess.4Reference implementationThe promoted spike, shortcuts marked.5ArchitectureData model, boundaries, interfaces, errors.6Test planNamed tests with inputs and outputs. Done.
Scroll sideways →Source: How to Do Spec-Driven Development.

Larridin’s reference implementation is a promoted spike; the lesson on the spike-first workflow covers how they build one. The lesson on test plans covers the sixth section in full.

Blend A PRD With An SRS

Osmani’s second principle is to treat the spec as a structured document with clear sections, not a loose pile of notes. A formal document gives a “literal-minded” agent a blueprint.

Osmani blends two document types. The Product Requirements Document (PRD) side keeps the user-centric “why” behind each feature. The Software Requirements Specification (SRS) side nails down specifics such as which database or API to use.

Use a consistent format. Many developers use Markdown headings or XML-like tags to mark sections, because models handle well-structured text better than free-form prose.

Be specific about your stack. Say “React 18 with TypeScript, Vite, and Tailwind CSS,” not “React project,” and include versions and key dependencies. Vague specs produce vague code.

Osmani warns that “minimal does not necessarily mean short.” Do not shy away from detail where it matters, but keep the spec focused.

Six Things Every Spec Must Cover

GitHub’s analysis of more than 2,500 agent configuration files found that the most effective specs cover six areas. Osmani turns them into a completeness checklist:

  1. Commands. Full commands with flags, early in the file, such as npm test, pytest -v, npm run build. The agent references these constantly.
  2. Testing. How to run tests, which framework, where test files live, and what coverage you expect.
  3. Project structure. Where source code, tests, and docs live, stated explicitly.
  4. Code style. One real code snippet that shows your style beats three paragraphs that describe it. Include naming conventions and formatting rules.
  5. Git workflow. Branch naming, commit message format, PR requirements.
  6. Boundaries. What the agent should never touch, such as secrets, vendor directories, production configs, and specific folders.

These six areas belong to the codebase-scoped layer, and Osmani blends them into one Project Spec document. The lesson on boundaries covers them in full.

Two Skeletons To Copy

Larridin’s six sections as a blank template:

spec.md
# Spec: [feature or system]

## Problem statement
## Non-goals
## Assumptions
## Reference implementation
## Architecture
## Test plan

Osmani’s skeleton for the config-side areas, from his guide:

spec.md
# Project Spec: My team's tasks app

## Objective
- Build a web app for small teams to manage tasks...

## Tech Stack
- React 18+, TypeScript, Vite, Tailwind CSS
- Node.js/Express backend, PostgreSQL, Prisma ORM

## Commands
- Build: `npm run build` (compiles TypeScript, outputs to dist/)
- Test: `npm test` (runs Jest, must pass before commits)
- Lint: `npm run lint --fix` (auto-fixes ESLint errors)

## Project Structure
- `src/` – Application source code
- `tests/` – Unit and integration tests
- `docs/` – Documentation

## Boundaries
- ✅ Always: Run tests before commits, follow naming conventions
- ⚠️ Ask first: Database schema changes, adding dependencies
- 🚫 Never: Commit secrets, edit node_modules/, modify CI config
2.3

How To Set Boundaries The Agent Will Respect

Osmani’s fourth principle is that the spec is both coach and referee. A good spec anticipates where the agent might go wrong and sets up guardrails. It also uses what you know, such as domain knowledge, edge cases, and gotchas, so the agent does not operate in a vacuum.

The Three Tiers: Always, Ask First, Never

GitHub’s study of 2,500 agent files found that the most effective specs use a three-tier boundary system rather than a flat list of do’s and don’ts.

The tiers tell the agent when to proceed, when to pause, and when to stop:

  • ✅ Always do
  • ⚠️ Ask first
  • 🚫 Never do

“Never commit secrets” was the single most common helpful constraint in the study.

✅ ALWAYS DOThe agent proceedswithout asking“Always run tests before commits.”⚠️ ASK FIRSTThe agent pauses fora human check“Ask before adding dependencies.”🚫 NEVER DOThe agent stops.Hard limit.“Never commit secrets or keys.”
Scroll sideways →Source: How to write a good spec for AI agents.

Put Your Domain Knowledge In The Spec

Your spec should reflect insights that only an experienced developer, or someone with context, would know. Osmani’s phrase for it is to pour your mentorship into the spec.

If you build an e-commerce agent and you know that products and categories have a many-to-many relationship, state that clearly. Do not assume the agent will infer it, because it might not.

If a library is notoriously tricky, mention the pitfalls to avoid. The spec can hold advice such as “when using library X, watch out for the memory leak issue in version Y and apply this workaround.”

Encode your style preferences too, such as “use functional components, not class components in React,” and the agent will emulate your style.

A small example anchors the agent to the exact format you want. Many engineers include one in the spec, such as “All API responses should be JSON,” followed by the error shape {"error": "message"}.

A Boundaries Block To Copy

Osmani’s three tiers with his example lines, as one block for your rules file or constitution, the codebase-scoped layer:

rules file
## Boundaries

✅ Always do
- Always run tests before commits.
- Always follow the naming conventions in the style guide.
- Always log errors to the monitoring service.

⚠️ Ask first
- Ask before modifying database schemas.
- Ask before adding new dependencies.
- Ask before changing CI/CD configuration.

🚫 Never do
- Never commit secrets or API keys.
- Never edit node_modules/ or vendor/.
- Never remove a failing test without explicit approval.

The Code: Your daily unfair advantage in software engineering.

Join 350,000+ software engineers, tech leads, and CTOs who start their morning with The Code.

Subscribe to Newsletter
2.4

How To Write Test Plans The Agent Cannot Misread

Larridin CTO Ameya Kanitkar makes the test plan part of the spec. Once you know what the system does, write the test plan before the implementation plan, not after. Test-driven development and spec-driven development fit together. Go over it and adjust it if something is wrong or missing.

There is no need to implement the tests yet. List them, name each one, and specify inputs and expected outputs. That list is a clear definition of done, and you measure the implementation plan against it.

Watch out for “ambiguous tests.” The two pairs at the end of this lesson show Ojstersek’s fix.

Keep What BDD Taught: Plain Language And Given/When/Then

Thoughtworks Technology Director Liu Shangqi argues that experience from behavior-driven development (BDD) is still valid, and this new technology should not change much in this area.

Specifications should still:

  • Use domain-oriented ubiquitous language to describe business intent, not tech-bound implementations.
  • Have a clear structure, with a common style to define scenarios with Given/When/Then.
  • Strive for completeness yet conciseness, and cover the critical path without an enumeration of all cases. With AI, that also saves tokens.
  • Aim for clarity and determinism. LLMs do not generate deterministic code, but clear specifications still reduce hallucinations.

Liu adds that the spec-by-example we use in BDD is essentially the few-shot prompt technique.

Put Tests In The Success Criteria, Or In A Conformance Suite

Anthropic engineer Addy Osmani advises that you incorporate a test plan, or even actual tests, into your spec and prompt flow. In the spec’s success criteria, say “these sample inputs should produce these outputs” or “the following unit tests should pass.”

He also relays Simon Willison’s case for conformance suites, which are language-independent tests, often YAML-based, that any implementation must pass. A conformance suite is more rigorous than ad-hoc unit tests, because it derives directly from the spec, and you can reuse it across implementations.

Willison sums it up. A robust test suite is “like giving the agents superpowers,” because they can validate and iterate quickly when tests fail.

Precise Tests, And A Scenario Skeleton To Copy

Ojstersek’s two pairs, side by side:

INSTEAD OF

Handles large inputs gracefully

WRITE

Processes 10k rows in under 2 seconds with memory under 500MB

INSTEAD OF

Fails safely on bad input

WRITE

Returns a 400 with a specific error code when the payload is missing the customer_id field

The standard Given/When/Then form Thoughtworks recommends for each scenario, as a blank skeleton:

scenario
Scenario: [what the user does, in the domain's own words]
  Given [the starting state]
  When  [the action]
  Then  [the observable result]

Osmani’s one-line conformance reference, for the spec’s success criteria:

success criteria
Must pass all cases in conformance/api-tests.yaml
2.5

Reverse Engineering A Real Spec

This lesson walks one real spec in its own order. The spec is on the left. Step through it, and the note beside each section says what it does for the agent and which rule it applies.

The spec is the zero-dependency brainstorm server design by Jesse Vincent, author of the superpowers skills library for coding agents. It is about 860 words and changes an existing system, not a greenfield one.

Zero-Dependency Brainstorm Server

Replace the brainstorm companion server’s vendored node_modules (express, ws, chokidar — 714 tracked files) with a single zero-dependency server.js using only Node.js built-ins.

Motivation

Vendoring node_modules into the git repo creates a supply chain risk: frozen dependencies don’t get security patches, 714 files of third-party code are committed without audit, and modifications to vendored code look like normal commits. While the actual risk is low (localhost-only dev server), eliminating it is straightforward.

Architecture

A single server.js file (~250-300 lines) using http, crypto, fs, and path. The file serves two roles:

  • When run directly (node server.js): starts the HTTP/WebSocket server
  • When required (require('./server.js')): exports WebSocket protocol functions for unit testing
WebSocket Protocol

Implements RFC 6455 for text frames only:

Handshake: Compute Sec-WebSocket-Accept from client’s Sec-WebSocket-Key using SHA-1 + the RFC 6455 magic GUID. Return 101 Switching Protocols.

Frame decoding (client to server): Handle three masked length encodings:

  • Small: payload < 126 bytes
  • Medium: 126-65535 bytes (16-bit extended)
  • Large: > 65535 bytes (64-bit extended)

XOR-unmask payload using 4-byte mask key. Return { opcode, payload, bytesConsumed } or null for incomplete buffers. Reject unmasked frames.

Frame encoding (server to client): Unmasked frames with the same three length encodings.

Opcodes handled: TEXT (0x01), CLOSE (0x08), PING (0x09), PONG (0x0A). Unrecognized opcodes get a close frame with status 1003 (Unsupported Data).

Deliberately skipped: Binary frames, fragmented messages, extensions (permessage-deflate), subprotocols. These are unnecessary for small JSON text messages between localhost clients. Extensions and subprotocols are negotiated in the handshake — by not advertising them, they are never active.

Buffer accumulation: Each connection maintains a buffer. On data, append and loop decodeFrame until it returns null or buffer is empty.

HTTP Server

Three routes:

  1. GET / — Serve newest .html from screen directory by mtime. Detect full documents vs fragments, wrap fragments in frame template, inject helper.js. Return text/html. When no .html files exist, serve a hardcoded waiting page (“Waiting for Claude to push a screen...”) with helper.js injected.
  2. GET /files/* — Serve static files from screen directory with MIME type lookup from a hardcoded extension map (html, css, js, png, jpg, gif, svg, json). Return 404 if not found.
  3. Everything else — 404.

WebSocket upgrade handled via the 'upgrade' event on the HTTP server, separate from the request handler.

Configuration

Environment variables (all optional):

  • BRAINSTORM_PORT — port to bind (default: random high port 49152-65535)
  • BRAINSTORM_HOST — interface to bind (default: 127.0.0.1)
  • BRAINSTORM_URL_HOST — hostname for the URL in startup JSON (default: localhost when host is 127.0.0.1, otherwise same as host)
  • BRAINSTORM_DIR — screen directory path (default: /tmp/brainstorm)
Startup Sequence
  1. Create SCREEN_DIR if it doesn’t exist (mkdirSync recursive)
  2. Load frame template and helper.js from __dirname
  3. Start HTTP server on configured host/port
  4. Start fs.watch on SCREEN_DIR
  5. On successful listen, log server-started JSON to stdout: { type, port, host, url_host, url, screen_dir }
  6. Write the same JSON to SCREEN_DIR/.server-info so agents can find connection details when stdout is hidden (background execution)
Application-Level WebSocket Messages

When a TEXT frame arrives from a client:

  1. Parse as JSON. If parsing fails, log to stderr and continue.
  2. Log to stdout as { source: 'user-event', ...event }.
  3. If the event contains a choice property, append the JSON to SCREEN_DIR/.events (one line per event).
File Watching

fs.watch(SCREEN_DIR) replaces chokidar. On HTML file events:

  • On new file (rename event for a file that exists): delete .events file if present (unlinkSync), log screen-added to stdout as JSON
  • On file change (change event): log screen-updated to stdout as JSON (do NOT clear .events)
  • Both events: send { type: 'reload' } to all connected WebSocket clients

Debounce per-filename with ~100ms timeout to prevent duplicate events (common on macOS and Linux).

Error Handling
  • Malformed JSON from WebSocket clients: log to stderr, continue
  • Unhandled opcodes: close with status 1003
  • Client disconnects: remove from broadcast set
  • fs.watch errors: log to stderr, continue
  • No graceful shutdown logic — shell scripts handle process lifecycle via SIGTERM
What Changes
BeforeAfter
index.js + package.json + package-lock.json + 714 node_modules filesserver.js (single file)
express, ws, chokidar dependenciesnone
No static file serving/files/* serves from screen directory
What Stays the Same
  • helper.js — no changes
  • frame-template.html — no changes
  • start-server.sh — one-line update: index.js to server.js
  • stop-server.sh — no changes
  • visual-companion.md — no changes
  • All existing server behavior and external contract
Platform Compatibility
  • server.js uses only cross-platform Node built-ins
  • fs.watch is reliable for single flat directories on macOS, Linux, and Windows
  • Shell scripts require bash (Git Bash on Windows, which is required for Claude Code)
Testing

Unit tests (ws-protocol.test.js): Test WebSocket frame encoding/decoding, handshake computation, and protocol edge cases directly by requiring server.js exports.

Integration tests (server.test.js): Test full server behavior — HTTP serving, WebSocket communication, file watching, brainstorming workflow. Uses ws npm package as a test-only client dependency (not shipped to end users).

Zero-Dependency Brainstorm Server, by Jesse Vincent, MIT license. Shown in full.

SECTION 1 OF 11

Title

One sentence states what changes and into what, as 714 vendored files become one file with no dependencies. The agent knows the problem and the shape of done before it reads anything else. This is Larridin’s problem statement, in one paragraph.

END OF MODULE 2

By this point you should have:

  • A kickoff prompt that turns a one-paragraph brief into a reviewed spec.
  • A check that makes the agent restate the task, list assumptions, and flag ambiguity.
  • Larridin’s six-section template and Osmani’s Project Spec skeleton, ready to copy.
  • A three-tier boundaries block for your rules file or constitution.
  • Two ambiguous-versus-precise test pairs, a Given/When/Then skeleton, and a one-line conformance reference.
  • A real, annotated brownfield spec to calibrate your own against.