---
title: "An Agent Ready Spec Checklist for PMs"
description: "A spec is agent ready when ten checks pass: scope fence, file paths, verification command, done condition, rollback, and the ambiguities you answered up front."
canonical_url: "https://builderscamp.com/guides/templates/agent-ready-spec-checklist"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Guilherme Salgueiro"
publisher: "Builders Camp"
guide_class: "templates"
---

# Agent ready spec checklist for product managers

**TL;DR:** Ten checks decide whether a spec is executable or still ambiguous, and the two most commonly failed are the verification command and the rollback line. Run them against a draft before an agent does, because every check that fails is a decision the agent makes without you. A passing checklist proves the spec is buildable, not that the feature is worth building.

## What is this checklist testing for?

Whether the spec can be finished without you in the room. A clear spec can be understood; an agent ready spec can be completed, checked, and stopped without a human supplying the missing judgement at the end.

Every failed check below turns into the same thing at run time: a decision made without you, quietly, in a diff that looks plausible. Run the list against a draft before an agent does. It takes about five minutes and it is the cheapest review in the whole workflow. If you want the writing process rather than the audit, [how to write a PRD for Claude Code](https://builderscamp.com/guides/templates/how-to-write-a-prd-for-claude-code) covers the drafting side.

## The ten checks

| # | The check | How you know it failed | What it costs at run time |
|---|---|---|---|
| 1 | A verification command is named and runnable | The spec says the feature should work, but names nothing the agent can execute | The only stopping signal is that the work looks done, so you become the verification loop |
| 2 | The done condition resolves without your judgement | It reads "the flow feels right" rather than "the new test passes and the suite still passes" | The agent stops at its own estimate of good enough, which is rarely yours |
| 3 | Every file you already know about is named | The spec describes behaviour with no path in it anywhere | An exploration pass spends context finding something you could have written in one line |
| 4 | A scope fence names what must not change | No section says what is off limits | An adjacent refactor arrives with the feature, and the diff looks reasonable |
| 5 | Empty, error and permission states are specified | Only the happy path is described | The blank page, the timeout and the unauthorised case get whatever seemed reasonable |
| 6 | Data access and write permissions are explicit | The spec assumes the agent knows what it may read or change | Either an over-cautious stall or a write you did not expect |
| 7 | Prior art is named, or its absence is stated | No existing pattern is referenced for a thing your codebase already does three times | A fourth way of doing it, inconsistent with the other three |
| 8 | A rollback line exists | Nothing says how the change is undone | The revert gets designed during the incident it caused |
| 9 | Nothing in the spec contradicts the standing rules | The task instruction and the repository's own conventions file disagree | Unpredictable precedence, and a result nobody can reproduce |
| 10 | The human gate is named before the run starts | No line says at which point a person must look | The review happens after the merge, or not at all |

## Which checks do people actually skip?

Numbers 8 and 9, and neither one for a good reason.

Rollback gets skipped because writing it feels like planning for failure on a change you expect to work. It is one sentence: how this is reverted, and what evidence would make you revert it. Teams that skip it do not save the minute; they spend it later, under worse conditions, with more people watching.

The contradiction check gets skipped because nobody thinks to look for it. A repository conventions file says never edit generated output, and this week's task says update the generated map to match the new schema. Both instructions are reasonable. Together they produce a run whose behaviour depends on which one the model weighted more heavily, and a result you cannot reproduce next week. Read the task against the standing rules before the run, and where they genuinely conflict, say in the spec which one wins for this task.

## How do you run it in practice?

Read the draft once with the checklist beside it and mark each item pass or fail. Do not fix anything on the first pass, because the useful signal is the pattern: a spec failing checks 1, 2 and 10 has no verification story at all, which is a different problem from a spec failing 4 and 7, which is a context problem.

Then fix in order of blast radius rather than in list order. The fence and the human gate bound what a bad run can do. The verification command and the done condition determine whether a good run stops in the right place. The context items, files and prior art, only make a correct run cheaper. If you have time for three repairs, take 1, 4 and 10.

- **Bound the damage first**: the fence and the human gate.
- **Close the loop second**: the verification command and the done condition.
- **Save context last**: file paths and prior art, which affect cost rather than correctness.

## What a passing checklist does not prove

That the feature should exist. The checklist tests executability, and a perfectly specified bad idea executes perfectly.

It also does not remove review. Anthropic's own guidance builds most of its patterns around giving the agent a check it can run, and adds an adversarial review step as a separate practice on top, because a test written in the same pass as the code is a weaker signal than it appears. Where the check has to be non-negotiable, a hook can run it as a script and block the turn from ending until it passes, which turns a written instruction into a mechanical gate. That is the same idea as a [quality gate in an AI pipeline](https://builderscamp.com/guides/glossary/quality-gate-ai-pipeline), applied to a single task rather than a whole workflow.

The honest limitation is that all ten checks are cheap to write and none of them are free to maintain. A spec with a stale verification command is worse than one with none, because it produces a green result nobody earned. If the command changes, the specs that name it change too, which is a good argument for keeping it in the repository's standing context file rather than retyping it into every task.

## Does this change for an agent that does not write code?

Seven of the ten transfer unchanged. A research agent, a data agent, or a feedback triage workflow still needs a fence, a done condition, explicit data permissions, and a named human gate, and it still fails in the same way when the empty case is unspecified.

Check 1 is the one that changes shape. When the output is a document rather than a build, the runnable check becomes a rubric applied by a second pass, or a small set of known inputs with known correct answers. That is closer to [acceptance criteria](https://builderscamp.com/guides/glossary/acceptance-criteria) than to a test suite, and it is written the same way: binary conditions, no partial credit. Builders Camp's AI Agents bootcamp covers that layer directly across 2 weeks and 3 live sessions: tool use, orchestration, memory, human approvals for high-stakes actions, and evaluating agent behaviour with scenarios and logs rather than by reading the output and nodding.

## Where the whole system is taught

Builders Camp's Building with Claude Code bootcamp, taught by Guilherme Salgueiro, treats this checklist as one piece of a standing setup rather than a per-task ritual: context architecture and a repository conventions file, reusable skills and agents, Product Requirement Prompts against traditional PRDs, then hooks and quality gates so the checks run themselves. The AI Agentic Builders Expert Track bundles it with agent design for people who want the ordered path.

[See the Building with Claude Code bootcamp](https://builderscamp.com/bootcamps/building-with-claude-code?utm_source=guide&utm_medium=organic&utm_campaign=agent-ready-spec-checklist)

The strongest signal that a team has internalised this list is not that their specs get longer. It is that their specs get shorter, because checks 3, 7 and 9 have moved into the repository's own standing context and stopped needing to be written at all.

## Frequently asked questions

### What makes a spec agent ready rather than just clear?

A clear spec can be understood. An agent ready spec can be finished without you. The difference is whether the document contains a check the agent can run and a done condition that resolves to true or false on its own, rather than one that needs your judgement at the end.

### Which check fails most often?

The verification command. Most specs describe the outcome and then leave proving it to a human, which means the agent's only stopping signal is that the work looks done. A command it can run and read the output of is what closes that loop without you sitting in it.

### Do I need to name every file the change will touch?

Name every file you already know about. Where you do not know, say so and ask for an exploration pass first. Guessing a path that does not exist is worse than admitting the gap, because the agent will spend a pass reconciling your guess with reality.

### What is a scope fence?

The explicit list of what must not change: generated files, migration directories, a shared component package, anything with its own review process. An agent fixing one thing will restructure something adjacent if nothing tells it not to, and the diff will look reasonable.

### Why does a spec need a rollback line?

Because deciding how to undo a change is much easier before it exists than during the incident it caused. One sentence naming how the change is reverted, and what evidence would trigger reverting it, costs a minute and saves an argument.

### Does a passing checklist mean the feature is a good idea?

No. The checklist tests whether the spec is executable, not whether the work is worth doing. That judgement belongs in a product document written for people, and no amount of specification precision substitutes for it.

### Does this checklist apply to non-coding agents too?

Most of it does. A research or data agent still needs a scope fence, a done condition, and a named human gate. The verification check is the item that changes shape, since the output is a document rather than a build, so the check becomes a rubric or a second pass rather than a test suite.

## Sources

- [Claude Code Docs: Best practices for Claude Code](https://www.anthropic.com/engineering/claude-code-best-practices)
- [Claude Code Docs: Common workflows](https://docs.claude.com/en/docs/claude-code/common-workflows)
- [Claude Code Docs: Set up hooks](https://docs.claude.com/en/docs/claude-code/hooks)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
