Builders Camp

Templates

How to write acceptance criteria with AI

Give a model the story and a state checklist, and it will enumerate the cases you forgot faster than any refinement session will. It will also produce criteria nobody can test, so every output gets one review pass that asks who observes this condition and with what. The thresholds are always yours, never the model's.

What does a model actually add here?

Coverage, not judgement. Handed a story and a list of states to consider, it enumerates the empty case, the permission case, the concurrency case and the migration case faster and more completely than a tired team at the end of a refinement session. What it cannot do is decide the bar, because nothing in the story tells it whether two seconds is acceptable for your users or whether losing one request in a hundred is tolerable for this particular action.

So the workflow splits cleanly. The model drafts the states. You set every number. Then one review pass catches the criteria that cannot be tested at all, which is where most of the damage hides, because an untestable criterion passes review by sounding reasonable and fails later by being unresolvable in the moment someone tries to use it. The definition and the collaboration model behind all of this are in acceptance criteria.

What does the transformation look like on a real story?

Take a vague one, deliberately: as a team admin, I want to remove someone from my workspace, so that former colleagues cannot see our data.

A first AI pass typically returns something like "the user is removed successfully", "the system handles errors gracefully", and "removal is fast". All three sound like criteria. None of them are. Nobody can look at the running product and agree on whether error handling was graceful, and "successfully" is doing every piece of work in the first one.

Now the same story after the states are enumerated and the numbers are yours:

  • Given an admin with at least one other member, when they remove that member, the member loses access on their next request and within 60 seconds on an existing session.
  • Given the removed member had documents assigned to them, those documents remain in the workspace and show an unassigned owner rather than disappearing.
  • Given an admin attempts to remove the last remaining admin, the action is refused with a message naming why.
  • Given the removal request fails partway, no partial state persists and the admin sees the failure rather than a success message.
  • The removal is written to the audit log with actor, subject and timestamp, readable by any admin.

Five statements, each observable by someone other than the person who wrote them. The model produced the shape of all five. The 60 seconds, the decision that documents survive, and the last-admin rule came from a person, and those are the three things the whole story actually turns on.

Which criteria cannot be tested, and how do you spot them?

Five patterns cover nearly all of them, and each has a repair rather than a deletion.

The subjective adjective. Graceful, intuitive, clean, smooth. Ask what the user does differently when the adjective is true, and write that instead.

The unbounded quantifier. Fast, reliable, most users, generally. Replace with the number you were avoiding choosing. If you genuinely do not know the number yet, that is an open question for the spec, not a criterion.

The hidden human judge. "The summary is accurate." Somebody has to decide accuracy, and unless you name the rubric and the reviewer, that somebody will be whoever is in the room at the wrong moment.

The missing observer. "The data is encrypted at rest." True or false, but who checks, and how? A criterion with no observer is an intention.

The compound condition. "The user can filter and sort results and the page loads quickly." Three criteria wearing one bullet, and it can only ever be partly true, which defeats the binary property that makes a criterion useful.

Run those five over any generated list and the untestable ones fall out in under two minutes. The check underneath all of them is the same question: who observes this, and with what.

What shape should the criteria take?

Given, When, Then where there is a trigger and an outcome, flat statements where there is not.

The Gherkin reference describes this structure as a way of expressing behaviour through concrete examples rather than abstract rules, and Cucumber's own guidance on writing better Gherkin pushes toward describing what happens rather than how the interface does it. That is the same rule the story itself follows, applied one level down.

The trap is forcing everything into the scenario shape because the prompt asked for it. A latency bound, a retention period, an accessibility requirement: these are conditions, not scenarios, and wrapping them in a Given clause adds a sentence and removes clarity. When the output contorts to fit the template, write the flat version and move on. For the two prompts that draft and then prune a criteria list, see AI prompts for writing user stories.

Why does the generated list keep getting longer?

Because a model asked for acceptance criteria produces acceptance criteria, and nothing in the request rewards restraint.

Two habits keep the list honest. Prune as a separate instruction rather than a clause inside the drafting request, since a combined ask reliably returns a longer list rather than a tighter one. And delete anything that is already in your Definition of Done, because a standing bar repeated inside every story's criteria buries the two or three conditions specific to this story under a dozen that apply to everything.

If the pruned list still runs past a dozen items, the problem is upstream. That is a story carrying several goals, and the right move is to split it rather than to keep specifying it. The list length is a useful diagnostic precisely because it is visible before anyone estimates anything.

What about acceptance criteria for an AI feature?

Split them. The deterministic surface around a model, whether the disclosure renders, whether the response arrives inside your latency budget, whether the user can edit and resubmit, whether the action is logged, takes ordinary criteria and they work exactly as above.

The quality of the model's own output does not. The same input can produce different text twice, so a binary condition cannot describe it, and a criterion written as though it can will be either always true or always arguable. That part needs an eval set scored against a rubric, which is a different artefact with a different cadence, and it belongs beside the criteria rather than inside them. The same separation is what a quality gate in an AI pipeline formalises.

The honest limitation across all of this: criteria coverage is not test coverage. A story can satisfy every criterion on the list and still break something nobody wrote a criterion for, and a generated list is particularly prone to this because it covers the states you asked about thoroughly and the states you did not ask about not at all.

Where the prompting discipline is taught

Builders Camp's AI Prompting for Product bootcamp covers the pattern underneath this page directly: prompt structure that frames roles, context, constraints and success criteria, writing prompts for specs and stakeholder communication, and evaluation loops with quick rubrics so you can tell whether a prompt change improved the output or just changed it. It runs 1 week across 2 live sessions with 19 self-paced microlessons.

If the real problem is that stories arrive underspecified and late rather than that the criteria are weak, Project Management for Product covers planning, dependency management and the delivery cadence that refinement sits inside, and the Product Delivery Specialist Track bundles it with the wider execution skillset.

See the AI Prompting for Product bootcamp

A small practice that pays for itself: after a story ships, read its criteria against what actually broke. The gap between the two is the state checklist your next prompt should carry, and it is the only version of that list tuned to your product rather than to software in general.

Bootcamps referred in this Guide

Frequently asked questions

What is AI genuinely good at when writing acceptance criteria?

Enumerating states you forgot. Ask it for the empty case, the error case, the permission-denied case, the concurrent-edit case and the migration case, and it will list them faster and more completely than a tired person at the end of a refinement session. That is coverage, not judgement.

What is it bad at?

Choosing the bar. A model has no way to know whether two seconds is acceptable for your users or whether a 5 percent failure rate is tolerable for this action, so it either omits the number or invents a plausible one. Every threshold in the output should be a number you put there.

How do you tell whether a criterion is testable?

Name who or what observes it, and with what. If the answer is a person using judgement, it is a preference. If the answer is a query, a test, a stopwatch, or a screen anyone can look at and agree on, it is a criterion.

Should acceptance criteria always be in Given, When, Then form?

No. That form fits behaviour with a trigger and an outcome. A performance bound, an accessibility requirement or a data retention rule is a flat condition, and forcing it into a scenario shape makes it harder to check rather than easier.

How many acceptance criteria should a story have?

Enough to cover the states that change behaviour, and no more. A long list is usually a signal the story needs splitting, not that the criteria are thorough. If the list runs past a dozen, look at the story before you look at the criteria.

Is this different from a Definition of Done?

Yes. Acceptance criteria are specific to one story. A Definition of Done is the standing bar every story clears regardless of what it does, which is why duplicating your Definition of Done into each story's criteria is wasted effort and makes real criteria harder to find.

Can AI write criteria for an AI feature?

For the deterministic parts, yes. The quality of a model's output cannot be captured by a binary criterion, so that part needs an eval set scored against a rubric instead, and the criteria cover the things around it: the disclosure, the latency budget, the edit path, the audit log.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Tiago Pedro da Costa

Tiago Pedro da Costa

As Co-founder & CTO of Zumer, Tiago builds platforms that leverage AI to automate knowledge, improve collaboration, and accelerate sustainability in the construction industry. His work ranges from Abaqus, a platform for project and site management, to an AI-powered assistant supporting BREEAM certification.

As Co-founder & CTO of Zumer, Tiago builds platforms that leverage AI to automate knowledge, improve collaboration, and accelerate sustainability in the construction industry. His work ranges from Abaqus, a platform for project and site management, to an AI-powered assistant supporting BREEAM certification.

LinkedInMore guides by Tiago Pedro da Costa

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the AI Prompting for Product bootcamp