Templates
AI prompts for writing user stories
Seven prompts cover the real user story workflow: extract, draft, split, interrogate the reason, draft criteria, prune criteria, and audit the batch against INVEST. Each one needs specific inputs, and each one has a failure it produces anyway, which is why the prompts are separate rather than combined into a single request.
Why seven prompts and not one good one?
Because the tasks pull against each other. Extraction wants breadth and no format. Splitting wants one item examined closely. Auditing wants the whole batch in view at once. A single prompt asked to do all three produces the average of three jobs, and the average is worse at each of them than a prompt written for one.
Everything below assumes you already know the three-part format and the INVEST checklist, which are covered in how to write a good user story. The workflow these prompts sit inside, and what to feed them, is how to write user stories with AI tools. This page is the prompts themselves, with the inputs each one needs and the failure each one still produces.
The extraction prompt withholds the format on purpose
Below is raw evidence from customer research. Do not write user stories yet.
List every distinct user need you can find. For each one give:
- the user segment, chosen only from this list: [your segments]
- the need, in the customer's own words where possible
- the exact line of evidence it came from
If a need appears in several places, say how many. If something reads
like a need but the evidence is one person's aside, mark it as thin.
Evidence:
[paste transcripts, ticket text, sales notes, verbatim]
Inputs it needs: the raw evidence, a closed segment list, and nothing else. Withholding the story template is the active ingredient. Given the format, a model starts producing well-shaped sentences and stops surfacing the observations that do not fit them, which are the ones worth having.
What it still gets wrong: it treats a loud single quote as a pattern. The "mark it as thin" instruction catches some of that, and you should still check the count yourself on anything you plan to build.
The thin markers are worth a second look rather than a delete. A need mentioned once by a customer in your largest segment is a different object from a need mentioned once by a trial user who churned the same week, and the model has no way to tell them apart. Sorting the thin list by who said it, before you decide what to drop, takes a minute and occasionally rescues the observation that turns out to matter most.
The drafting prompt is the least interesting one
Write this need as a user story in the format: As a [segment], I want
[goal], so that [reason].
Rules: the goal states an outcome, never an interface element. Match the
voice of these existing stories from our backlog: [paste 2 or 3 real ones].
Leave the reason clause blank if the evidence does not support one.
Need: [one item from the extraction output]
Inputs it needs: one extracted need and two or three real stories from your own backlog. Pasting real examples fixes tense, length, and segment naming in a single move, which is the practical form of few-shot prompting and more reliable than describing your house style in adjectives.
What it still gets wrong: it fills the blank reason clause anyway, about half the time, with something circular. Delete those rather than editing them.
The splitting prompt does the work drafting cannot
This story is probably more than one story. Do not rewrite it.
Identify every distinct goal inside it. For each, say whether it could
ship on its own and still be worth something to the user. Then propose
the smallest slice that is independently valuable, and say explicitly
what that slice does NOT include.
Story: [paste]
Inputs it needs: one story, and your honest sense of sprint length if you want the sizing opinion to mean anything. The Agile Alliance defines Independent and Small as separate tests within INVEST, and this prompt is aimed squarely at both. A model asked to split an existing item performs noticeably better than a model asked to generate correctly sized items up front, because splitting is a judgement about one concrete thing rather than a constraint applied while writing.
What it still gets wrong: it splits along technical seams, frontend and backend, rather than along value. A slice nobody would pay for is not a slice, and you have to reject those yourself.
When the output splits badly, name the axis in a follow-up rather than asking again. There are three that produce valuable slices reliably: by step in the workflow, where the first slice covers the most common path end to end and later slices add the branches; by data variation, where the first slice handles one record type and later slices add the rest; and by rule complexity, where the first slice implements the simple rule and later slices add the exceptions. Telling the model which axis to use turns a vague retry into a specific instruction, and it is a judgement you are better placed to make than it is, because you know which path is actually the common one.
The reason clause prompt is an interrogation
Attack the reason clause in this story. Answer three questions directly:
1. Who is measurably worse off if we never build this?
2. Does the reason restate the goal in different words? Quote the overlap.
3. What evidence in [paste your evidence] supports the reason, if any?
If the reason survives all three, say so plainly and stop.
Story: [paste]
Inputs it needs: the story and the evidence it came from, in the same message. Ask a model to write a reason and you get a reason. Ask it to attack one and you get information.
What it still gets wrong: it is agreeable under pressure. Push back twice and it will abandon a correct assessment, so treat the first answer as the real one.
Two prompts for criteria, because drafting and pruning are different instructions
Write acceptance criteria for this story in Given, When, Then form.
Cover the empty state, the error state, and the permission-denied state
as separate criteria. Do not restate the story sentence.
Story: [paste]
Then, as a separate message:
Delete every criterion above that cannot be checked independently of the
story sentence, or that duplicates another. Return only the survivors,
and say what you removed and why.
Inputs both need: the story and, ideally, your existing conventions. The Given, When, Then structure comes from Gherkin, whose reference describes it as a way of writing behaviour as concrete examples, and it is what makes each criterion mechanically checkable rather than merely agreed.
What they still get wrong: the first prompt writes criteria that restate the story in fresh vocabulary. The second one is what removes them, and combining the two into one request reliably produces a longer list instead of a tighter one.
Given, When, Then is also the wrong shape for some criteria, and the prompt will force it anyway. A performance bound, an accessibility requirement, or a data retention rule is a condition rather than a scenario, and wrapping it in a Given clause makes it harder to verify rather than easier. When the output contorts itself to fit the template, take that as a signal to write the criterion as a flat statement and move on, rather than accepting a scenario that nobody will ever run.
The audit prompt reads the batch, not the story
Here are [n] stories for one sprint. Do not improve them.
Produce a table with one row per story and these columns: Independent,
Negotiable, Valuable, Estimable, Small, Testable. Mark each pass or fail
with a five-word reason. Then answer two questions below the table:
which stories depend on the same unbuilt thing, and are any two of these
the same story written twice?
Stories: [paste all]
Inputs it needs: the whole batch in one message. This is the prompt where a longer context window genuinely changes the output quality, because the dependency question cannot be answered one story at a time.
What it still gets wrong: it marks Valuable as a pass whenever a reason clause exists, regardless of whether the reason is any good. Run the interrogation prompt on anything it passes too easily.
The last prompt finds the story nobody wrote
Given this set of stories and this system boundary: [paste both], list
the user-facing states that appear nowhere in the set. Consider at least:
first-run with no data, partial failure, permission denied, cancellation
mid-flow, and what happens to existing users when this ships.
For each gap, say whether it needs its own story or belongs in an
existing one.
This is the prompt that earns its place fastest. The states it names are cheap to specify, easy to forget, and the ones a team discovers in production. Its weakness is the mirror image: it will propose a story for every edge it can imagine, including ones that do not exist in your product, so the output is a list to prune rather than a list to add.
What all seven have in common
- They take evidence or an artefact as input, never a description of one.
- They name the thing they must not do, because a model told only what to do will do more of it.
- They end with an instruction to stop, which is what stops a helpful answer becoming a longer one.
The honest limitation is that none of these prompts improve your discovery. A well-prompted batch built on two interviews is a well-formatted version of two interviews. The speed is real; the evidence problem is unchanged, and a fast drafting loop makes weak evidence easier to hide rather than easier to see.
Where a shared prompt standard is taught
Builders Camp's AI Prompting for Product bootcamp covers prompt structure that reliably works, research and synthesis prompts, writing prompts for specs and stakeholder updates, and evaluation loops with quick rubrics so you can tell whether a change to a prompt actually improved the output. It runs 1 week across 2 live sessions with 19 self-paced microlessons, and it names cross-functional teams rolling out a shared prompting standard as one of its audiences, which is the setting where a prompt library beats individual skill.
If the bottleneck is the backlog rather than the prompting, Product Manager Foundations covers prioritisation and scoping inside the wider product process, and the Product Delivery Specialist Track covers the planning and dependency work a well-split backlog exists to support. Prompt engineering is the underlying skill all seven of these lean on.
See the AI Prompting for Product bootcamp
Version the prompts the way you version anything else that produces output you act on. When a batch comes back worse than last sprint's, the first question worth asking is which of the seven changed, and that question is unanswerable if the prompts live in chat histories.
Bootcamps referred in this Guide
Frequently asked questions
Why seven prompts instead of one good one?
Because the tasks pull in opposite directions. Extraction wants breadth and no format. Splitting wants one item examined closely. Auditing wants the whole batch at once. A single prompt asked to do all three averages them, and the average is worse at each job than a prompt written for it.
What do these prompts need pasted into them?
Raw evidence for the extraction prompt, a closed list of the user segments you are allowed to name, and your system boundaries. The later prompts take the output of the earlier ones. None of them work well on a summary you wrote yourself, because a summary already contains your conclusions.
Should I paste an example story into the prompt?
Two or three real ones from your own backlog, yes. Showing the model the house style it should match is more reliable than describing that style in adjectives, and it fixes tense, length, and segment naming in one move rather than three corrections.
Do these prompts work in any AI tool?
They are plain text with no tool-specific syntax, so they run anywhere. What changes between tools is context length and whether the tool can read your existing backlog directly, which matters most for the audit prompt where the batch is the input.
What is the reason clause prompt actually testing?
Whether anyone is worse off if the story is never built. It is written as an interrogation rather than a generator, because asking a model to write a reason produces a reason, and asking it to attack one produces the information you need.
How do I stop the acceptance criteria prompt restating the story?
Run a second prompt over its output that deletes any criterion that does not name a condition checkable independently of the story sentence. Pruning is a separate instruction from drafting, and combining them produces a longer list rather than a tighter one.
Do I still need to know the user story format myself?
Yes. Every prompt here assumes you can tell a good output from a plausible one, and the model is consistently better at plausible. The format and the INVEST checklist are the standard you are judging against, not something the prompt supplies.
Sources

Andre Albuquerque
CEO of Builders Camp, SuperOperator, and other companies. Building products.
CEO of Builders Camp, SuperOperator, and other companies. Building products.
LinkedInMore guides by Andre Albuquerque
Inês Lourenço
CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.
CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.
LinkedInMore guides by Inês LourençoLast updated 2026-09-18
Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
Related guides
How to Write a Good User Story
A good user story follows the format 'As a [user], I want [goal], so that [reason],' passes the INVEST checklist...
Andre AlbuquerqueWhat Is Prompt Engineering?
Prompt engineering is the practice of structuring the instructions you send a language model, role, context...
Andre AlbuquerqueWhat Is Few-Shot Prompting?
Few-shot prompting means putting a few worked examples, typically 3 to 5, into the prompt before the real input, so the...
Andre AlbuquerqueChatGPT for product managers
ChatGPT for product managers works best as a set of Projects, each with its own instructions and attached files, rather...
Andre Albuquerque

