---
title: "How to Write User Stories With AI Tools"
description: "Draft user stories with AI in three passes: feed it real evidence, force a split, then rewrite the reason clause yourself. Here is what the model gets wrong."
canonical_url: "https://builderscamp.com/guides/templates/how-to-write-user-stories-with-ai"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Inês Lourenço"
publisher: "Builders Camp"
guide_class: "templates"
---

# How to write user stories with AI tools

**TL;DR:** Draft user stories with AI in three passes: give it raw evidence rather than your summary, make it propose a split before it polishes anything, then rewrite the reason clause yourself. The model is reliably good at the user and the goal and reliably bad at why the work matters, which is the clause that decides whether the story should exist.

## What is AI actually good at here?

Turning a pile of raw evidence into candidate users and candidate goals, fast enough that you notice the gaps in your evidence while they are still cheap to fill. It is unreliable at the third clause, the reason, and that is the clause that decides whether the story deserves a place on the backlog at all.

This page covers the drafting workflow. The craft itself, the three-part format, the INVEST checklist, and what a real acceptance criterion looks like, is [how to write a good user story](https://builderscamp.com/guides/templates/how-to-write-a-user-story), and the rest of this assumes you already know it. For the literal prompts with their inputs and failure modes, see [AI prompts for writing user stories](https://builderscamp.com/guides/templates/ai-prompts-for-writing-user-stories).

## What do you feed it before it writes anything?

Raw material, not your summary of the raw material.

Paste the interview transcript section, the support ticket text with the customer's own wording intact, the funnel step with its number, the sales note as it was written. A model handed your summary writes stories about your summary, which returns your existing conclusions in a new format and feels productive while adding nothing. A model handed the underlying evidence occasionally produces a story you did not expect, and that story is the whole reason to run the exercise.

Give it the constraints too, in the same message. The segments it may use, as a closed list taken from your actual research. The system boundaries, so it does not write stories for a surface your team does not own. And the current sprint's theme, so the output is a backlog slice rather than a wish list. Left open on any of the three, the model fills the gap with something plausible, and plausible is the specific failure mode worth guarding against.

## The three passes

**Pass one, extract.** Ask for candidate users and candidate goals only, with the evidence line each one came from, and no story format yet. Withholding the template at this stage matters more than it sounds: given the format, the model starts producing well-shaped sentences and stops surfacing the messy observations that do not fit them. You want the mess first.

**Pass two, split.** Now hand it the extracted list and ask which items are actually more than one story, and where each one breaks. This is the pass that pays for itself. The Agile Alliance's own framing of INVEST puts Small and Independent as separate tests for good reason, and a model asked to split is noticeably better than a model asked to write well-sized stories in the first place, because splitting is a judgement about one concrete thing rather than a constraint applied while generating.

**Pass three, rewrite.** Take the split list and write the reason clause yourself, one story at a time. The model's version tends to be circular: "so that I can filter the list" as the reason for "I want a filter on the list." That sentence passes a glance and fails the only test that matters, which is whether anyone is worse off if you never build it.

## What does AI reliably get wrong?

Five things, and they repeat across tools rather than being specific to one.

It invents personas when you leave the segment list open, and the invented persona then quietly sets the scope of everything downstream. It writes the interface into the goal clause, producing "I want a dropdown showing my saved cards" where the GOV.UK Service Manual's rule applies exactly: describe the need, not the solution. It merges goals with "and", which is the same failure a hurried human makes and just as easy to spot. It produces stories of suspiciously uniform size, because nothing in the prompt rewards the honest answer that one item is three times bigger than the rest. And it writes acceptance criteria that restate the story in different words instead of naming a separately checkable condition, which is worth pruning against a real [acceptance criteria](https://builderscamp.com/guides/glossary/acceptance-criteria) standard rather than accepting as a bonus.

None of these are subtle. All of them survive a casual read, which is why the review pass has to be deliberate rather than a skim.

## How do you review twenty stories without reading twenty stories?

Review the batch as a batch, before you read any single item closely.

Three passes over the whole set catch most of what goes wrong, and each one takes under a minute:

- **Dependency sweep**: does more than one story wait on the same unbuilt thing? If so, that thing is the real first story and nothing else is independent yet.
- **Size spread**: are they all the same size? A uniform batch usually means the model normalised the estimate rather than the work.
- **Reason audit**: read only the reason clauses, in a list, with nothing else. Circular ones become obvious the moment they sit next to each other, and they are invisible when each story is read alone.

Only then read the survivors individually.

## What you should not delegate

The reason clause, the split decision when it is genuinely contested, and the call that a story should not exist.

That last one is the part a model will never volunteer. Asked to produce stories, it produces stories, including for work that should not be on the backlog. Nothing in the interaction rewards the answer "there is no story here, the evidence does not support one", so that answer has to come from you. A model that generates twelve items from an input that honestly supports four is not malfunctioning; it is doing what it was asked.

The honest limitation on all of this is that a faster drafting loop makes weak discovery easier to hide, not easier to spot. Twelve well-formatted stories built on two interviews look more finished than two rough stories built on the same two interviews, and they are not. The speed is real and the evidence problem is unchanged.

## Does this work for a team, or only for one PM?

It works better for a team, on one condition: the prompt lives in a shared file rather than in individual chat histories. When everyone uses the same extraction and splitting prompt, an odd story is a signal about the story. When everyone has their own, an odd story is a signal about whose prompt made it, and nobody can tell the two cases apart.

Builders Camp's AI Prompting for Product bootcamp is built around exactly that standard: prompt structure that reliably works, research and synthesis prompts for turning messy notes into decisions, writing prompts for specs and stakeholder communication, and evaluation loops with quick rubrics so you can tell whether a prompt actually improved. It runs 1 week across 2 live sessions with 19 self-paced microlessons, and its stated audience includes cross-functional teams rolling out a shared prompting standard rather than individuals improving alone.

If the underlying problem is the backlog rather than the drafting, Product Manager Foundations covers prioritisation and scoping across the wider product process, and the Product Delivery Specialist Track goes further into planning, dependencies, and the delivery cadence that a well-split backlog exists to serve.

[See the AI Prompting for Product bootcamp](https://builderscamp.com/bootcamps/ai-prompting-for-product?utm_source=guide&utm_medium=organic&utm_campaign=how-to-write-user-stories-with-ai)

One habit worth adding once the workflow settles: keep the rejected stories in a separate file with the reason each was cut. Six weeks later, when the same request arrives from a different stakeholder, that file answers it in ten seconds instead of restarting the argument from the beginning.

## Frequently asked questions

### What part of a user story should AI actually draft?

The first two clauses and the split. A model is good at turning a paragraph of evidence into a candidate user and a candidate goal, and good at proposing where an oversized story breaks apart. The reason clause is the part you write, because it encodes a judgement about why the work matters and the model has no access to that.

### What do you have to give it before it writes anything?

Raw evidence rather than your summary of it: interview quotes, support ticket text, a funnel step with a number, the actual feature request. A model handed a summary produces a story about your summary, which is a restatement of what you already decided rather than a draft you can test.

### Why do AI-drafted stories keep failing the INVEST check?

Mostly on Independent and Small. A model asked for stories from one input will happily produce items that all depend on the same unbuilt foundation, or one story containing three goals joined by 'and'. Both are visible in seconds once you run the check, and neither is visible if you read the stories one at a time.

### Does AI write good acceptance criteria along with the story?

It writes plausible ones, which is the problem. The generated criteria tend to restate the story in different words rather than state a separately checkable condition. Treat them as a starting list to prune, not as the finished set.

### How do you stop the model from inventing users?

Name the segments it may use, in the prompt, as a closed list drawn from your actual research. Left open, it will produce a plausible persona that nobody in your product has ever been, and that persona will then quietly shape the scope of the story.

### Is it faster than writing stories by hand?

Faster at the first draft, slower at the review, and the review is where the value sits. The gain is not typing speed. It is that a fast draft makes the gaps in your evidence visible early, while they are still cheap to go and fill.

### Should the whole team use the same prompt?

Yes, and keep it in a shared file rather than in individual chat histories. A shared prompt makes the output comparable across people, which means a story that looks odd is a signal about the story rather than about whose prompt produced it.

## Sources

- [Agile Alliance: What does INVEST stand for?](https://agilealliance.org/glossary/invest/)
- [GOV.UK Service Manual: Writing user stories](https://www.gov.uk/service-manual/agile-delivery/writing-user-stories)
- [Atlassian: User stories with examples and a template](https://www.atlassian.com/agile/project-management/user-stories)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
