---
title: "How to Use AI for Backlog Refinement"
description: "Triage a stale backlog in batches with AI: the delete rule, the inputs that make clustering work, a worked pass over 400 tickets, and the calls you keep."
canonical_url: "https://builderscamp.com/guides/tools/ai-for-backlog-refinement"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Tiago Pedro da Costa"
publisher: "Builders Camp"
guide_class: "tools"
---

# How to use AI for backlog refinement on a backlog nobody has read in a year

**TL;DR:** A model can read 400 backlog tickets in one pass and return duplicate clusters, dead items and tickets too thin to size, which is a week of human work compressed into an afternoon of review. The rule that matters is the delete rule: if nobody can name the customer, the problem or the decision the ticket would change, it gets closed rather than rewritten.

## Why does a stale backlog resist human refinement?

Because refinement scales linearly and backlogs do not. The Scrum Guide describes Product Backlog refinement as an ongoing act of breaking items down and adding detail such as description, order and size, which is fine at the rate of a few items a week. It says nothing about the 400 tickets your team accumulated across two reorgs, three product renames and a departed PM who wrote every title as a feature name with no body.

That pile has a specific shape. Roughly, it is duplicates written by different people, tickets describing a screen that no longer exists, requests whose customer has since churned, and a genuine minority of good items buried among them. Reading it end to end takes a person somewhere between two days and never, and the honest outcome for most teams is never. A model reading all 400 at once is the first tool that changes the arithmetic rather than the willpower required.

## What does a batch pass actually look like?

Work in batches of 50 to 100 tickets, scoped to one component or one theme. For each batch, paste the full ticket bodies, the created date and the reporter, plus three or four sentences describing what the product does today. Then ask for four outputs, separately, in this order:

1. Duplicate and near duplicate clusters, with the item IDs and the sentence in each that makes them a match.
2. Tickets describing behaviour that the product description says no longer exists.
3. Tickets too thin to size, with the specific question each one fails to answer.
4. Everything left over, grouped by the underlying problem rather than the proposed solution.

That fourth output is the one people skip and the one that pays. Fifteen tickets asking for a different filter on the same table are one problem, not fifteen. Grouping by problem rather than by requested solution is the move that turns a backlog into something a roadmap can be built from, and it is mechanical enough that a model does it reliably. Atlassian's own guidance on user stories makes the same case from the writing side: a story states a problem for a user, not a feature name.

## The delete rule, stated plainly

A ticket gets closed, not refined, when nobody can name the customer, the problem, or the decision it would change.

That is three independent tests and failing any one of them is enough. "Add dark mode to settings" has a solution but no named customer or problem. "Improve onboarding" has a problem so vague that no decision follows from it. "Investigate the Stripe thing" has neither. None of those is refinable, because refinement means adding detail and there is no detail to add, only invention.

The counterweight is real: teams delete tickets and then discover the ticket was the only record of a regulatory commitment somebody made in a meeting in March. Handle that by making deletion a batch with a reason column rather than a series of individual judgement calls. The model produces the candidate list with one line of justification each, a human reads the reasons rather than the tickets, and anything that makes a reviewer hesitate goes back to the refine pile. A 200 item delete list reviewed on reasons takes about 20 minutes. Reviewed ticket by ticket, it never gets done.

## A worked pass over roughly 400 tickets

Assume a backlog of 400 items, seven components, two years old. Split into batches by component, run the four outputs above, and expect a rough shape like this.

| Category | Typical share | What you do with it |
|---|---|---|
| Duplicate or near duplicate | 15 to 25 percent | Merge into the clearest one, close the rest |
| Describes something that no longer exists | 10 to 20 percent | Close with a reason |
| Too thin to size | 20 to 30 percent | Delete under the rule, or send back to the requester |
| Real problem, badly written | 20 to 30 percent | Rewrite, this is the refinement work |
| Ready as written | under 10 percent | Order it |

Those percentages are the shape of one exercise, not a benchmark, and your split will differ with how disciplined your intake has been. What they show reliably is the ratio nobody expects: the majority of a stale backlog is not work waiting to be done, it is a record of conversations. Once you see that, the refinement question stops being "how do we get through this" and becomes "what do we stop letting in".

## Where the model is confidently wrong

Duplicate detection on titles alone. Two tickets titled "export not working" can describe the CSV export on the billing page and the PDF export on the reporting page, written eight months apart by different people. A model merging on title similarity throws away one real bug and leaves a misleading merged ticket that will confuse whoever picks it up.

The fix is input shape, not prompting. Pass full bodies, require the model to quote the matching sentence from each ticket in the cluster, and treat any cluster where the quoted sentences are not obviously about the same thing as a non match. This is the same evidence discipline that keeps a drafted plan honest in [sprint planning](https://builderscamp.com/guides/glossary/sprint-planning-ceremony), and the same reason an [AI hallucination](https://builderscamp.com/guides/glossary/ai-hallucination) is a manageable risk when every claim arrives with its source attached.

## Which judgements stay with the team?

Ordering, and what gets cut. A model can group by problem and flag what is unwritable. It has no view on which of two real problems your company should solve next quarter, because that depends on strategy, on a commercial commitment nobody wrote in a ticket, and on what the team can absorb given everything else in flight.

Keep three calls human: what gets deleted (the model proposes, you approve the batch), what order the survivors sit in, and whether a rewritten ticket is ready to start. The Scrum Guide ties that last one to your Definition of Done, which is local to your team and invisible to any model that has not been shown it.

## What a refined backlog changes downstream

The payoff is not a tidier tracker. It is that planning stops being an archaeology exercise. Beyond Agile argues for keeping the rituals that produce a feedback loop and killing the ones that only produce documents, and a refinement session where the team spends 40 minutes deciphering tickets is the second kind. After a batch pass, planning starts from items that already state a problem and a customer, which is when the ceremony becomes a decision rather than a reading.

Project Management for Product runs one week on the delivery side of this: slicing work without false certainty, managing dependencies across teams, and writing status that names blockers and decisions rather than activity. It sits inside the Product Delivery Specialist Track for anyone who wants the full path rather than a single bootcamp.

[See the Project Management for Product bootcamp](https://builderscamp.com/bootcamps/project-management-for-product?utm_source=guide&utm_medium=organic&utm_campaign=ai-for-backlog-refinement)

Once the backlog is readable, the next questions are ordering and sizing: [AI for prioritization](https://builderscamp.com/guides/tools/ai-for-prioritization) covers scoring what survives, [AI for sprint planning](https://builderscamp.com/guides/tools/ai-for-sprint-planning) covers the commitment itself, and the [Product Delivery Specialist Track](https://builderscamp.com/tracks/product-delivery) bundles the delivery skills around both.

## Frequently asked questions

### What is AI good at during backlog refinement?

Reading everything at once. A model handles 400 tickets in a single pass and returns duplicate clusters, items whose described problem no longer exists in the product, and tickets too thin to size. Humans are better at every one of those judgements individually and hopeless at doing 400 of them in a row.

### What is the rule for deleting rather than refining a ticket?

Delete when nobody can name the customer, the problem or the decision the ticket would change. A ticket that survives only because someone might want it one day is a cost with no owner. Refine when the problem is real and only the description is bad.

### Should AI ever delete a ticket by itself?

No. Have it produce a delete candidate list with a one line reason per item, then close them in a batch a human has read. The reason column is what makes the review fast, and the review is what stops you closing the one ticket that was holding a compliance commitment.

### How many tickets should go in one batch?

Work in batches of 50 to 100 and keep each batch to one theme or one component. Bigger batches produce duplicate clusters that span unrelated areas and are harder to check than the tickets themselves.

### What inputs make the duplicate clustering actually correct?

Full ticket bodies rather than titles, the created date, the reporter, and a short description of what the product does today. Titles alone produce false merges, because two tickets can be worded identically and describe different screens.

### Does refinement in Scrum mean grooming meetings?

The Scrum Guide describes Product Backlog refinement as the ongoing act of breaking items down and adding detail such as description, order and size. It is an activity, not a scheduled ceremony, which is precisely why a batch pass done asynchronously fits it well.

### Does Builders Camp teach a specific refinement tool?

No. Project Management for Product covers slicing work, dependency management and delivery cadence without naming a product, so the method transfers to whatever tracker and model you already use.

## Sources

- [Scrum.org: The 2020 Scrum Guide](https://scrumguides.org/scrum-guide.html)
- [Atlassian: Product backlogs](https://www.atlassian.com/agile/scrum/backlogs)
- [Atlassian: User stories](https://www.atlassian.com/agile/project-management/user-stories)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
