---
title: "How to Review an AI Written PRD Properly"
description: "An AI written PRD fails in four places: invented requirements, missing edge cases, an empty non goals section, and a metric chosen after the launch happened."
canonical_url: "https://builderscamp.com/guides/other/how-to-review-an-ai-written-prd"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Inês Lourenço"
publisher: "Builders Camp"
guide_class: "other"
---

# How to review an AI written PRD

**TL;DR:** Reviewing an AI written PRD means running four specific passes: provenance, edge cases, non goals and metrics. The draft will be fluent everywhere, including in the places where nobody has any evidence, so the review has to supply the doubt the document does not signal. Budget more time than you would for a human draft of the same length.

## What actually goes wrong in an AI written PRD?

Not the prose. A generated PRD is usually better organised and better written than a hurried human one, which is exactly what makes reviewing it harder: the document gives you no signal about where it is guessing. Google Cloud's definition of a hallucination covers the mechanism plainly, a model producing incorrect or misleading output because of gaps and biases in what it learned from, and the output carries no marker separating those completions from the ones grounded in what you actually supplied.

So the review is not a proofread. It is four passes, each looking for a different kind of gap, and each one faster than it sounds once you know what you are looking for.

The four failure modes are consistent enough to name in advance: requirements nobody asked for, edge cases nobody described, a non goals section left empty, and a success metric that cannot be measured until after the thing has shipped.

## The provenance pass: every requirement traces to a person or it goes

Take the requirements section and, line by line, name the source. A support ticket, an interview, a named stakeholder, a legal constraint, an analytics figure. Anything that traces to nothing is a candidate for deletion, not because it is necessarily wrong but because nobody has yet decided it is right.

This pass catches the specific thing a model does well and dangerously: completing the pattern. Feed it a spec for a billing feature and it will produce requirements about invoice export, dunning emails and tax handling, because those appear in billing specs. They may be good ideas. They are not, on the evidence in front of you, requirements, and shipping them costs the same engineering weeks as the ones somebody actually asked for.

The efficient version of this pass is to run it with the model itself first. Ask it to annotate each requirement with the exact input it derived from and to list the ones it inferred. It is reliable at that task in a way it is not reliable at judging its own quality, because the question has a checkable answer.

## The edge case pass: read for what the document does not mention

A generated PRD writes the happy path well and the exceptions barely at all. Go looking for four specific absences: the empty state, the partial failure, the concurrent action, and the user who is in an unusual but legal position (no permissions, expired plan, one record instead of a hundred). Then ask whether the spec says anything at all about each.

Builders Camp's AI Product Management bootcamp builds its practical challenge around what happens when nobody ran that pass. An AI triage feature has been live for six months at a primary care platform, routing around 1,200 patient intakes a week. It clears the accuracy threshold set at launch, 81 percent against a 75 percent bar. It also deprioritises patients from the lowest income quartile at 2.3 times the rate of the highest, after controlling for symptom severity, because patients with less exposure to clinical vocabulary describe symptoms in plain language and the model learned from four years of human triage decisions that contained the same pattern. The number that passed is real. The number that mattered was never in the spec.

That case is deliberately uncomfortable, and the point of it is narrower than "AI is risky." Every requirement in that original PRD could have been written by a careful human. The gap was a category of user behaviour nobody thought to describe, which is precisely the category a model cannot invent on your behalf.

## An empty non goals section is the tell, not a clean bill of health

Generated drafts are long where elaboration is rewarded and thin where refusal is. Requirements, user stories and background come back rich. Non goals comes back as two vague lines or nothing at all, because the model has no way of knowing what you decided not to do.

Treat that section as the review's load-bearing check. For every requirement in the document, ask what was considered and dropped. If the answer is nothing, the scope conversation has not happened yet and the PRD is documenting an intention rather than a decision. The [PRD template](https://builderscamp.com/guides/templates/prd-template) treats the same section as the strongest defence against scope creep for the same reason: it is the only part of the document that makes the trade visible.

## A metric you can only measure after launch is not a success metric

The last pass is the shortest. Find the success metrics, and for each one ask two questions: what is it today, and who computes it. A generated PRD will happily write "increase activation rate" with no baseline, which means whatever happens after launch can be narrated as a win.

For an AI feature this pass goes further, and Builders Camp's AI Product Management curriculum states the shift directly: evals replace acceptance criteria and the launch gate. A requirement like "the assistant should give accurate answers" is unreviewable prose until it becomes a labelled set of traces with expected behaviour attached, which is what [writing evals for AI products](https://builderscamp.com/guides/tools/how-to-write-evals-for-ai-products) covers in practice. Until that set exists, you are reviewing a description of a feature rather than a definition of done.

## Who this is for, and who it is not for

This fits a product manager whose team has started drafting specs with a model and has noticed that review now takes longer, not less, which is the honest cost of the change. Builders Camp's AI Product Management bootcamp covers the surrounding discipline across 2 weeks and 4 live sessions, from defining good with evals to designing the environment an agent operates in, and its framing is that the deliverable has moved from the document to the harness around it. People who want that path sequenced alongside prompting, agents and shipping usually take it inside the AI Product Expert Track.

It is not the right page if your actual problem is that nobody on the team writes specs at all. That is a process gap, and Product Manager Foundations covers the end-to-end product process it sits inside. It is also not a substitute for engineering review: a PRD that survives all four passes can still be infeasible, and the person who tells you that is not you.

## Keep the annotations after the review ends

The habit worth forming is not the review itself, it is what you keep from it. Leave the provenance annotations in the document as a short list at the bottom: these requirements came from research, these were inferred and accepted anyway, these were inferred and cut. Six weeks into a build, when someone asks why the spec says what it says, that list is the difference between a decision you can defend and a sentence that has quietly become a fact because it has been in the document a long time. A [Product Requirement Prompt](https://builderscamp.com/guides/glossary/product-requirement-prompt) inherits this problem too, since an agent reading the spec will treat an inferred line exactly as seriously as an observed one.

[See the AI Product Management bootcamp](https://builderscamp.com/bootcamps/ai-product-management?utm_source=guide&utm_medium=organic&utm_campaign=how-to-review-an-ai-written-prd)

For why a fluent draft can still be wrong, see [AI hallucination](https://builderscamp.com/guides/glossary/ai-hallucination). For the document being reviewed, the [PRD template](https://builderscamp.com/guides/templates/prd-template).

## Frequently asked questions

### How long should reviewing an AI written PRD take?

Longer than reviewing a human one of the same length, at least the first few times. A human draft signals its own uncertainty through hedged sentences and unfinished sections. A generated draft arrives uniformly finished, so you have to supply the doubt yourself, section by section.

### What is the fastest way to spot an invented requirement?

Ask who asked for it. Every requirement should trace to a person, a ticket, a research session or an explicit constraint. A requirement that traces to nothing is usually the model completing a pattern it has seen in other specs, not a need anyone expressed.

### Why do generated PRDs miss edge cases so consistently?

Because edge cases come from knowing the system, not from knowing what a PRD looks like. A model writes the happy path fluently because the happy path is what most specs describe. The empty state, the partial failure and the concurrent edit are absent from the draft in the same way they are absent from most of its training material.

### Is an empty non goals section a real problem?

It is the single most reliable tell. A generated PRD tends to fill sections that reward elaboration and leave thin the one section whose value is refusal. If nothing has been ruled out, nobody has made a scope decision yet.

### Should I ask the model to review its own PRD?

As a first pass, and never as the last one. Asking it to list every claim it cannot trace to your input is genuinely useful. Asking whether the document is good produces agreement, because agreeableness is cheaper for it to produce than a real objection.

### Does a PRD for an AI feature need a different review?

Yes, on one axis: an AI feature's requirements cannot be verified by reading them. You need an eval set, a labelled collection of inputs and expected behaviour, before the requirement means anything, because 'the model should prioritise urgent cases correctly' is a sentence, not a test.

### What do you do when the generated draft is genuinely good?

Ship it, and keep the provenance annotations. The risk is never that a generated PRD is badly written. It is that six weeks later nobody remembers which lines were observed and which were inferred, and the inferred ones get defended as if they had evidence behind them.

## Sources

- [Google Cloud: What are AI hallucinations?](https://cloud.google.com/discover/what-are-ai-hallucinations)
- [Builders Camp: AI Product Management bootcamp](https://builderscamp.com/bootcamps/ai-product-management)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
