Builders Camp

Other Guides

The AI product requirements document guide

An AI product requirements document runs eight sections, and each has to prove something rather than describe it: that the task tolerates a probabilistic answer, that good is defined as an eval set, that failures are enumerated by kind, that autonomy has a boundary, that the economics survive per use cost, that the data is yours to use, and that the product still works with the feature turned off. Prose alone cannot carry any of those.

What does an AI PRD have to prove that an ordinary one does not?

That the feature can be judged at all. A conventional requirement is checkable by reading it: either the invoice loads in one click or it does not. Cagan and Nika name the reason this breaks for AI features directly, generative systems are probabilistic rather than deterministic, so the same inputs can produce different outputs and weights shift over time. "The assistant answers accurately" is not a requirement. It is a hope with a grammar.

So the document changes shape. Each section below carries a burden of proof rather than a description, and the test for whether a section is finished is whether somebody could disagree with it using evidence. Builders Camp's AI Product Management curriculum frames the same shift in a sentence worth stealing: the deliverable has moved from the document to the environment around it, the evals, the harness and the business design that captures the value.

Eight sections. Each one below states what it has to establish before the next one means anything.

The job section has to prove the task tolerates being wrong

Start by describing the job the model does in one paragraph, then immediately answer the question that decides everything downstream: what happens when it is wrong, and can the experience absorb that?

Cagan and Nika draw the line with two examples that make the point faster than an argument. A personalised news feed can survive a recommendation that misses, because the interface can manage it. A system controlling an insulin dose cannot, because a value outside medical guidelines is unacceptable at any frequency. Most real features sit somewhere between, and the spec's job is to say where, explicitly, before anyone builds.

If the honest answer is that the task cannot tolerate an error, the document should say so and propose the version that can, usually by moving the model from deciding to suggesting.

The definition of good is an eval set, not a sentence

This is the section that replaces acceptance criteria, and it is the one most commonly written as prose and left there. A labelled set of real inputs with the expected behaviour attached is a thing you can run against a change. A paragraph describing quality is not.

Builders Camp's AI Product Management bootcamp puts the method in its own curriculum: generate and label your own traces, cluster the failures you find, and turn those clusters into repeatable tests. The practical detail that matters is that the set comes from your traffic rather than from imagination, because the failures you can imagine are rarely the ones that show up. Writing evals for AI products covers the mechanics.

State the pass bar as a number in the document, and state what it does not cover in the next sentence. An 81 percent score against a 75 percent bar is a real result and it says nothing about which 19 percent failed or who they were.

The failure section has to name kinds of wrong, not a rate

One error rate hides everything useful. Break the failures into kinds and attach the product's response to each, because they are different problems with different fixes.

A generative feature typically fails in at least three distinct ways: it omits something true, it asserts something false, or it declines a request it should have handled. Google Cloud's description of hallucination covers the middle one, output that is incorrect or misleading because of limits and biases in what the model learned, and it is the failure most likely to be trusted by a user, because a confident wrong answer looks exactly like a confident right one. Omission is quieter and often worse in a summarisation product. Over-refusal costs you nothing in accuracy metrics and a great deal in adoption.

Write the handling decision next to each: what the user sees, whether the system retries, whether anything is logged for review. That list is the actual specification of the feature's behaviour, more than the description of the happy path above it.

The autonomy section has to say where a human stands

Every AI feature makes a claim about how much it is allowed to do on its own, and the claim gets made whether or not the spec states it. Write it down as three lines: what the system may do without asking, what requires a human to approve first, and what it may never do.

Anthropic's guidance on building effective agents is useful here for its restraint. Its three stated principles are to keep agent design simple, to show planning steps rather than hiding them, and to invest in the interface between the agent and its tools through documentation and testing. The recommendation running through it is to add autonomy only when simpler approaches demonstrably fall short, which is close to the opposite of how most specs are written, since a spec tends to describe the most capable version of the idea rather than the smallest one that works.

Builders Camp's AI Agents bootcamp makes the same point from the failure side, with a three agent pipeline that sends a false churn alert on a company's largest account because no approval step sat between a classification and a message to a customer success lead. Guardrails are the implementation of this section; the section itself is the decision.

The economics section has to survive the per use cost

Accuracy is bought with money and latency. A more accurate approach usually means more data, more processing, a slower response, or all three, and those costs land on the user experience and the unit economics at the same time.

Put three numbers in the document: expected cost per use, acceptable latency at the ninetieth percentile, and the volume you expect at steady state. Then multiply. A feature that is excellent and costs more per active user than the plan they are on is not a quality success waiting for scale, it is a pricing decision nobody has made yet.

The data section has to prove you are allowed to use it

Name every source the feature reads or learns from, the permission you have for each, and what you know about its limits. Cagan and Nika are direct that product managers need a clear understanding of the training data and how a model was trained and tuned, and that every large dataset carries potential biases and limitations. Treating that as a data science concern is how a bias that was already in your historical decisions gets automated and scaled.

The uncomfortable version of this section names what you cannot check. Write that down too.

The rollback section has to prove the product works without the feature

Last, and shortest, and skipped most often. Describe what the product does with the feature turned off: which screens change, what a user sees instead, whether anything already generated stays visible. Then confirm someone can flip it without a deploy.

If the answer is that the main flow breaks, you do not have a rollback plan. You have a launch with no exit, and the decision to accept that should be made deliberately in a document rather than discovered at 11pm by whoever is on call.

Who this is for, and who it is not for

This fits a product manager specifying an AI feature for a team that will actually ship it. Builders Camp's AI Product Management bootcamp covers the surrounding craft across 2 weeks and 4 live sessions, from evals and agent environments to the business design that captures the value, and its practical challenge puts you in front of a live triage feature that passed its accuracy bar and failed a bias audit anyway. People who want the sequence rather than a single bootcamp usually take it inside the AI Product Expert Track.

It is the wrong document if you are still deciding whether AI belongs in the feature at all. That question comes first, and how to build an AI product is a better starting point than a spec template.

Write the section you are least able to fill in first

Whichever of the eight sections you would rather postpone is the one carrying the risk. For most teams it is the eval set, because building one means admitting nobody has defined good yet and that defining it will take a week of labelling. Start there anyway. Every other section in the document is cheap to rewrite later, and that one determines whether any of the rest can be checked at all.

See the AI Product Management bootcamp

For the agent design layer underneath a feature like this, see how to design an AI agent and the AI Product Expert Track.

Bootcamps referred in this Guide

Frequently asked questions

Why does an AI feature need a different document?

Because the usual requirement format assumes a deterministic system. Cagan and Nika put it plainly: generative AI is probabilistic, not deterministic, so the same input need not produce the same output. A requirement written as 'the system does X' stops being checkable the moment that is only true most of the time.

What is the first question the document has to answer?

Whether the task tolerates being wrong sometimes. A personalised feed can absorb an occasional bad recommendation inside the experience. A system calculating an insulin dose cannot. That judgement decides whether the rest of the spec is worth writing.

Can acceptance criteria still be used for an AI feature?

For the deterministic parts, yes: the button exists, the response renders, the audit entry is written. For the model's own behaviour, no. That part needs an eval set, a labelled collection of inputs with expected outcomes, because a sentence describing good behaviour cannot be run.

What is a failure taxonomy?

A list of the distinct ways this feature can be wrong, each with the product's response attached. Not a single error rate. A summariser that omits a fact, invents a fact, or refuses a valid request has three different problems, and they call for three different handling decisions.

Should the spec include cost per use?

Yes, as a budget with a number in it. Accuracy trades against cost and latency, and that trade shapes the user experience and the unit economics together. A feature that is right 99 percent of the time and costs more per call than the customer pays per month is a business problem disguised as a quality result.

How do you specify the data the feature runs on?

Name the sources, the permission you have to use each one, and what you know about their limitations. Every large dataset carries biases and gaps, and the product manager is expected to understand how those might show up in the finished product rather than treating training data as someone else's concern.

What does the rollback section have to prove?

That the product still works with the feature off. If turning it off breaks the main flow, you have no rollback, you have a hope that nothing goes wrong. Write down what users see in the off state before launch, not during the incident.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Inês Lourenço

Inês Lourenço

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

LinkedInMore guides by Inês Lourenço

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the AI Product Management bootcamp