Tools
How to Write Evals for AI Products
An eval set is a fixed list of real inputs and the pass or fail rubric you score them against before any prompt or model change ships. Start by reading 20 to 50 real outputs from actual usage, not invented ones, write one binary check per failure you find, and weight the set toward the cases that already break your product. Run it before every change, not once at launch, or you are shipping AI features on vibes.
What is an eval set, exactly?
An eval set is a fixed list of real inputs, paired with what a correct or acceptable output looks like, that you score every time a prompt, model, or retrieval step changes. Think of it as the thing that stands in for "did this actually get better" once you no longer have a single obvious right answer to check against. A traditional feature ships when the code passes its tests. An AI feature ships when its outputs pass a rubric a person defined in advance, because the model's behavior is probabilistic, not fixed.
The set itself is not complicated. Eugene Yan's write-up on product evals describes the simplest version as a spreadsheet: one column for the input, one for the model's output, one for any context that helps a grader judge it, and one for the label, pass or fail. You do not need a framework to start. You need a frozen list of cases and a written definition of what passing means for each one.
Why are evals called the new PRD?
For an AI feature, the most precise specification you can write is an eval, because a paragraph describing good behaviour cannot be run and an eval can. Anthropic's documentation on defining success criteria puts the order plainly: "Building a successful LLM-based application starts with clearly defining your success criteria and then designing evaluations to measure performance against them." The same page asks for criteria that are specific and measurable, so "good performance" becomes something a test can check.
Used as a spec practice, the loop has four steps:
- Pick the hero use case. Name the one job the feature must do well for the launch to count, for example "answer a billing question from the customer's own invoice history."
- Write the ideal answer. For 10 to 20 real inputs to that job, write what a great response looks like, including what it must not do.
- Turn the answers into an eval. Each ideal answer becomes one or more pass or fail checks that a person, a code rule or a validated LLM judge can apply.
- Hill-climb. Change the prompt, retrieval or model, re-run the set, keep the change only if the pass rate goes up without new failures.
The document still has a job: it explains why this use case, for whom, and what the team decided not to build. What moves out of the document is the acceptance criteria, which is also why reviewing an AI written PRD ends with the eval set rather than the prose.
How many examples do you actually need to start?
Twenty to fifty real outputs is the floor. Hamel Husain and Shreya Shankar's AI evals FAQ describes a minimum viable setup as spending 30 minutes manually reviewing 20 to 50 outputs whenever you make a significant change, and recommends a larger pool of about 100 diverse traces when you are hunting for failure modes. It is worth taking seriously because it is written for teams that already shipped and got burned skipping this step. The number that matters more than the count is where the examples come from: real usage logs, real support tickets, or your own attempts to break the feature, not a list you imagined in a planning meeting. A handful of real, hard cases teaches you more than a hundred synthetic ones, because synthetic inputs rarely reproduce the exact way your product actually fails in front of a user.
Weight the set toward failures on purpose. If your AI feature is 90 percent accurate on easy inputs and falls apart on the remaining 10 percent, an eval set built from a random sample of traffic will mostly test the easy 90 percent, and you will feel confident right up until a user hits the gap. Go looking for the hard cases instead: the ambiguous request, the edge-case input, the thing a less capable model gets wrong that a more capable one might paper over.
What makes a good pass or fail criterion?
A pass or fail criterion is a written yes or no question, specific enough that two different people would grade the same output the same way. "Is this a good answer" is not a criterion. "Does this answer address what the user actually asked, without inventing a fact that is not in the source material" is a criterion, because it names the failure mode you are actually worried about.
Binary labels beat a 1-to-5 scale for the same reason a clear rubric beats a vague one: a scale lets a grader avoid the hard call in the middle, and the hard call in the middle is exactly the judgment you need someone to make out loud. Write the rubric before you look at any outputs, or you will unconsciously write it to match whatever the model already produced. Treat it the same way you would treat acceptance criteria in a PRD: specific enough that someone else could apply it without asking you what you meant.
| Grading method | Best for | What it costs | Where it breaks |
|---|---|---|---|
| Human review | Defining the first rubric, ambiguous or high-stakes failure modes | Slow, does not scale past a few hundred cases a week | Reviewer fatigue produces inconsistent grading over time |
| Code-based check | Objective, deterministic rules: does the JSON parse, is the citation present, is the tool call formatted correctly | Fast, cheap, runs on every commit | Cannot judge tone, factual accuracy, or anything requiring interpretation |
| LLM-as-judge | Scaling a rubric a human already validated, subjective quality checks at volume | Cheap and fast once set up, needs ongoing spot checks | Drifts from human judgment if the rubric or the underlying model changes and nobody re-checks agreement |
How do you read traces to find failure modes?
A trace is the full record of one user session: the input, any retrieved context and tool calls, and the output. Reading them is the habit that keeps an eval set honest, because the checks you write should come from failures you have seen, not failures you imagined in a planning meeting. Husain and Shankar's FAQ is direct about this: "Write evaluators for errors you discover, not errors you imagine."
The routine is simple enough to run in an afternoon:
- Read about 100 real traces. The FAQ recommends reviewing at least 100 and continuing past that while you are still learning something new. Write a short note on every trace that is wrong.
- Name the failure modes. Group the notes into named patterns, such as "answers from the wrong customer's invoice" or "promises a refund policy we do not have," and count how often each appears.
- Write one binary check per failure mode. One pass or fail question per pattern, never a single check that tries to judge everything. Eugene Yan makes the same point about LLM judges: one evaluator per dimension is easier to align than a single evaluator that grades many dimensions at once.
The strongest objection to "evals are the new PRD" comes from the same FAQ, which answers "Should I practice eval-driven development?" with "Generally no," because an LLM can fail in ways nobody predicts before building. Both positions hold once you separate the two kinds of check. The hero use case and its ideal answers are the target you write up front. The failure-mode checks come afterwards, from reading traces, and they are where most of the eval set's value ends up.
How do you run an eval set without slowing down every ship?
Automate it into the same place a prompt or model change already gets reviewed, the same way a test suite runs before a pull request merges. The eval run itself should take minutes: feed the frozen input list through the new prompt or model, score each output against the rubric, and flag anything that flips from pass to fail compared to the last run. That comparison, not the raw pass rate, is what tells you whether a change made things better or worse.
The part that actually takes time is building the first set and getting the rubric right, not running it afterward. Budget a real afternoon for that step rather than trying to improvise it the night before a launch, because a rushed rubric produces a set that passes everything and catches nothing.
Should a person or a model grade the output?
Start with a person, always, for any new failure mode. You cannot hand a rubric to an LLM judge until a person has proven the rubric actually separates good outputs from bad ones on real examples. Once the rubric holds up on real examples, an LLM judge can take over the repetitive grading at volume, scored with a clear yes or no question. Budget the labels honestly: Husain and Shankar's FAQ plans for 100 to 200 examples per failure mode, labeled by a trusted domain expert, before trusting a judge on a subjective failure, and it measures the judge by how many of the human-labeled failures it catches.
Keep checking the judge's agreement against a fresh batch of human-labeled cases every so often. An LLM judge that agreed with your team last quarter can quietly drift once the underlying model updates or your product's inputs shift, and nobody notices until the pass rate looks great and users are still unhappy.
When should you re-run evals and add new cases?
Re-run the set on every prompt change, retrieval change and model change, including a provider's model update you did not ask for. A new model can fix three failure modes and open two you had closed, and only a fixed set run before and after will show you which. Husain and Shankar's FAQ lists the triggers for a fresh round of trace reading: "Re-run error analysis when making significant changes: new features, prompt updates, model switches, or major bug fixes."
Every production failure is also a new eval case. When a user reports a bad answer, or a trace review turns one up, add that input and the correct behaviour to the set before you fix it, then confirm the fix turns it from fail to pass. Over a few months, the set becomes a record of every way the feature has broken, which is exactly the regression suite a probabilistic product needs. Hamel Husain's earlier post, Your AI Product Needs Evals, frames the same idea as three levels, unit tests, human and model review, and A/B testing, run at different cadences because each costs more than the one before.
What happens when you skip evals entirely?
You ship on vibes, which works exactly until the first prompt tweak that looks like an obvious improvement in the three examples someone tried by hand and turns out to break a case none of them thought to check. The AI Product Management bootcamp's own practical challenge is built around a version of this failure: an AI feature passes its stated accuracy threshold and still causes real harm, because the threshold measured the wrong thing and nobody had a rubric that would have caught it earlier. Evals do not prevent every failure. They prevent the specific failure of finding out from a user instead of from your own test run.
Who this guide is for, and who it is not for
Writing evals as described here suits a product manager, founder, or builder who has an AI feature already shipped or close to shipping and needs a repeatable way to know whether the next change made it better or worse, not a data scientist looking for a paper on evaluation metrics. It assumes you can write a specific yes or no question about what a good answer looks like for your product, because that judgment call cannot be automated away.
It is not for someone looking for a guaranteed accuracy number to hit before launch. No such number exists in the general case, and Builders Camp does not publish one because the right bar depends entirely on what a wrong answer costs your specific users. It is also not a substitute for talking to users directly. Evals check a rubric you already trust; user research is how you find out when the rubric itself needs to change.
Where evals fit next to the rest of the AI product job
Evals are one piece of a larger discipline that also includes designing what an AI agent is allowed to do on its own, deciding when retrieval-augmented generation actually improves an answer instead of just adding latency, and knowing how to build an AI product end to end rather than as a series of disconnected demos. If evals are the piece you are missing, Builders Camp's AI Product Management bootcamp walks through building an evaluation toolkit as one of its four live sessions, alongside the same responsible-AI patterns this guide only has room to summarize.
Builders Camp runs the AI Product Management bootcamp live over 2 weeks, 4 sessions, with a certification quiz and a practical challenge built around a realistic AI incident scenario, not a toy example.
Bootcamps referred in this Guide
Frequently asked questions
What is the difference between an eval and a QA test?
A QA test checks that code does what it was written to do. An eval checks that a model's output is still good, which is a judgment call, not a pass or fail on syntax. That is why an eval needs a rubric and a human-labeled reference set, where a QA test just needs an assertion.
How many examples should a first eval set have?
Hamel Husain and Shreya Shankar's evals FAQ describes a minimum viable setup as 30 minutes reviewing 20 to 50 real outputs whenever you make a significant change, pulled from actual usage, support tickets, or your own dogfooding, not invented from imagination. For finding failure modes they recommend a larger pool of about 100 diverse traces. A small set of real cases beats a large set of synthetic ones, because synthetic inputs rarely reproduce the specific way your product actually breaks.
Should a person or a model grade the output?
Use a person for the first pass on every new failure mode, because you need to define the rubric before anything can apply it. Use code for anything an objective rule can check. For a subjective failure mode, Husain and Shankar's FAQ plans for 100 to 200 examples labeled by a trusted domain expert before an LLM judge takes over, and it validates the judge by measuring how many human-labeled failures it catches.
What is a good pass mark for an AI feature?
There is no universal number. Builders Camp's own certification quizzes use a 70 percent pass mark as a working example of a bar that is neither trivial nor unreachable, but your feature's bar should come from what a wrong answer actually costs a user, not a borrowed number.
Can evals replace talking to users?
No. An eval set tells you whether a model's output meets a rubric you already believe in. Talking to users is how you find out the rubric itself is wrong. Skipping user research because you have evals running is a common way to ship a feature that is technically passing and still wrong.
Do evals slow down shipping?
A five-minute automated run before every prompt or model change is faster than the alternative, which is shipping a regression and finding out from a user complaint. The slow part is building the first eval set, not running it afterward.
What tools do people use to run evals?
Eugene Yan's write-up on product evals recommends starting with a spreadsheet: columns for input, output, useful metadata, and a pass or fail label. Move to a dedicated framework when running the set by hand starts to slow down changes. The tool matters less than having a frozen, labeled set to score against in the first place.
What does 'evals are the new PRD' mean?
It means the specification for an AI feature is a runnable test, not a document. You pick the use case that matters most, write the ideal answer for real inputs, turn those answers into pass or fail checks, and improve the prompt or model until the checks pass. Anthropic's own documentation starts LLM application work the same way: define success criteria, then design evaluations to measure against them.
How often should you read production traces?
Husain and Shankar's FAQ suggests reviewing at least 100 fresh traces each review cycle, with cycles of 2 to 4 weeks being typical in the teams they have seen, plus 10 to 20 traces a week in between, focused on outliers. Always re-run the review after an incident, a prompt change or a model switch.
Sources

Andre Albuquerque
CEO of Builders Camp, SuperOperator, and other companies. Building products.
CEO of Builders Camp, SuperOperator, and other companies. Building products.
LinkedInMore guides by Andre AlbuquerqueLast updated 2026-09-26
Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
Related guides
How to Design an AI Agent
Designing an AI agent means deciding what it can see, what it may use, when a human has to approve its action, and what...
Andre AlbuquerqueHow to Build an AI Product from Scratch
Building an AI product from scratch starts with naming what a wrong answer costs, not what a right one looks like...
Andre AlbuquerqueHow to Become an AI Product Manager
Becoming an AI product manager means adding an evaluation and guardrail layer on top of the core PM skills you already...
Andre AlbuquerqueRAG for Product Managers
Retrieval-augmented generation, or RAG, retrieves the most relevant passages from your own data at answer time and...
Andre AlbuquerqueHow to review an AI written PRD
Reviewing an AI written PRD means running four specific passes: provenance, edge cases, non goals and metrics. The...

Andre Albuquerque & Inês LourençoAI-Native vs AI-Enabled Products
Classify your product with one test: switch the model off. If the product still does its core job, it is AI-enabled...
Andre Albuquerque