Builders Camp

Tools

How to verify AI-generated code as a product manager, before it merges or reaches a decision

Verify AI-generated code by asking for evidence instead of a summary: the test output, the build result, a screenshot of the change and the list of files touched, then set review depth by what breaks if the code is wrong. In the Stack Overflow 2025 Developer Survey, 66 percent of developers named AI solutions that are almost right, but not quite, as their biggest frustration, which is exactly the failure a quick read does not catch.

Why is AI-generated code harder to catch when it is wrong?

A broken formula announces itself. It returns a negative revenue figure, or a date in 1900, or nothing at all, and you notice in the first second. Agent output fails differently: it produces a fluent, well-structured, correctly formatted answer that happens to be untrue, and every surface signal you normally use to judge quality is intact. The code compiles. The summary is confident. The diff looks tidy.

Developers see this constantly. In the Stack Overflow 2025 Developer Survey, 66 percent of the 31,476 people who answered the question named "AI solutions that are almost right, but not quite" as a frustration, the most common one on the list, and 45 percent said debugging AI-generated code takes more time. The survey measures professional developers, not product managers, so it says nothing about how often your own agent is wrong. What it does show is that the people closest to the code treat "almost right" as the normal failure, not the rare one.

Builders Camp's Building with Claude Code practical challenge is built on the same failure: a junior engineer accepts a diff that passes all the tests and breaks staging, because the agent was confident and nobody had defined what done actually meant. The scenario is teaching material rather than an incident report, so treat it as an illustration rather than evidence about frequency. The mechanism it shows is that confidence and correctness are separate properties, and only one of them is visible by default.

What counts as evidence, and what only looks done?

Anthropic's own guidance for Claude Code puts it plainly: "Have Claude show evidence rather than asserting success." The best practices page lists the test output, the command Claude ran and what it returned, or a screenshot of the result, and adds that reviewing evidence is faster than re-running the check yourself. For a product manager, that turns review from reading code you may not be fluent in into reading an evidence packet you can judge.

Ask for What it proves What it does not prove
The exact test command and its pasted output The tests that exist pass right now That the tests cover the behaviour you care about
Build, typecheck and lint output The code is syntactically and structurally sound That it does the right thing
A screenshot or short recording of the changed screen The visible behaviour matches what you asked for Anything off the happy path
The list of files touched, with one line per file on why The change stayed inside the scope you expected That the untouched files did not need changing
A plain-language story of the diff The agent can explain its own change That the explanation matches the code

The last row is the one to be suspicious of. A diff story is useful because a mismatch between the story and the file list is easy to spot, not because the story is true. "Looks done" is a statement about the summary. Evidence is a statement about the software.

How much review does a change actually need?

Less than you fear for most changes, and far more than you want for a few. Match the depth to what breaks if the code is wrong, not to how long the diff is. A one-line change to a payment rule deserves more attention than four hundred lines of new internal dashboard.

Risk level Example Minimum review before merge
Low: reversible, internal, no data written Copy change, a new chart on an internal page Evidence packet, glance at the screenshot, merge
Medium: user-facing, reversible New onboarding step, a filter on a public list Evidence packet, run the change yourself on a preview, read which tests were added or changed
High: money, access, deletion, anything irreversible Pricing logic, permissions, a migration that drops a column Evidence packet, a human engineer reads the diff, a named person approves, a rollback plan exists before merge

Write the tier down before the agent starts, not after it finishes. If the agent is working on a high-risk change, say so in the task, and ask for the packet in that tier's shape. Deciding the depth after you have seen a clean-looking summary is how a high-risk change gets a low-risk review.

What are you actually reviewing: the claim, the path, or the source?

Three different things, and conflating them is why review often feels thorough and catches nothing. The distinction applies to code and to the analysis an agent produces around it.

The claim is the sentence that will reach a decision. "The checkout bug is fixed" is a claim, and so is "pricing confusion accounts for 37 percent of tickets this quarter." Reviewing it means checking it against something you can reproduce independently, which is usually cheaper than it sounds.

The path is how the agent got there: which files it read, which filter it applied, what it did with the ambiguous rows or the failing test. You review the path when the claim is not directly checkable, which is common in synthesis work where there is no ground truth to compare against.

The source is whether the underlying material says what the summary says it says. For code, the source is the diff itself, and the classic failure is an agent that changes a test to match its code rather than the code to match the test. For analysis, it is a summary that merges two customers' complaints into one theme. Neither is hallucinating. Both are quietly optimising for "the check passes."

A worked example: the agent says 37 percent of tickets mention pricing

You have handed Claude Code an export of 220 support tickets and it has come back with six themes, percentages for each, and a recommendation. Here is the review that takes fifteen minutes and is worth doing every time.

  1. Recompute one slice. Filter the export to a single week and count the pricing mentions yourself. If the agent's proportion holds roughly on that slice, the total is plausible. If it does not, you have disproved the number without checking the other 200 rows. The asymmetry is the whole value: a matching slice is weak evidence, a mismatched slice is conclusive.
  2. Ask what it counted as a mention. Does a ticket saying "this is more than I expected to pay" count? Does one asking how to change a billing address? The definition is where most of the error lives, and the agent made a choice here whether or not you gave it one.
  3. Pull three quotes and read the originals. Ask for the ticket identifiers behind one theme, then open those tickets. You are checking whether the theme exists in the source or was assembled from adjacent complaints that sounded similar.
  4. Name the decision, then ask what would change it. If the recommendation is to reprice, ask what proportion would have to be true for the recommendation to flip. If the answer is "anything above 20 percent", then a number between 20 and 37 changes nothing and you can stop reviewing. If the answer is "anything above 35", you are one rounding decision away from a different quarter and the review is not optional.

The fourth step saves the most time, because it tells you how much precision the decision actually needs before you spend effort chasing it.

Why is a rule in CLAUDE.md not a control?

Anthropic's memory documentation is direct about this: CLAUDE.md files and auto memory are loaded as context, not as enforced configuration, and the more specific and concise the instruction, the more consistently Claude follows it. Follows, not obeys. A rule in a context file raises the probability of the behaviour you want and never reaches certainty.

That distinction has a practical consequence for how you write rules. Anything where a violation is recoverable belongs in the context file, because a strong prior is enough and the cost of being occasionally ignored is a correction. Anything where a violation is not recoverable belongs in a hook, which is a command that runs at a fixed point and does not negotiate. The hooks reference documents exit code 2 from a PreToolUse hook as the signal that blocks a tool call. "Prefer the house tone in release notes" is a context rule. "Never write to the production configuration file" is a hook. The guide to CLAUDE.md vs skills vs hooks vs subagents sorts the rest of your rules the same way.

Claude Code's own defaults work the same way. The security documentation says that in Manual mode it starts read-only, asks before editing files or running commands that can modify your system, and cannot write outside the folder it was started in without explicit permission. Those are boundaries on blast radius, not checks on correctness. They stop the agent from doing damage you did not authorise. They have nothing to say about whether the code it just wrote is right.

The honest limitation: this does not scale, and pretending otherwise is the trap

A fifteen-minute review per output is fine at three outputs a week and impossible at thirty. The instinct at that point is to review less carefully, which is the worst available option because it keeps the appearance of a control while removing the control. The two real options are to review a sample properly, accepting that some unchecked output reaches decisions, or to convert the repeatable parts of the review into something automatic.

The second is what a quality gate is: an automated checkpoint that verifies a specific, measurable condition before work advances. The build passes. No file outside the agreed folders changed. Every claim carries a ticket identifier. None of those checks judge whether the work is smart, and that is the point: they clear the mechanical failures so your fifteen minutes goes to the judgment call rather than to counting fields.

Where is the review discipline taught?

Building with Claude Code lists hooks, CI/CD and quality gates among its six public learning topics, taught by Guilherme Salgueiro across 2 live sessions, and its practical challenge asks you to define done as a set of objective acceptance criteria, a quality gate and a stop condition rather than the agent's confidence. That definition is the practical core of everything above.

AI Product Management moves the same thinking to products you ship to customers, where evals replace acceptance criteria and the launch gate: you generate and label your own traces, cluster the failures, and turn them into repeatable tests, across 4 live sessions. AI Agents covers the adjacent question of how much autonomy to grant in the first place, including autonomy levels, approvals and constraints for high-stakes actions, and monitoring, which is the design decision that determines how much review you will be doing later.

Decide the review before you read the output

The move that changes the most is deciding, before you look, what would make you reject it. Write the rejection condition down: a test that must appear in the diff, a file that must not change, a number outside a range, a theme with no traceable source. Then read. A review that starts after you have already been persuaded by the prose is not a review, it is a second reading of a document you liked the first time.

See the Building with Claude Code bootcamp

For the failure mode this whole discipline exists to catch, what an AI hallucination actually is covers why fluent output and correct output come apart, and how to write evals for AI products covers turning a manual check into a test you can run every time.

Bootcamps referred in this Guide

Frequently asked questions

What counts as evidence that AI-generated code works?

Something you can read that the agent could not have produced by guessing: the test command and its pasted output, the build or typecheck result, a screenshot of the changed screen next to the design, and the list of files touched compared with the files you expected. A summary saying the feature is done is a claim, not evidence.

Is a green CI run enough to merge?

Only for the behaviour the tests actually check. CI proves the existing tests pass; it says nothing about a path nobody wrote a test for, and an agent that also wrote the tests may have tested its own assumptions. Read which tests changed in the diff before you trust the green tick.

How do I check a number an agent produced without redoing the work?

Recompute one slice by hand. If the agent reports that 37 percent of tickets mention pricing, filter the export for a single week and count. A matching slice does not prove the total, but a mismatched slice disproves it immediately, and that asymmetry is what makes spot-checking worth the five minutes.

Does putting the rule in CLAUDE.md stop the agent from breaking it?

No. Anthropic's documentation states that CLAUDE.md files and auto memory are treated as context, not enforced configuration. They make compliance likely, not certain. To block an action regardless of what Claude decides, use a PreToolUse hook, which runs as a shell command and does not negotiate.

What does Claude Code block on its own?

In Manual mode it starts read-only, asks before editing files or running commands that could modify your system, and can only write inside the folder it was started in and its subfolders. It runs a built-in set of read-only commands such as ls, cat and git status without asking. Those are safety boundaries, not correctness checks.

Should a product manager review the reasoning or just the output?

The output first, because that is what reaches a decision, then the path for anything you cannot verify directly. If the conclusion is checkable in two minutes, check it and skip the transcript. If it is not checkable, the path is the only evidence you have, and an answer whose path you cannot follow should not carry a decision.

When is a second agent a better reviewer than a person?

When the check is mechanical and repeatable: format conformance, missing fields, a rule that every claim cites a source. A separate agent that did not produce the work has no stake in defending it, which is the argument for role separation. It is a poor substitute for a person on anything requiring judgment about customers or money.

Which Builders Camp bootcamp covers this?

Building with Claude Code, taught by Guilherme Salgueiro across 2 live sessions, lists hooks, CI/CD and quality gates among its six public learning topics, and its practical challenge asks you to define done as evidence rather than the agent's confidence. AI Product Management covers evals for products you ship to customers.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Inês Lourenço

Inês Lourenço

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

LinkedInMore guides by Inês Lourenço

Last updated 2026-09-27

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Building with Claude Code bootcamp