Builders Camp

Tools

How to run a design review with Claude Code

Give Claude Code the spec and the component file, and ask for a table of every requirement with present or absent and a file and line as evidence. Missing states are the most common finding, and contrast against the WCAG 2.2 minimums of 4.5:1 for normal text and 3:1 for large text is the cheapest check to automate. What it cannot judge is whether the screen is any good, which is still the review that matters most.

What is an agent actually reviewing here?

The code, not the screen. That distinction sets the boundary for everything below. Claude Code reads the component file, the spec you wrote, the styles and the copy, and reports what those files do and do not contain. It does not see the rendered page, so a review it runs is a review of implementation completeness and mechanical accessibility, not a review of whether the thing looks right.

This is less limiting than it sounds, because most of what gets caught late in a design review is mechanical. The error state nobody built. The empty state that ships with placeholder text. The button label that says Continue in the spec and Next in the build. Those all live in the source, and finding them before a designer opens the branch is worth about an hour of somebody's week.

Start from the spec, not from a screenshot

A screenshot shows one state. A specification lists all of them, which is why the comparison runs in that direction. Point the session at both files and ask for a table: one row per requirement, a column for present or absent, and a column with the file and line that proves it.

The evidence column is the part that makes this trustworthy. A review that says "the error state is implemented" is an assertion. A review that says "the error state is implemented, CheckoutForm.tsx line 142" is a claim you can open and verify in four seconds, and the difference shows up immediately in how often the claim is wrong.

Ask for absences explicitly too. Models are better at confirming what exists than at noticing what is missing, so the prompt should name the four states you expect on any screen that loads data, and require a verdict on each: loading, empty, error, success. That framing is what turns a vague review into a list with holes in it.

Which accessibility checks are worth running every time?

Contrast and labels, because both are arithmetic rather than opinion. W3C's WCAG 2.2 sets a minimum contrast ratio of 4.5:1 for normal text and 3:1 for large text at Level AA, with a separate 3:1 minimum for user interface components and graphical objects such as icons and input borders. Given two colour values, that check has a single right answer, and a machine is better at it than you are.

Labels are the other one. Every form input needs a programmatic label rather than a placeholder standing in for one, every meaningful image needs alt text, and every decorative image needs an explicitly empty alt rather than a missing attribute. All three are visible in the markup and all three are commonly missed, because nothing looks wrong on screen when they are absent.

  • Contrast, computed from the declared colour tokens, against the two WCAG minimums above.
  • Programmatic labels on every input, checkbox and select, with placeholders not counting.
  • Alternative text present on meaningful images and deliberately empty on decorative ones.

Those three cover a real share of what an audit would flag, and they cost one instruction per screen. They are a starting point, not a substitute for testing with assistive technology and real users.

What the agent will get confidently wrong

Computed contrast is only true if the token it computed is the token that renders. A colour set on a parent element, a utility class applied conditionally, or a theme that swaps values at runtime will all produce a correct calculation of the wrong pair. Treat a contrast pass as a strong signal and a contrast failure as a thing to verify in the browser before you file it.

Anything that depends on layout is out of scope entirely. Whether the primary action falls below the fold on a small screen, whether a long product name wraps into three lines and pushes the price out of view, whether the focus ring is actually visible against the background it lands on: none of that is knowable from source. That work belongs in a browser, on the devices your users actually hold.

For the design tooling on the other side of this handoff, Figma for product managers covers the file that the spec usually came from, and v0 covers generating the screen in the first place rather than checking it afterwards.

When in the cycle should this run?

On the branch, before the designer looks at it, and before anyone calls a review meeting. That placement is the whole economic argument. A designer's attention is expensive and finite, and spending the first fifteen minutes of it discovering that the error state was never built is the worst possible use of it. Run the mechanical pass first so the human review starts on a screen where the states exist, the copy matches and the labels are present.

The same logic applies to an AI-generated prototype, which is where a lot of screens now start. A prototype built in an afternoon from a prompt has the same failure pattern as a rushed implementation: the happy path is complete and everything else is absent. Checking a prototype built in Lovable against the spec before you put it in front of a user costs one instruction and prevents the test session where three of five participants get stuck on an unhandled error.

What does not work is running this after merge. At that point the findings arrive as bug tickets competing with everything else in the queue, which is how a two minute fix becomes a three week wait. The value of the check is entirely a function of when it runs.

What a PM should review that no tool catches

Whether the screen does the job. An implementation can match the spec completely and still be wrong, because the spec was wrong, and that is the failure mode a checklist is structurally incapable of catching. Read the empty state copy and ask whether it tells somebody what to do next or just reports that there is nothing here. Read the error message and ask whether a user could act on it. Look at where the eye lands first and whether that is the action you want taken.

Builders Camp's Prototyping with AI bootcamp spends 1 week on this half, across 2 live sessions and 9 microlessons: user flows and information architecture, UI concepts, and prototype copy for onboarding and empty states specifically. The point of prototyping fast is to have this argument before the screen is built, which is cheaper than having it in review.

Turn the output into severity, not a list

A review that hands a designer twenty flat observations gets sorted by whoever reads it, and they sort it badly because they are guessing at your intent. Give every finding a severity and a recommendation instead: blocks the task, slows the task, or cosmetic inconsistency. Three levels is enough, and the argument about which level something belongs in is the useful conversation.

Usability Testing for Product Managers covers that translation directly over 1 week and 6 microlessons, including how to write findings as issues with severity ratings and clear recommendations rather than as a transcript of what you noticed. Claude Code for Product Managers covers the workflow that produced the raw findings, in 1 week with 2 live sessions and 8 microlessons, including how to validate agent output before acting on it.

Run it on a screen that already shipped

The first pass is more informative on something live than on something in review, because nothing is at stake and nobody is defending it. Take a screen that went out three months ago, hand over the original spec and the component, and read the absent column. The states that turn out to be missing on a screen everyone considered finished are a better argument for adding this step than any list of benefits, and they are usually the same three states missing everywhere else.

See the Claude Code for Product Managers bootcamp

For the build and validate halves around it, Prototyping with AI runs 1 week on flows, UI concepts and microcopy, and Usability Testing for Product Managers covers task design, moderation and turning findings into decisions.

Bootcamps referred in this Guide

Frequently asked questions

What can an agent verify on a built screen?

Things that exist in the code: whether the loading, empty, error and success states are implemented, whether the copy matches the spec word for word, whether form inputs have programmatic labels, whether images carry alt text, and what contrast ratio the declared colour tokens produce. It reads the source, so anything the source states is checkable and anything only visible when rendered is not.

What can it not check?

Whether the screen is any good. Visual hierarchy, whether the primary action sits where attention already is, whether the flow makes sense to somebody who has not read the spec, and whether the empty state says something useful or just says no results. Those are judgements, and an agent that offers an opinion on them is generating prose, not evidence.

Which accessibility rules are worth automating first?

Contrast and labels. W3C's WCAG 2.2 sets a minimum contrast ratio of 4.5:1 for normal text and 3:1 for large text at Level AA, and a separate 3:1 minimum for user interface components and graphical objects. Those are arithmetic once you know the two colours, which makes them exactly the kind of check worth running on every screen instead of occasionally.

Why start from the spec instead of a screenshot?

Because a screenshot only shows one state, and missing states are the most common finding in any design review. A spec lists what should exist; the code shows what does. Comparing those two produces a table of present and absent, and the absent rows are usually the error handling nobody built.

How should findings be written up?

As issues with a severity and a recommendation, not as a list of observations. An issue that blocks a task ranks above one that slows it down, which ranks above a cosmetic inconsistency. A review that hands a designer twenty flat bullet points gets triaged by whoever reads it, which means it gets triaged badly.

Does this replace a designer or an accessibility audit?

Neither. It catches the mechanical failures that waste a designer's review time, so the human review starts from a screen where the states exist and the labels are present. A full accessibility audit involves assistive technology and real users, and no static code check substitutes for that.

Which Builders Camp bootcamp fits this work?

Prototyping with AI covers the build side across 1 week, 2 live sessions and 9 microlessons, including user flows, UI concepts and prototype microcopy for onboarding and empty states. Usability Testing for Product Managers covers what happens after the review, turning observations into severity ratings and prioritised recommendations, in 1 week.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Ricardo Luiz

Ricardo Luiz

He is an accomplished Product Director, bringing a wealth of experience in driving innovation, building high-performing teams, and fostering collaborative environments.

He is an accomplished Product Director, bringing a wealth of experience in driving innovation, building high-performing teams, and fostering collaborative environments.

LinkedInMore guides by Ricardo Luiz

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Claude Code for Product Managers bootcamp