Tools
How to use AI for design feedback that is not just taste
Ask for a review against Jakob Nielsen's 10 usability heuristics, one row per heuristic, pass or fail, with the specific element named as evidence. Refined in 1994 from an analysis of 249 usability problems, the set is the closest thing design has to a checkable standard. Leave the severity column empty, because that number is yours to fill in.
Why a named standard beats asking what it thinks
Ask a model what it thinks of a screen and you get an essay: three things that work, three suggestions, a closing sentence about balance. Every line of it is unfalsifiable, which means none of it survives contact with a designer who disagrees. Ask for a review against Jakob Nielsen's 10 usability heuristics and every observation arrives with a category and a claim attached, and a claim is something two people can check together.
The set is worth knowing because it is stable. Nielsen refined the 10 heuristics in 1994 from a factor analysis of 249 usability problems, and Nielsen Norman Group states they have remained unchanged since, with the article itself expanded in 2020. Thirty years of stability in an area that mostly runs on opinion is the reason to anchor a review to it rather than to whatever the model feels about the screen.
The ten, and what each one asks of a screen
Visibility of System Status, Match Between the System and the Real World, User Control and Freedom, Consistency and Standards, Error Prevention, Recognition Rather than Recall, Flexibility and Efficiency of Use, Aesthetic and Minimalist Design, Help Users Recognize, Diagnose, and Recover from Errors, and Help and Documentation.
Read as a review checklist they turn into ten specific questions. Does this screen tell the user what is happening right now, in language borrowed from their world rather than from the database? Can they get out of a state they entered by accident? Does it look and behave like the rest of the product? Does it prevent the mistake rather than explain it afterwards? Is anything required to be remembered from a previous screen? Is there a faster route for the person doing this for the fortieth time? Is anything on screen competing for attention with the thing that matters? When something fails, does the message say what happened and what to do? And is help available at the moment it is needed rather than in a document nobody opens?
Nine of those ten questions can be asked of a static screen. That is the practical reason this works as a prompt.
The prompt shape that produces findings instead of prose
One row per heuristic. A verdict of pass, fail or not applicable. For every fail, the specific element it refers to, quoted or named. No positive findings, no summary paragraph, no recommendations in the same pass.
The evidence requirement is what does the work. A model asked for a review will produce a fluent observation about visual hierarchy whether or not the screen has a hierarchy problem, because fluent observations are what it is for. A model asked to name the element it means either names one you can go and look at, or produces a row you can immediately discard. That single constraint cuts most of the noise.
Banning praise matters more than it sounds. Asked for balanced feedback, a model manufactures the balance, and you end up reading three invented compliments for every real finding. You are not running this review to feel good about the screen.
- One row per heuristic, so nothing is skipped and nothing is invented outside the frame.
- A named element for every fail, so each row is checkable in under ten seconds.
- No severity column, because that number is yours and filling it is the review.
Screenshot or source, and why the answer is both
A screenshot and a multimodal model give you the review of what is visible: what the eye lands on, whether the labels read as the user's language or the database's, whether two actions are competing for the same attention. That review cannot tell you the error state was never implemented, because the error state is not in the picture.
The source gives you the opposite. It shows which states exist, whether inputs carry real labels, whether images have alt text, and what contrast the declared colour tokens produce against the WCAG 2.2 Level AA minimum of 4.5:1 for normal text. Running that pass against code is covered in detail in running a design review with Claude Code, and it is the cheaper of the two to automate because the answers are arithmetic.
Run the screenshot pass for the heuristics that turn on perception, and the source pass for the ones that turn on existence. Neither is a substitute for the other and both take about five minutes.
The three findings to throw away without reading twice
A hierarchy comment with no element named. "The visual hierarchy could be clearer" is the design-review equivalent of a horoscope: true of most screens, actionable on none.
A row that restates the heuristic instead of applying it. If the Error Prevention row says that preventing errors is better than handling them, the model had nothing to say about your screen and filled the cell anyway. Delete it rather than arguing with it.
A severity rating. Nielsen's 0 to 4 scale runs from cosmetic problem to usability catastrophe and combines three factors: how often the problem occurs, how hard it is for users to overcome, and whether it persists once they have learned about it. A model knows none of those for your product. It will still produce a confident 3, and that confident 3 is the single most dangerous output in the review because it looks like a decision.
What survives, and what you do with it
The rows that survive are specific, mechanical and dull, which is exactly right for a first pass: a submit button with no disabled state during submission, a filter that silently resets on navigation, an error message naming an internal field, a destructive action with no undo. Nobody needed a model to be clever to find those. They needed something to look at all ten categories without getting bored at the fourth.
Then you fill in the severity yourself, using the three factors, and the argument about which finding is a 3 and which is a 1 is the conversation worth having with a designer. Usability Testing for Product Managers covers that translation over 1 week with 1 live session and 6 microlessons, including how to write findings as issues with severity ratings and clear recommendations rather than a transcript of what you noticed.
Prototyping with AI covers the other end, over 1 week with 2 live sessions and 9 microlessons: generating the UI concepts and the prototype copy for onboarding and empty states, which is where several of these findings should have been caught before the screen was built at all. Prototyping for Product Managers covers choosing the fidelity that makes a given question answerable, in 1 week with 1 live session and 7 microlessons.
See the Prototyping with AI bootcamp
For the work upstream of the review, mapping the flow first prevents a class of finding entirely, and assembling the testable version in Figma is where the surviving issues get fixed. Run the heuristic pass on a screen that shipped last quarter before you run it on one in review: the findings land harder when nobody is defending the screen, and they tend to be the same three categories failing everywhere in the product.
Bootcamps referred in this Guide
Frequently asked questions
Why review against a heuristic set instead of asking what the model thinks?
Because an open question produces prose and a named standard produces findings you can check. Jakob Nielsen's 10 usability heuristics, refined in 1994 from a factor analysis of 249 usability problems and unchanged since, give every observation a category and a claim. A finding filed under Visibility of System Status either points at a missing status signal or it does not, and you can look.
What are the 10 heuristics?
Visibility of System Status, Match Between the System and the Real World, User Control and Freedom, Consistency and Standards, Error Prevention, Recognition Rather than Recall, Flexibility and Efficiency of Use, Aesthetic and Minimalist Design, Help Users Recognize, Diagnose, and Recover from Errors, and Help and Documentation. Nielsen Norman Group publishes the full set with examples.
Should I send a screenshot or the code?
Both answer different questions. A screenshot lets a multimodal model comment on what is visible: hierarchy, labelling, what is competing for attention. The source lets it check what exists: states, labels, alt text, declared colour values. A screenshot review cannot tell you the error state was never built, and a source review cannot tell you the page looks crowded.
How do I stop it from flattering the design?
Ban the positive findings from the output and require evidence for the negative ones. Ask for a table with one row per heuristic, a verdict of pass, fail or not applicable, and for every fail the specific element it refers to. A model asked for balanced feedback will manufacture praise to balance with, and that praise costs you reading time for nothing.
Which findings should I distrust?
Anything about hierarchy without an element named, anything that restates the heuristic definition instead of applying it, and any severity rating the model assigned itself. The first two are padding. The third is a judgement about your users and your business that nothing in the prompt gave it the information to make.
Can it rate severity?
It can produce a number, and the number is not grounded. Nielsen's 0 to 4 scale combines frequency, impact and persistence, and a model knows none of the three for your product. Take the findings, keep the severity column empty, and fill it yourself in the review. That column is the part of the review where product judgement actually lives.
Does this replace a design review with a designer?
It replaces the first fifteen minutes of one. Heuristic violations are the mechanical layer, and clearing them before a designer looks means their attention goes to hierarchy, craft and whether the screen does the job. Builders Camp's Usability Testing for Product Managers covers what comes after, turning observations into severity ratings and prioritised recommendations, over 1 week with 1 live session and 6 microlessons.
Sources

Andre Albuquerque
CEO of Builders Camp, SuperOperator, and other companies. Building products.
CEO of Builders Camp, SuperOperator, and other companies. Building products.
LinkedInMore guides by Andre Albuquerque
Ricardo Luiz
He is an accomplished Product Director, bringing a wealth of experience in driving innovation, building high-performing teams, and fostering collaborative environments.
He is an accomplished Product Director, bringing a wealth of experience in driving innovation, building high-performing teams, and fostering collaborative environments.
LinkedInMore guides by Ricardo LuizLast updated 2026-09-18
Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
Related guides
How to run a design review with Claude Code
Give Claude Code the spec and the component file, and ask for a table of every requirement with present or absent and a...

Andre Albuquerque & Ricardo LuizFigma for product managers
Figma for product managers means using its prototyping mode to connect static frames into a clickable flow you can test...
Andre AlbuquerqueWhat an AI wireframe generator is actually good for
A generated wireframe settles two things: what is on the screen and in what order. It settles nothing about whether the...

Andre Albuquerque & Ricardo LuizHow to use AI for user flows without losing the edge cases
A model will draft a plausible happy path in under a minute, which is the boring half of the work and worth automating...

Andre Albuquerque & Ricardo Luiz

