Builders Camp

Tools

How to use AI for usability test analysis without losing the evidence

A model will code five transcripts against your code list in minutes and tally how many participants hit each problem, which is the frequency third of a severity rating. The other two thirds, impact and persistence, stay with you. Require a timestamp on every quote, and spot check five of them before anything reaches a stakeholder.

What is actually being automated here?

Coding, and counting. You run five sessions, you have five transcripts, and the job in front of you is tagging every passage where something went wrong and then working out which problems recurred. That work is mechanical, it is slow, and human coders drift: by the fourth transcript you are tagging slightly differently than you were in the first, usually toward whatever theory you formed in session two.

A model does not drift. Give it a fixed code list and five transcripts and it applies the same definitions to all of them, then reports how many participants hit each code. That count is worth more than it looks, because most teams estimate frequency from memory and memory over-weights the session that was most uncomfortable to watch.

The code list has to be yours, and it has to name behaviours

Asked to find the themes, a model clusters by topic. You get back onboarding, navigation, pricing clarity: areas, not problems, and an area cannot be fixed. Write the code list yourself before the first session, from the test plan, and phrase every code as a behaviour somebody could observe.

"Applied a filter without noticing it persisted to the next screen" is a code. "Navigation confusion" is a folder. The first one tells an engineer what to change; the second one starts a meeting. Keep the list to roughly eight to twelve codes, add a code for unanticipated failures, and expect to add two or three real codes after the first session, which is normal and worth doing before the other four.

Then hand the model the list, the transcripts, and one instruction that matters: apply only these codes, and put anything that does not fit under unanticipated rather than stretching a code to cover it. Stretching is how a model produces a tidy analysis of a messy test.

The transcript has to contain behaviour, or there is nothing to analyse

This is the failure that quietly ruins more AI-assisted analyses than any prompt problem. A speech-to-text transcript records what people said. Usability testing is about what people did. Feed a model pure speech and it will report, accurately, that three of five participants described the new layout as clean, while missing entirely that two of them scrolled past the primary action twice.

Fix it while moderating, not afterwards. Type observed actions into the transcript as they happen, in a consistent format: the participant scrolls past the primary action, the participant clicks the logo expecting it to be a back button, the participant pauses for eleven seconds on the empty state. Those lines are what the coding pass has to bite on, and they take no more effort than the notes you were taking anyway.

The moderation discipline underneath this, asking neutral questions and capturing observations rather than reactions, is covered directly in Usability Testing for Product Managers, which runs 1 week with 1 live session and 6 microlessons across planning, recruiting, task design, moderation and synthesis.

Every quote carries a timestamp, or it does not exist

Require a participant identifier and a timestamp on every quoted line in the output, and delete any quote that arrives without one. This is not a formality. A quote is the thing that gets pasted into a deck, repeated in a leadership meeting and remembered six months later, and a fabricated one does damage in proportion to how good it sounds.

Then spot check five of them against the recordings. Five is enough to tell you whether this model, on these transcripts, with this prompt, is reliable, and it takes about ten minutes. Telling the model in advance that timestamps will be verified measurably reduces the number of unverifiable lines it produces, which tells you something about what it does when nobody is checking.

  • A participant ID and timestamp on every quote, with unlocatable quotes deleted rather than chased.
  • Five verified quotes per analysis, checked against the recording before anything is shared.
  • A count of participants per code, not a count of mentions, because one talkative participant is not a pattern.

Frequency is countable, impact and persistence are not

Nielsen's severity scale runs 0 to 4, from not a usability problem at all through cosmetic, minor and major to usability catastrophe, and it combines three factors: how often the problem occurs, how hard it is for users to overcome, and whether it persists once the user has learned about it.

A model can give you the first one honestly, because frequency is arithmetic over your transcripts. It cannot give you the second or third, because neither is in the transcripts. How hard a problem is to overcome depends on what else the user could have done, what your support experience is like, and whether they had an alternative. Whether it persists depends on whether the interface teaches, which you only learn from repeat use rather than a first session.

So take the frequency column from the analysis and leave the severity column empty until a human fills it. A problem that two of five participants hit and recovered from in four seconds is not automatically less serious than one that four of five hit and shrugged at, and working out which is which is the review where product judgement earns its keep.

What a stakeholder should receive, and what they should not

Not the coded transcript. A stakeholder who receives forty tagged passages will read the first six and form a view from whichever two were most vivid, which is worse than receiving nothing because it arrives with the authority of research attached.

Send the problem list instead: each problem named as a behaviour, the count of participants who hit it, one verified quote with its timestamp, and your severity call with the reasoning in a clause. Six to ten problems is a readable list. If the analysis produced thirty, you coded topics rather than behaviours somewhere, or you merged two rounds that should have stayed separate.

Keep the raw output. When somebody disputes a finding three weeks later, being able to open the transcript at the timestamp and play it ends the argument in forty seconds. That is a better use of a recording archive than the assumption that anyone will rewatch a session.

What this does not replace

The test itself. Five participants is still five participants: Nielsen Norman Group's guidance holds that five users surface roughly 85 percent of usability problems and a single user about 31 percent, and running three rounds of five with design changes between them beats one round of fifteen. Faster analysis makes three rounds affordable, which is the actual win here, rather than making one round more conclusive than it was.

It also does not replace deciding what to test next. The analysis tells you what broke; it has no view on whether the flow should exist. That question belongs upstream with prototyping and discovery, and Prototyping with AI covers the build side over 1 week with 2 live sessions and 9 microlessons, including prototype copy for onboarding and empty states, which is where several of your codes will keep firing. The Discovery Expert Track sequences the wider practice, from problem discovery and interviewing through synthesis, opportunity mapping and validation.

See the Usability Testing for Product Managers bootcamp

For the review that happens before a participant ever sees the screen, getting design feedback from AI against the usability heuristics catches the mechanical failures that would otherwise consume a session. How to run customer interviews covers the adjacent discipline where the same transcript-quality problem applies, and research synthesis defines the wider practice this coding pass sits inside.

One thing worth doing after the second round: run the same code list against the previous round's transcripts and compare the participant counts. Codes that dropped are evidence your fix worked, codes that held are evidence it did not, and that comparison is the cheapest before-and-after measurement available to a product team that is not yet running experiments.

Bootcamps referred in this Guide

Frequently asked questions

What part of usability analysis can AI actually do?

Coding. Given transcripts and a fixed list of codes you wrote, a model will tag every passage consistently across five sessions in a few minutes, and consistency across sessions is exactly where human coders drift. It will also tally how many participants hit each code, which is the frequency number you need and the one most teams estimate from memory instead of counting.

Why does the code list have to come from me?

Because a model asked to find the themes will cluster by topic and you need clustering by failure. Topic clusters produce findings like navigation and onboarding, which are areas rather than problems. A code list you wrote from the test plan produces findings like did not notice the filter had been applied, which names a behaviour somebody can fix.

How do I stop it inventing quotes?

Require a timestamp and a participant identifier on every quoted line, then spot check five of them against the recording. A fabricated quote cannot survive a timestamp check, and a model told in advance that timestamps will be verified produces noticeably fewer of them. Any quote without a location gets deleted rather than investigated.

Can it tell the difference between what a participant did and what they said?

Only if your transcript records both. A speech-to-text transcript captures opinions, not behaviour, so a model reading it will faithfully report that a participant liked the new layout while missing that the same participant took ninety seconds to find the button. Note observed actions in the transcript as you moderate, or the analysis has nothing behavioural to work with.

How many sessions should I analyse this way?

The same number you would run anyway. Nielsen Norman Group's long-standing guidance is that five users surface roughly 85 percent of usability problems in a design, with a single user revealing about 31 percent, and that running three studies of five beats one study of fifteen. AI changes the cost of analysis, not the mathematics of saturation.

Should AI assign severity?

No. Nielsen's 0 to 4 scale combines frequency, impact and persistence. A model can count frequency from your transcripts, which is the mechanical third. Impact means how hard the problem is for users to overcome and persistence means whether they keep hitting it once they know about it, and both are judgements about your product and your users that the transcript does not contain.

Which Builders Camp bootcamp covers this?

Usability Testing for Product Managers runs 1 week with 1 live session and 6 microlessons and covers the whole chain: planning, recruiting, task design, moderation, and synthesis into prioritised issues. It sits inside the Discovery Expert Track, which covers problem discovery, interviewing, synthesis, opportunity mapping and validation as a sequence.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Mihaela Draghici

Mihaela Draghici

Through the Language Mapping Workshops & The Language Mapping Blueprint, Mihaela helps product leaders and teams get clear on how they talk about problems, priorities, ownership, outcomes, and success. She believes that when teams align on language, collaboration speeds up, trust increases, and execution becomes calmer and more effective.

Through the Language Mapping Workshops & The Language Mapping Blueprint, Mihaela helps product leaders and teams get clear on how they talk about problems, priorities, ownership, outcomes, and success. She believes that when teams align on language, collaboration speeds up, trust increases, and execution becomes calmer and more effective.

LinkedInMore guides by Mihaela Draghici

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Usability Testing for Product Managers bootcamp