Builders Camp

Tools

How to use AI for customer feedback analysis without losing the evidence

Customer feedback analysis with AI holds up when the codebook exists before the model runs, every theme carries retrievable quotes, and 5 to 15 percent of items land in an explicit unclassified bucket. Nielsen's severity model treats a single evaluator's ratings as too unreliable to trust, and a model is one evaluator. The quarterly job is not re-running the analysis, it is checking the themes still mean what they meant last quarter.

What should a feedback analysis actually produce?

A table, not a paragraph. One row per piece of feedback, with the theme it was assigned, the verbatim wording, the source it came from, and the date. The themes sit on top of that table as a count and a definition, never as free-standing prose. If the analysis cannot answer "show me the 14 tickets in this theme and what each one said", it is a summary, and a summary is unfalsifiable by design.

That format costs nothing extra when a model is doing the assignment, and it is the whole difference in what happens next. A stakeholder who disagrees with a theme can read the rows. A PM revisiting the analysis in six months can see whether a theme grew or the definition loosened. Neither is possible from a paragraph that says customers are frustrated with onboarding.

Why does the codebook come before the model?

Because a model asked to find themes will find themes, every time, in any corpus, including a corpus with no structure in it. The clustering is not a discovery, it is an output format. What makes it a discovery is that you defined what you were looking for and it came back at a volume you did not expect.

So run the first pass open, on a sample rather than everything, purely to learn the vocabulary of your own corpus. Then write the codebook by hand: eight to twelve themes, each with a one-line definition saying what it includes, one line saying what it excludes, and one real example. Add an explicit unclassified bucket. Every pass after that uses the fixed codebook, and the model's job stops being invention and becomes assignment, which is a task you can check.

Nielsen Norman Group's thematic analysis method runs six phases, and the two that matter most here are the ones people skip: the deliberate break before evaluating theme fit, and the fit evaluation itself with a second reader. Its practical advice is to budget at least as much time for analysis as you spent on collection. AI changes the cost of phases three and four. It does not remove the need for phase six, and the codebook is what makes phase six possible at all.

How do you check a theme is real rather than plausible?

Three checks, in this order, and all three are cheap:

  • Retrieval. Pick three quotes attached to the theme and search the source file for the exact string. A paraphrase that has lost a hedge is the common failure, and it matters more than an outright fabrication because it survives a casual read.
  • Counts that reconcile. Sum the per-theme counts and compare against the row count of the file you supplied. A model that thinned out the middle of a long file still returns confident themes, and only the arithmetic shows it.
  • Distinct sources. Record how many separate customers, not how many items, sit behind each theme. One vocal account filing nine tickets is one customer, and the theme table should say so.

Anthropic's own guidance for reducing hallucinations describes the first check as a prompt rather than a review step: extract word-for-word quotes before the analysis, and after drafting, make the model find a supporting quote for each claim and drop any claim it cannot support. Doing it as both a prompt instruction and a human spot check is not duplication, because a model auditing itself will sometimes produce a quote to pass the audit.

What does volume tell you, and what does it hide?

Volume tells you what people complain about in writing. It does not tell you what costs you money, and the two diverge in a predictable direction. Low-effort annoyances generate a lot of tickets and little churn. The thing that quietly loses accounts often generates almost no feedback at all, because people who leave stop writing to you first.

Nielsen's severity model is the useful correction: severity combines frequency, impact if it occurs, and persistence, rated on a 0 to 4 scale from cosmetic problem to catastrophe. Ask a model for that rating alongside the theme and it will produce one, confidently. Treat it as one evaluator's opinion. Nielsen is explicit that ratings from a single evaluator are too unreliable to be trusted and that the mean of three independent evaluators is the practical standard, and a model has no exposure to your support costs, your contract values or your renewal calendar.

The working compromise is to let the model produce frequency, which it can genuinely count, and have a human supply impact and persistence for the top ten themes only. That is an hour a quarter, not a project.

How do you stop the themes drifting between quarters?

Freeze the codebook and version it. Each quarter, run the new feedback against the existing definitions, and record three numbers: the size of the unclassified bucket, the count per theme, and the number of themes you had to add. A healthy run leaves 5 to 15 percent unclassified. Zero means the model is forcing items into whatever fits closest, which is exactly how a genuinely new problem stays invisible for two quarters.

Then do the part nobody schedules. Take last quarter's top three themes, pull ten rows each, and read them against this quarter's ten. If the wording has shifted while the label stayed the same, the trend line you have been showing leadership is measuring your codebook, not your customers.

Which sources can share a number, and which cannot?

Code everything with the same codebook so themes are comparable, and report counts separately by source. Support tickets are a self-selected sample of people annoyed enough to write in. Sales call notes are filtered through what a seller thought was worth recording. Interviews are a recruited sample. App store reviews skew to the two extremes of the distribution. A single blended count across those four is a number with no population behind it, and it will be the number that gets quoted in the roadmap review.

Builders Camp's Voice of the Customer covers exactly this in 1 week: feedback sources and capture, taxonomy and tagging designed for decisions rather than reporting, synthesis into insights, prioritisation, closing the loop with customers, and the operating cadence that keeps the system from decaying into a shared inbox. Product Analytics sits next to it for the behavioural half of the picture, where activation, cohort retention and account-level engagement tell you what the tickets never will. AI Prompting for Customer Discovery covers the interview end of the same pipeline.

What this looks like when someone challenges it

The quarterly readout is not where feedback analysis is won. It is won in the ten minutes afterwards when a sales lead says a theme is wrong, and you open the table, filter to that theme, and read the fourteen rows aloud. Either the rows support the label, and the conversation moves to what to do about it, or two of them clearly belong somewhere else, and you fix the definition in front of everyone.

Both outcomes are fine. What is not fine is having nothing to open, which is the state every summary-shaped analysis leaves you in. For the tool-specific versions of this workflow, see triaging customer feedback with Claude Code and synthesising user research with Claude Code. For the survey half of the same problem, AI for survey analysis covers open-text coding and the sampling traps behind it.

See the Voice of the Customer bootcamp

Bootcamps referred in this Guide

Frequently asked questions

What is the difference between summarising feedback and analysing it?

A summary compresses. An analysis assigns every item to a theme and keeps the assignment reversible, so you can open any theme and read the original wording behind it. The test is whether you can answer the question 'which 14 tickets are in this theme, and what did each one actually say', which a summary cannot answer at all.

Should I let the model invent the theme names?

For the first pass on a new corpus, yes, because you do not yet know what is in it. For every pass after that, no. Fix the codebook, define what each theme includes and excludes in writing, and add an 'unclassified' bucket. A model given free rein on the second pass will rename yesterday's themes and your trend line becomes meaningless.

How do I know the model read everything?

Count. Ask for the total number of items processed and the number assigned to each theme, and check the sum matches the row count of the file you supplied. A model that silently skipped the middle of a long file will still produce a confident set of themes, and the count is the only cheap signal that it did.

How big should the unclassified bucket be?

Between 5 and 15 percent on a healthy run. Zero means the model forced every item into an existing theme, which hides the new problem you most wanted to find. Above a quarter means your codebook no longer describes the feedback you are receiving, and it needs rewriting rather than patching.

Can AI rate severity as well as assign themes?

Treat its rating as one evaluator's opinion, not a verdict. Nielsen's severity model combines frequency, impact and persistence on a 0 to 4 scale, and his guidance is that ratings from a single evaluator are too unreliable to trust, with the mean of three evaluators being the practical standard. A model is one evaluator with no exposure to your support costs.

Should support tickets and interview transcripts go in the same theme counts?

No. Tickets are a self-selected sample of people annoyed enough to write in, interviews are a recruited sample, and blending them produces a count that means nothing. Code them with the same codebook so the themes are comparable, then report the two counts separately.

Where does Builders Camp teach this end to end?

Voice of the Customer is the direct match. It runs 1 week and covers feedback sources and capture, taxonomy and tagging designed for decisions rather than reporting, synthesis into insights, prioritisation, closing the loop with customers, and the operating cadence that keeps the system alive.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Mihaela Draghici

Mihaela Draghici

Through the Language Mapping Workshops & The Language Mapping Blueprint, Mihaela helps product leaders and teams get clear on how they talk about problems, priorities, ownership, outcomes, and success. She believes that when teams align on language, collaboration speeds up, trust increases, and execution becomes calmer and more effective.

Through the Language Mapping Workshops & The Language Mapping Blueprint, Mihaela helps product leaders and teams get clear on how they talk about problems, priorities, ownership, outcomes, and success. She believes that when teams align on language, collaboration speeds up, trust increases, and execution becomes calmer and more effective.

LinkedInMore guides by Mihaela Draghici

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Voice of the Customer bootcamp