Builders Camp

Tools

What AI for product discovery actually speeds up, and what it quietly breaks

AI safely speeds three parts of a discovery loop and should not touch the other two. Recruiting messages, interview guides and first-pass transcript coding are the accelerations worth taking; deciding what a finding means and choosing what to build stay human, because that is where an invented theme does the most damage. The check that separates the two is whether every theme carries a verbatim quote you can find in the source file.

Which parts of discovery does AI actually speed up?

Discovery is five jobs, not one, and AI is genuinely good at three of them. Recruiting and screening messages, interview guide drafting, and first-pass coding of transcripts all involve producing a lot of structured text from a clear brief, which is the shape of work a language model handles well. The two it should not own are the judgement calls: deciding what a pattern means for your product, and choosing which opportunity to pursue next.

The middle of that list is where the real time goes. Nielsen Norman Group's thematic analysis method runs six phases, and its practical advice is to budget at least as much time for analysis as you spent on collection. Eight 45-minute interviews therefore carry six hours of analysis behind six hours of talking. AI collapses the descriptive coding and the first clustering pass, which is most of that load, into minutes. It does not collapse the deliberate break the method puts between clustering and evaluating fit, because that step exists to give you distance from your own first read.

So the honest version of the promise is narrow. You are not buying discovery without talking to customers. You are buying back the afternoon you would have spent highlighting transcripts, on the condition that you spend some of it checking the highlights.

What does an invented finding look like in practice?

It looks like the most quotable line in the deck. A model asked to cluster eight transcripts into themes will produce a clean set of five, each with a confident label and a plausible summary, whether or not five themes exist in the data. The failure is rarely a factual error you can spot. It is a theme that is true of the category, true of products like yours, and absent from the transcripts in front of it.

The tell is the quote. Ask for a verbatim excerpt and a speaker identifier beside every theme, then search the source file for that exact string. Three things happen at that point: most themes check out, one or two have quotes that are near-verbatim rather than exact, and occasionally a theme has no retrievable quote at all. The near-verbatim ones matter more than they look, because a smoothed quote has usually had its hedge removed, and the hedge was the finding.

Anthropic's own guidance for reducing hallucinations lands on the same mechanism from the other direction: ask for word-for-word quotes before the analysis, have the model cite a supporting quote for each claim after drafting, and make it retract any claim where it cannot find one. Same discipline, expressed as a prompt instead of a review step. Doing both is not redundant, because a model asked to audit its own output will sometimes manufacture a quote to satisfy the audit.

How do you keep every theme traceable to a real quote?

Give each participant a stable code before anything reaches the model. P1 through P8, one per transcript, written into the file name and repeated at the top of every speaker turn. Then require every output row to carry three columns: the theme, the verbatim quote, and the code of the person who said it. A theme without all three does not get into the document you show anyone.

That structure also solves the privacy question, which most discovery workflows handle badly. Names, emails and account identifiers come out before the transcript leaves your systems, and the code replaces them. Your quotes stay traceable to a real person through a mapping file you keep, and the third-party tool never sees who that person is.

The last piece is a count. Next to each theme, record how many distinct participants it appears in. A theme carried by one person out of eight is a quote, not a theme, and labelling it accurately is the difference between a finding and a preference dressed as one.

Does AI change how many interviews you need?

No, and the thresholds people already use are the ones to keep. Torres recommends three to four story-based customer interviews before a team builds its first opportunity solution tree. Nielsen's usability work puts a single participant at around 31 percent of the problems in a design and five participants at roughly 85 percent, which is why five is the number most teams have internalised for evaluative testing. Neither number is about analysis cost. Both are about when new conversations stop telling you new things, and a faster coding pass does not move that point.

What does change is your appetite. When synthesis was six hours, eight interviews felt like a project. When it is forty minutes, twelve feels reasonable, and the constraint moves to recruiting. That is a real gain, as long as you do not quietly start counting a longer transcript pile as stronger evidence than it is.

What should stay entirely human, no matter how good the tooling gets?

Three things, and they are the three that decide what you build:

  • Deciding whether a theme is a problem worth solving or a complaint worth absorbing, which depends on strategy the transcripts do not contain.
  • Choosing which opportunity to pursue next, which is a bet about where your team's effort compounds.
  • Telling a stakeholder that the evidence killed the thing they asked for, which requires someone accountable for the call.

Everything above those lines is preparation. A discovery workflow that hands any of the three to a model has not saved time, it has moved the judgement somewhere nobody reviews.

Where does this fit in a weekly discovery cadence?

Run it as a loop with fixed slots rather than a project. Two interviews a week, coded the same day while the conversation is fresh, themes updated against the running codebook on Friday, and one decision recorded per fortnight with the evidence that supported it. Builders Camp teaches this shape in AI Prompting for Customer Discovery, whose practical challenge asks members to run one full loop: an assumption turned into a hypothesis with a stated falsification bar, an AI-drafted interview guide edited to cut leading questions, three themes each tied to a supporting quote and a confidence rating, and one product decision at the end.

The wider version of the same job is a Voice of the Customer system, which extends the loop beyond interviews to support tickets, sales calls and reviews. Those sources are higher volume and lower context, so the traceability rule matters more there, not less. If you want the whole sequence rather than a single bootcamp, the Discovery Expert Track bundles interviewing, synthesis, opportunity mapping and validation into one path.

The two-minute check that decides whether any of this was worth it

Before the readout, open the source file and search for three quotes from your deck at random. Not the ones you remember writing down, three chosen without looking. If all three return an exact match, the synthesis held. If one comes back missing or paraphrased, every theme in that document goes back through the quote check before anyone sees it, because a model that smoothed one quote smoothed others.

That check costs two minutes and it is the only thing standing between a fast discovery loop and a confident one that is wrong. For the tool-specific version of the workflow, see synthesising user research with Claude Code and triaging customer feedback with Claude Code. For the interview craft underneath all of it, how to run customer interviews covers the part AI does not touch, and AI for customer interviews covers the part it does.

See the AI Prompting for Customer Discovery bootcamp

Bootcamps referred in this Guide

Frequently asked questions

Can AI replace customer interviews in discovery?

No. Teresa Torres' opportunity solution tree method is explicit that opportunities have to come from story-based customer interviews, because generating them from what the team already knows imports the team's own biases. AI can draft the guide, transcribe the session and do the first coding pass. The conversation with a real person is the part that produces new information, and nothing in the workflow substitutes for it.

What is the fastest honest win from AI in discovery?

First-pass coding of transcripts. Nielsen Norman Group's thematic analysis method runs six phases and advises budgeting at least as much time for analysis as you spent collecting the data. AI collapses phases three and four, the descriptive coding and the first clustering, from hours to minutes, while you keep phases five and six, the break and the fit evaluation, entirely human.

How do I tell an invented theme from a real one?

Ask for a verbatim quote and a speaker identifier next to every theme, then search the source file for that exact string. A theme with no retrievable quote is an invention, regardless of how plausible it reads. Anthropic's own guidance for reducing hallucinations is to have the model find a supporting quote for each claim after drafting, and retract any claim where it cannot.

Should I let AI generate opportunities I have not heard in an interview?

Use them as interview prompts, never as tree entries. Torres makes the same distinction for sales conversations and support tickets: useful as inspiration for what to explore next, not as direct inputs, because they are missing the context an interview gives you.

Does AI change how many interviews I need?

It changes the cost of analysing them, not the number you need. Torres recommends three to four story-based interviews before you build a first opportunity solution tree, and Nielsen's usability research puts five participants at roughly 85 percent of the problems in a design. Those thresholds are about saturation in what people tell you, which AI does not affect.

Is customer data safe to paste into a general AI tool?

Check the retention and training terms of the specific tool and plan you are on before any transcript leaves your systems, and strip names, emails and account identifiers first. Replace each participant with a stable code like P4 so your quotes stay traceable internally without carrying personal data into a third-party tool.

Which Builders Camp bootcamp covers this workflow?

AI Prompting for Customer Discovery is the closest match. It runs 1 week, covers hypothesis-driven discovery, interview planning prompts, thematic synthesis and insight-to-decision, and its practical challenge requires each theme to be recorded against a supporting quote with a stated confidence level.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Mihaela Draghici

Mihaela Draghici

Through the Language Mapping Workshops & The Language Mapping Blueprint, Mihaela helps product leaders and teams get clear on how they talk about problems, priorities, ownership, outcomes, and success. She believes that when teams align on language, collaboration speeds up, trust increases, and execution becomes calmer and more effective.

Through the Language Mapping Workshops & The Language Mapping Blueprint, Mihaela helps product leaders and teams get clear on how they talk about problems, priorities, ownership, outcomes, and success. She believes that when teams align on language, collaboration speeds up, trust increases, and execution becomes calmer and more effective.

LinkedInMore guides by Mihaela Draghici

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the AI Prompting for Customer Discovery bootcamp