Builders Camp

Tools

How to synthesise user research with Claude Code

Synthesise in two passes rather than one: code each transcript on its own, then cluster those codes across transcripts. The two pass split is what keeps every theme attached to a participant ID and a verbatim quote, which is the only property that makes synthesis defensible when someone disagrees with the conclusion. It does nothing about a biased sample, so recruitment still decides what you can find.

What does good synthesis actually produce?

A theme list where every line can be traced back to a person who said something. Not a set of well-phrased insights, not a tidy affinity map, and not a summary of what users want. The deliverable is the chain: theme, participant IDs, verbatim quotes, source file. Anything that cannot survive somebody asking "who said that, exactly?" is a hypothesis wearing the clothes of a finding.

This is the specific thing an agent can either help with enormously or destroy quietly, depending on how you run it. Handed twelve transcripts and asked for themes, a model will produce five plausible ones in thirty seconds, and you will have no way to tell which of them came from the data.

Why does the two pass split matter so much?

Because compression and traceability pull in opposite directions. A single session asked to read twelve transcripts and report themes has to compress, and what gets dropped first is the link between a claim and the sentence that supports it. The output reads beautifully and cites nothing.

Splitting the work fixes it. Pass one takes each transcript on its own and produces a per-participant code file: short labels for what this person said, each with the verbatim line and its position in the transcript. Pass two reads only those code files, not the raw transcripts, and clusters codes into themes. Because every code already carries its quote, the clustering pass cannot invent evidence, only group it. The themes that come out the far end arrive with citations attached because they never had a chance to lose them.

The practical setup is a folder per study: one transcript file per participant with a consistent name, one codes file per participant, one themes file at the end. Claude Code reads from the folder you start it in, so the whole study is one directory and the second pass is one instruction.

What rules should the coding pass follow?

Three, and they are worth writing into a file the session loads rather than retyping each time.

  • Every code carries a verbatim quote and a participant ID. No quote, no code. This is the rule that makes everything downstream checkable.
  • Mark each code spontaneous or prompted. Spontaneous means the participant raised it; prompted means you asked about it directly. Frequency only means something inside the spontaneous set.
  • Preserve hedges and contradictions. When somebody says they would probably pay for it but then describes a workaround they already use, both halves are the finding. A coding pass optimised for clean labels will keep the first and drop the second.

The third rule is the one people skip and the one that most often changes a decision, because the contradiction between what someone says and what they do is usually the most informative thing in the transcript.

How do you check the output before you present it?

Grep three quotes. Pick three at random from the final theme list, search the transcript files for the exact string, and see whether they exist. This takes two minutes and catches both of the failure modes that matter: the invented quote, which does not appear at all, and the tidied quote, where filler words were removed and the meaning quietly shifted with them.

Then run one adversarial pass. Ask for the three strongest pieces of evidence against the conclusion you are about to present, with participant IDs and quotes. An agent asked to support a hypothesis will support it, because that is what the request describes. Asked to attack it, it will find the two participants whose transcripts you skimmed, and one of them usually matters.

For the vocabulary underneath all of this, the research synthesis entry covers the coding and clustering terms, and confirmation bias in data covers the specific way this goes wrong when the person running the analysis already has a preferred answer.

How many interviews before this is worth setting up?

Below five, read them yourself. Five transcripts is a couple of hours of careful reading, and the reading is where you notice the thing nobody said out loud: the pause before an answer, the two participants who described the same workaround in different words, the question that made someone change the subject. A coding pass flattens all of that into labels, and at small volumes the flattening costs more than it saves.

The setup starts paying somewhere around eight to twelve, and it compounds from there. At twelve transcripts the per-participant code files are faster to re-read than the transcripts, and at thirty they are the only way anyone revisits the study at all. The other threshold is repetition: if you run the same kind of study every quarter, build the folder structure and the coding rules once even on a small first study, because the second one costs almost nothing after that.

None of this argues for skipping the raw transcripts entirely. Read two in full, whatever the volume, and let the rest go through the pipeline. Those two are your calibration for whether the codes coming out the other end resemble the conversations that went in.

What this cannot fix, no matter how well you run it

Sampling. If the twelve people you spoke to were recruited from your most engaged accounts, a flawless, fully cited synthesis gives you a flawless, fully cited description of your most engaged accounts, and the fact that every quote checks out makes that conclusion feel stronger rather than weaker. Traceability is a defence against fabrication, not against a skewed sample.

It also cannot decide what to do. An opportunity list phrased as customer problems is the output; picking which one to pursue is a judgement about strategy, effort and where the business needs to win. Builders Camp's AI Prompting for Customer Discovery bootcamp covers that whole arc across 1 week, 2 live sessions and 10 microlessons: hypothesis-driven discovery, interview planning, recruiting and screeners, thematic analysis, and turning insights into opportunities you can rank. The Discovery Expert Track bundles 7 bootcamps for people who want the interviewing and validation craft alongside the synthesis, rather than the synthesis on its own.

Claude Code for Product Managers is the narrower one: how to structure files and folders so an agent can reason across research, specs and decisions without losing context, and how to validate what it hands back before acting on it. Its practical challenge is built around exactly that reflex, using a case where the generated work looked correct and was not.

When is a different tool the better call?

When you want to interrogate a body of research rather than process a new batch. NotebookLM is built for asking questions of documents you have already collected, and it answers with citations into the source set by design. Claude Code is the better fit when the work is producing structured artefacts in a repeatable shape, in files you will search again, next to the specs those findings are going to change.

If the study has not started yet, the synthesis question is premature anyway. A discovery interview script that avoids leading questions removes more downstream analysis problems than any coding workflow can repair afterwards.

Run the adversarial pass first next time

The habit worth stealing from all of this is not the two pass structure, useful as it is. It is asking for the evidence against your own conclusion before you ask for the themes that support it. Do that on the study you finished last month, on transcripts you have already read, and see how long the list is. That is the honest measure of how much your current synthesis process is telling you something new versus telling you what you walked in with.

See the Claude Code for Product Managers bootcamp

For the discovery craft that sits upstream of synthesis, the AI Prompting for Customer Discovery bootcamp runs 1 week, and the Discovery Expert Track covers interviewing, synthesis and validation end to end.

Bootcamps referred in this Guide

Frequently asked questions

Why code each transcript separately instead of asking for themes across all of them at once?

Because a single pass over twelve transcripts compresses, and compression is where traceability dies. Coding one transcript at a time keeps each code attached to the participant and the line it came from. The clustering pass then works on codes that already carry their evidence, so the final theme list can cite a real quote rather than a plausible one.

How do I check that a quote in the output is real?

Search for it in the transcript files, verbatim. Pick three quotes at random from the theme list and grep for the exact string. Two failure patterns show up this way: a cleaned-up quote where filler words were removed and the meaning shifted, and a quote that does not exist at all. If three out of three match exactly, the run is probably sound. If one does not, redo the pass with a stricter citation rule.

Does a theme mentioned by more participants matter more?

Not on its own. Seven of twelve participants mentioning onboarding friction means very little if five of them were asked directly about onboarding. Ask the coding pass to mark each code as spontaneous or prompted, then read frequency only within the spontaneous set. That single distinction changes which themes look strong more often than any other check.

Can Claude Code fix a biased sample?

No, and this is the limit worth stating plainly. Perfect traceability over a sample of twelve power users produces a perfectly evidenced description of power users. Recruitment and screener design decide what the synthesis can possibly find, which is why they sit upstream of it in any serious discovery process.

How do I stop it agreeing with the hypothesis I already have?

Ask for the evidence against it, by participant, before you ask for themes. A prompt that names your hypothesis and requests supporting themes will get supporting themes. A prompt that asks for the three strongest pieces of contradicting evidence, with quotes, gets something you can actually use, and often kills a feature two weeks earlier than it would have died otherwise.

Should customer transcripts go through an AI tool at all?

That is a consent and policy question before it is a tooling one. Check what your participants agreed to and what your plan terms say: Anthropic states that under commercial terms on Team and Enterprise plans it does not train generative models on code or prompts sent to Claude Code, while Free, Pro and Max accounts have a setting that controls whether their data is used for model improvement. Replacing names with participant IDs before synthesis costs nothing and removes most of the exposure.

What does the finished artefact look like?

A theme list where each theme has a one line statement, the participant IDs that support it, two or three verbatim quotes with their source file, and a note on whether the evidence was spontaneous or prompted. Underneath it, an opportunity list phrased as customer problems rather than features, each pointing back at the themes it rests on.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Inês Lourenço

Inês Lourenço

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

LinkedInMore guides by Inês Lourenço

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Claude Code for Product Managers bootcamp