Builders Camp

Tools

How to triage customer feedback with Claude Code and keep every quote traceable

Triage works when every theme cites the row identifiers behind it, which means the export has to stay a file with identifiers rather than text pasted into a chat. Strip the personal data, require three supporting rows before a theme exists, and read the unclustered remainder yourself, because that is where anything new is hiding.

What does a defensible triage output actually look like?

A short list of themes, each with a count, each with the row identifiers behind it, plus a section of everything that did not fit. That last section is the tell. An output with five clean themes and no remainder has either been given unusually tidy data or has quietly forced outliers into the nearest bucket, and the outliers are the part you were hoping to find.

The reason to run this in a file-based environment rather than a chat window comes down to that citation requirement. Claude Code reads the export off disk and writes its analysis to a second file beside it, so the themes and the evidence live in the same folder and a colleague can check any claim without asking you for the original. A pasted export produces the same-looking summary with no path back to the source, which is fine until someone senior asks how many customers actually said that.

How should you prepare the export?

Strip the personal data first. Names, emails, account identifiers and free-text fields that contain them should come out before the file reaches the agent, replaced by the row identifier you intend to cite. This costs one spreadsheet pass and buys two things: an output you can circulate without a privacy review, and a discipline that keeps the analysis about problems rather than about which customer is loudest.

Keep everything that carries signal about weight: plan tier, date, channel, whether the ticket was resolved. These are what let you ask later whether a theme is concentrated in one segment, which is usually the question that decides whether it matters. Check your own data policy before the first run, and read Anthropic's published documentation on how Claude Code handles data rather than assuming it matches the tool your company already approved.

What does the triage prompt need to contain?

Four instructions, and the last two are the ones that separate a usable output from a slide.

Give the agent the file and the decision the analysis serves. Require a minimum of three supporting rows before anything is allowed to be a theme, which kills the single loud complaint that would otherwise arrive dressed as a trend. Require each theme to be phrased as the customer's problem rather than as a category: "cannot tell which invoice a charge belongs to" rather than "billing". And require the leftovers, everything that did not reach three rows, listed in full at the end.

The phrasing rule does more work than it looks like. Categories such as usability, pricing and performance describe the shape of your support taxonomy, not what anybody wants, and a theme list made of them cannot be argued with because it does not say anything. A theme phrased as a problem can be wrong, which is what makes it useful.

How do you keep the quotes traceable?

Make the identifier part of the claim, not part of a bibliography. Every theme line should carry the row numbers behind it inline, and every quoted sentence should carry the single row it came from. Then spot check: pick three rows at random from a theme, open the export, and confirm each one actually supports the theme it was filed under.

That check catches the specific failure that matters here, which is a row assigned to a theme it only loosely fits because the theme needed a third example. Anthropic's best-practice guidance makes the general version of this point, that an agent stops when the work looks done, so unless a verification step exists the only signal available is whether the output looks finished. A theme list always looks finished.

What do you do when the export is too big for one pass?

Batch it, and make each batch write its own file. A run of 400 rows behaves differently from a run of 4,000, and the difference is not file size but context: Anthropic's guidance treats the context window as the most important resource to manage and warns that performance degrades as it fills, which in practice means a long run produces careful themes for the first few hundred rows and increasingly lazy ones after that.

The pattern that works is two passes. First, split the export and have each batch produce its own themes file with the row identifiers intact. Then run a second pass over those files only, merging themes that are the same problem under different words and keeping the row identifiers as they merge. The second pass reads a few hundred lines rather than a few thousand rows, so the merge happens with the whole picture in view.

Write the instructions down once rather than rebuilding the prompt each quarter. A CLAUDE.md file in the folder can hold the rules that never change, the three-row minimum, the problem-phrasing requirement, the leftovers section, so the next run starts from your standard rather than from whatever you remember. Anthropic's own framing for that file is exactly this: it is where you put what you would otherwise re-explain every session.

Where does automated triage get it wrong?

  • It flattens intensity. Ten mild mentions and two enraged ones become a count of twelve, and the churn risk sitting in those two disappears into an average.
  • It over-reads recency. A batch of tickets from the week an incident happened produces a theme that is really an event, and nothing in the text marks the difference.
  • It cannot see who is missing. Customers who churned silently, or who never contact support at all, generate no rows, and the analysis will confidently describe a population that excludes them.

None of these is fixed by a better prompt. They are fixed by knowing they are there and by treating the output as a map of what was said to you, not of what is true about your users.

How does this fit a real feedback system?

Triage is one step in a loop, and running it in isolation produces a document that is read once. Builders Camp's Voice of the Customer bootcamp lays out the fuller pipeline, from sources and capture through taxonomy and tagging to synthesis, prioritisation and closing the loop back to customers, with an operating cadence so the work stays alive rather than becoming a one-time project. It runs 1 week, with 2 live sessions and 10 self-paced microlessons.

The discovery half sits next to it. Builders Camp's AI Prompting for Customer Discovery bootcamp covers turning transcripts and notes into themes, tensions and opportunity statements with traceability to evidence, which is the same discipline applied to interviews rather than to an export, and it is explicit that the point is speeding up thinking rather than replacing user conversations.

What triage is really for

It is a pointer, not a verdict. The output should end with you knowing which four customers to call and which assumption to go and test, and if it ends instead with a ranked list you are about to put into a roadmap, the tool has been asked to make a judgement it has no standing to make. Builders Camp's Claude Code for Product Managers bootcamp teaches the reviewing habit that keeps that line visible, in 1 week across 2 live sessions and 8 self-paced microlessons, including a microlesson on validation techniques for assessing the reliability and risk of an agent's output before acting on it.

See the Claude Code for Product Managers bootcamp

For working against a fixed set of sources with citation built in, see NotebookLM for product managers, and for the chat-based version of the same synthesis job, ChatGPT for product managers. The failure mode the traceability rule exists to catch is covered in AI hallucination.

Bootcamps referred in this Guide

Frequently asked questions

Why use Claude Code rather than pasting the export into a chat window?

Because the export stays a file. The agent reads it off disk row by row, writes the themes into a second file next to it, and can cite row identifiers because the rows have identifiers. A pasted export loses its structure at the moment of pasting, which is exactly when traceability disappears.

What should I strip out of the export before I start?

Names, email addresses, account identifiers and anything else that identifies a person, unless you have a specific reason and a legal basis to keep them. Replace them with the row identifier you will cite instead. This makes the output safer to circulate and does not weaken the analysis, because a theme never depended on who said it.

How do I stop the themes being generic?

Ban the abstract ones in the prompt. Themes like usability, performance and pricing describe the column the feedback arrived in, not what customers want. Ask for themes phrased as the problem the customer is having, and require at least three supporting rows before a theme is allowed to exist.

What do I do with feedback that fits no theme?

Keep it, in its own section, and read it yourself. The unclustered remainder is where new problems live, and an analysis that silently discards it will report a stable picture of the world for months after the world changed.

How many rows can this handle at once?

Fewer than you would like, and the limit is context rather than file size. Anthropic's best-practice guidance treats the context window as the resource to manage and notes that performance degrades as it fills, so process in batches and have each batch write its output to a file rather than holding everything in one conversation.

Is it safe to put customer feedback through this?

Treat it as you would any third-party processing of customer data: check your own policy first, strip identifiers, and read Anthropic's published data usage documentation rather than assuming. The safest default for a first run is an export with personal fields already removed.

Does this replace talking to customers?

No, and Builders Camp's discovery bootcamp is explicit about the same line: use AI to speed up thinking, not to replace talking to users. Triage tells you which conversations to go and have, which is a different and smaller claim than knowing what customers want.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Inês Lourenço

Inês Lourenço

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

LinkedInMore guides by Inês Lourenço

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Claude Code for Product Managers bootcamp