---
title: "Analyse Product Metrics With Claude Code"
description: "Point Claude Code at a CSV export, ask for the number, then run the four sanity checks that catch a confidently wrong metric before it reaches your roadmap."
canonical_url: "https://builderscamp.com/guides/tools/claude-code-metric-analysis-for-pms"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Mário Araújo"
publisher: "Builders Camp"
guide_class: "tools"
---

# How to analyse product metrics with Claude Code

**TL;DR:** Claude Code analyses a metrics export by writing a script against the file and running it, so you get the number and the code that produced it in one answer. Budget more time for the checks than for the question: row reconciliation, a named denominator, the partial final period and duplicate IDs catch most confidently wrong results. What it cannot catch is broken instrumentation, because the file looks fine either way.

## What can Claude Code actually do with a metrics export?

It reads the file, writes a script against it, runs that script, and shows you the code and the output together. That is the whole mechanism, and it is the reason this works better than pasting rows into a chat window. A model summarising 40,000 rows by reading them is estimating. A model that writes six lines of Python, executes them in your terminal and prints the result is computing, and you can rerun the same script tomorrow and get the same number.

The setup is smaller than people expect. Make a folder, put the export in it, start Claude Code from that folder, and describe the question in a sentence. Claude Code starts with read-only permissions in its manual mode and asks before it runs anything that changes your machine, so the first pass over a CSV is about as low risk as an agentic tool gets.

## How do you ask for a number you can defend?

Name the denominator before you name the metric. "What is our activation rate?" has at least three defensible answers depending on whether the denominator is signups, verified accounts or accounts that reached day 7, and an agent asked the bare question will silently pick one. "Of accounts created in August, what share fired the first project-created event within 14 days of signup?" has exactly one answer, and you can argue with the definition instead of arguing with the number.

Then ask for three things in the same request: the answer, the script that produced it, and the row counts at each filtering step. The row counts are the part people skip and the part that catches errors. If 12,400 rows go in, 9,100 survive the date filter and the segments underneath sum to 8,600, then 500 rows vanished somewhere and nobody knows where yet.

Keep the definition somewhere the next session can read it. A short project file with your metric definitions, your date conventions and the names of the events that matter turns a one-off answer into a repeatable one, which is the difference between a useful habit and a party trick.

## What does a first prompt look like when it works?

A weak one names the metric and stops: "analyse this export and tell me how activation is doing." A working one names the file, the definition, the window, the cut and the evidence you want back. Something closer to: read signups.csv and events.csv, join them on account_id, and for every account created between 1 August and 31 August, tell me what share fired a project_created event within 14 days of the signup timestamp, split by acquisition channel. Show me the script, the row count after each filter, and the five accounts closest to the 14 day boundary.

That last clause is the one worth copying. Asking for the rows nearest a threshold surfaces off-by-one errors in the window logic, which are the most common silent bug in this kind of analysis and the hardest to spot in an aggregate. Everything else in the prompt is bookkeeping; the boundary rows are the test.

## Which sanity checks catch a confidently wrong number?

Four, and they take about two minutes each.

- **Row reconciliation.** Total rows in, total rows per segment out, and the segments must sum back. A gap means a filter dropped rows quietly, almost always nulls or unparsed dates.
- **The partial final period.** A weekly chart whose last bucket holds three days of data always looks like a collapse. Ask what date range each bucket actually covers before you read the trend.
- **Duplicate identity.** Count rows, then count distinct user or account IDs. When those two numbers differ and your analysis assumed one row per user, every per-user average is wrong by the duplicate factor.

The fourth is the cheapest and the one people trust least: pull five raw rows from one segment and read them yourself. If a segment claims 900 accounts activated in week two, look at five of those nine hundred and check that each one actually did the thing. A single mislabelled row in a sample of five is a strong signal that the join is wrong.

## What it cannot see, and why that matters more than the maths

An agent answers the question you asked about the file you gave it. It cannot tell you the event stopped firing on Android three weeks ago, that a rename split one funnel step into two, or that the export was generated in a different timezone than the dashboard you are comparing against. The arithmetic will be correct and the conclusion will be wrong, which is worse than an obvious error because nothing looks broken.

That gap is a measurement problem, not a tooling problem. Builders Camp's Product Analytics bootcamp, run by Mário Araújo over 1 week with 2 live sessions and 9 microlessons, spends most of its time on exactly this layer: what to instrument, how user-level events roll up into account-level truth in a business-to-business product, and why teams with good tools still read retention curves wrong. Data for Product Managers goes wider across 2 weeks and 4 live sessions, covering funnels, cohorts and the correlation-versus-causation trap that survives any amount of clean computation.

If you want a definition of a term before you argue about it, [cohort analysis](https://builderscamp.com/guides/glossary/cohort-analysis) and the [metrics tree](https://builderscamp.com/guides/glossary/metrics-tree) entries are the two that come up most often in these conversations.

## When is a spreadsheet still the faster tool?

When the question is one filter deep and you already have the file open. Sorting a 300-row export by revenue and looking at the top ten is not a job for an agent, and reaching for one adds a minute of setup to a ten second task.

The crossover point is roughly where the answer needs a join, a repeat or an audit trail. Two exports that need matching on account ID, a question you will ask again next month, or a number that will end up in a document someone else challenges: those are the cases where having the script beats having the answer. A [ChatGPT session](https://builderscamp.com/guides/tools/chatgpt-for-product-managers) will happily discuss the same data, but it cannot execute a script against a local file and hand you the code, and that is the specific thing being bought here.

## How does this fit the rest of the analysis workflow?

Treat the agent as the step between "I have a question" and "I have a defensible chart", not as the whole pipeline. Discovery happens before it, in the decision about which number would change what you do. Communication happens after, because a correct cohort table convinces nobody on its own. For the reading and summarising half of that workflow, [NotebookLM](https://builderscamp.com/guides/tools/notebooklm-for-product-managers) is a better fit than a terminal agent, and the two do not compete.

Claude Code for Product Managers is the 1 week bootcamp that teaches the terminal half directly: 2 live sessions, 8 microlessons and a practical challenge built around reviewing agent output critically rather than accepting it. Its certification quiz has 10 questions and a pass mark of 70 percent, which is a modest bar and not the point; the practical challenge is, and it is deliberately a case where the generated work looked clean and was wrong.

## Start with one number you already argue about

Pick the metric your team disagrees on most, export it, and write the definition down before you open anything. Half the time the argument dissolves at that step, because two people were computing different denominators and neither had said so out loud. The other half you now have a script, and the argument moves to where it belongs, which is what the number should change.

[See the Claude Code for Product Managers bootcamp](https://builderscamp.com/bootcamps/claude-code-for-product-managers?utm_source=guide&utm_medium=organic&utm_campaign=claude-code-metric-analysis-for-pms)

For the measurement layer underneath, the [Product Analytics bootcamp](https://builderscamp.com/bootcamps/product-analytics) covers account-level analysis and activation definitions, and [Data for Product Managers](https://builderscamp.com/bootcamps/data-for-pm) covers funnels, cohorts and experiment design across 2 weeks.

## Frequently asked questions

### Does Claude Code need database access, or is a CSV export enough?

A CSV export is enough for most PM questions. Claude Code reads files in the folder you start it from, so an export dropped into a working folder is something it can open, group, join and count. Database access through an MCP server is a different setup with a different risk profile, and it is not required to answer questions like activation by cohort or feature adoption by account.

### How do I know it computed the number instead of estimating it?

Ask for the script and the raw output, not just the answer. A model reading tens of thousands of rows and reporting a total is estimating. A model that writes a few lines of Python, runs them and prints the result is computing, and you can rerun that same script yourself. If the reply contains a number but no code and no command output, treat it as a guess.

### Which sanity check catches the most wrong numbers?

Row reconciliation. Ask for the total row count in the file, then the row count in every segment of the breakdown, and check that the segments sum back to the total. When they do not, a filter dropped rows silently, usually null values or dates that failed to parse, and every percentage built on that breakdown is wrong.

### Can I put customer data in a CSV that Claude Code reads?

Check your plan and your own policy first. Anthropic states that under commercial terms on Team and Enterprise plans it does not train generative models on code or prompts sent to Claude Code, while Free, Pro and Max accounts have a setting that controls whether their data is used for model improvement. Separately, strip or hash direct identifiers you do not need for the analysis, because the analysis rarely needs a real email address to produce a cohort curve.

### Why did the same question give two different answers on two runs?

Usually because the question was underspecified, not because the tool is unreliable. Activation rate has at least three defensible definitions depending on the denominator and the time window, and an agent asked for it without those will pick one. Write the definition down in the prompt, or in a project file the session loads, and the answer stops moving.

### Does this replace a data analyst?

It removes the queue for small questions and it does not replace the person who owns the instrumentation. An agent answers the question you asked about the file you handed it. It cannot tell you that the event stopped firing on Android three weeks ago, which is exactly the class of problem an analyst catches and a CSV never shows.

### Which Builders Camp bootcamp covers this?

Claude Code for Product Managers is the direct match: 1 week, 2 live sessions and 8 self-paced microlessons, one of which is specifically about analysing qualitative and quantitative product data while keeping the reasoning auditable. Product Analytics and Data for Product Managers cover the measurement side that decides whether the number was worth asking for.

## Sources

- [Anthropic: Claude Code security](https://docs.claude.com/en/docs/claude-code/security)
- [Anthropic: Claude Code data usage](https://docs.claude.com/en/docs/claude-code/data-usage)
- [Builders Camp: Claude Code for Product Managers](https://builderscamp.com/bootcamps/claude-code-for-product-managers)
- [Builders Camp: Product Analytics](https://builderscamp.com/bootcamps/product-analytics)
- [Builders Camp: Data for Product Managers](https://builderscamp.com/bootcamps/data-for-pm)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
