---
title: "How to Use AI for Market Research Work"
description: "Size a market with AI without inventing the number: build bottom-up from counts you can name, open every citation, and test which assumption moves the answer."
canonical_url: "https://builderscamp.com/guides/tools/ai-for-market-research"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Mário Araújo"
publisher: "Builders Camp"
guide_class: "tools"
---

# How to use AI for market research you can actually defend

**TL;DR:** Decide the threshold before you size anything: nobody needs the market number, they need to know whether the opportunity clears a bar that a specific decision depends on. Build bottom-up as a chain of five or six terms, each carrying a named source and a range, so a reviewer can disagree with one step rather than with the whole figure. Then open every citation the model produced, because a fabricated source is indistinguishable from a real one until you click it.

## What is a market size actually for?

Almost never for the number. It is for a threshold: whether this segment can carry a team, whether a second product line clears the bar that justifies the hiring, whether the opportunity is large enough to be worth losing focus over. Name that threshold first, in a sentence, before any research starts.

Doing so changes the work. If the bar is 10 million in annual revenue and the honest range of the estimate is 4 million to 40 million, you have not answered the question and should say so, rather than reporting the midpoint. Most sizing exercises produce a single number precisely because nobody wrote down what it had to clear.

## Why top-down sizing cannot be argued with, and why that is the problem

A top-down estimate takes a published category total and multiplies by a share. Every objection to it is an objection to the share, and the share was chosen to make the result look reasonable, so the conversation has nowhere to go. It is a number that cannot be wrong in any specific way.

Bottom-up sizing has parts. Count the addressable entities, take the fraction you can actually reach, apply the price you can actually charge, apply a realistic adoption rate over a stated period. Four or five terms, each of which a reviewer can challenge on its own without rejecting the whole exercise. That is the property that makes a business case survive contact with a CFO.

A worked shape, using invented figures so the arithmetic is visible: 40,000 registered firms in the target activity codes, of which 15 percent sit in the size band the product is built for, giving 6,000; a reachable fraction of 30 percent given the channels that exist today, giving 1,800; an annual contract value of 4,000 euros; and a five-year adoption ceiling of 8 percent, giving 144 customers and roughly 576,000 euros of annual recurring revenue at the end of the horizon. Every one of those five numbers is either a sourced count or an assumption with a name. The last two are assumptions, and a reviewer who disagrees with the 8 percent can change it without touching anything else.

## Where the counts actually come from

Official statistics and registers, in that order. Eurostat's structural business statistics publish enterprise counts by economic activity and size class across the EU, which is exactly the shape a bottom-up base needs, and national statistics offices publish comparable breakdowns for their own countries. Company registers go one level deeper: Companies House in the UK, and its equivalents elsewhere, hold filings that carry legal consequences for being wrong.

For narrower populations, the useful sources are regulator lists, professional body directories, trade association membership, and public procurement registers. For software categories, integration directories and public job adverts naming a specific tool give a floor rather than a total, which is still a real number and should be labelled as a floor.

What all of those share is that someone can check them. That is the entire selection criterion.

## The failure mode: a confident number with a fabricated source

A model asked for a market size will answer. It will often attach a publisher, a report title and a year, and the combination looks exactly like a citation because that is the shape it learned. Sometimes the report exists and the figure does not. Sometimes neither exists.

The mechanism behind this is worth understanding rather than moralising about. A model answering from recall is answering from a fixed point in the past, and vendors publish that point: Anthropic's model documentation lists both a training data cutoff and an earlier reliable knowledge cutoff per model, which is a direct statement that recency degrades before the cutoff is even reached. Market figures are exactly the kind of fact that changes annually and appears in thousands of secondary sources, which is the worst possible combination for recall.

So the rule is mechanical. Open every citation. Search the page for the number. If the URL does not resolve, or resolves and does not contain the figure, delete the claim. Do not go looking for a different source that agrees with it, because that is how a fabricated number acquires real-looking backing.

## Build the model so one number can be changed

The deliverable is a chain, not a figure. Each term gets its source, its value, and a plausible low and high. Then run the sensitivity: change one term at a time to its low and high and see which one moves the answer most.

In almost every sizing exercise the answer is dominated by one assumption, usually adoption rate or reachable fraction, and finding out which one tells you where to spend the next week of research. That is more valuable than the estimate itself, because it converts an unanswerable question into a specific one you can go and test with ten customer conversations.

## Where AI is genuinely strong: reading the market's own language

Sizing is the part everyone asks for. The part that more often changes a product is qualitative, and it is where a model earns its place. Public review sites, support forums, community threads, conference talks and job adverts contain how a market describes its problem in its own words, at a volume nobody reads by hand.

Write the taxonomy yourself from thirty examples you have read, have the model code the rest against it, then check a sample of its labels against your own. The output is a vocabulary you can use in positioning and messaging, grounded in what buyers said rather than in what your team believes they meant. The definitions for turning that into a position are in [ideal customer profile](https://builderscamp.com/guides/glossary/ideal-customer-profile) and [product positioning](https://builderscamp.com/guides/glossary/product-positioning).

## The sampling problem nobody mentions

Everything public is self-selected. Review sites over-represent the delighted and the furious, forums over-represent people with an unresolved problem, and job adverts describe what a company intends rather than what it does. A synthesis built on those sources is a synthesis of a skewed sample, and running it at ten times the volume makes it more precise about the same skew.

State the bias in the write-up rather than correcting for it silently, and treat the output as hypotheses for interviews. [How to validate a startup idea](https://builderscamp.com/guides/other/how-to-validate-a-startup-idea) covers the step after this one.

## Should you use a model to simulate the respondents?

Asking a model to role play a segment and answer your survey is fast, free and tempting, and it produces the answers most consistent with what has been written about that segment publicly. That is a description of published discourse, not of your buyers, and the two diverge exactly where a new product lives. Treat simulated responses as a way to pressure-test your question wording before real people see it, which is a real use, and never as a data source in the write-up.

The tell that this has gone wrong is agreement. Simulated respondents rarely contradict each other, and a survey where nobody disagrees has measured a model, not a market.

## Size the decision, then go and test the one assumption that matters

Pick the term in your chain with the widest range, design the cheapest test that narrows it, and run that before refining anything else. A sizing model with one assumption tested beats a prettier one with five assumed.

[See the Product Strategy bootcamp](https://builderscamp.com/bootcamps/product-strategy?utm_source=guide&utm_medium=organic&utm_campaign=ai-for-market-research)

Product Strategy runs two weeks, 4 live sessions and 8 microlessons, aimed at choices you can defend rather than analysis for its own sake. Business for Product Managers covers the revenue, cost and unit economics thinking that a sizing chain plugs into, in one week with 2 live sessions and 8 microlessons, and Product Marketing with AI covers the market to narrative to launch loop once the research is done. For the research tooling itself, see [Perplexity for product managers](https://builderscamp.com/guides/tools/perplexity-for-product-managers).

## Frequently asked questions

### Can AI produce a market size I can put in a business case?

Only if you build the arithmetic and it fills in the sourced counts. A model asked directly for a market size will produce a figure with a plausible-looking citation, and the figure is often a rounded recollection of a secondary source. A model asked to find the number of registered firms in a specific activity code, with a link, is doing research.

### What is wrong with a top-down market size?

It cannot be falsified. A number reached by taking a published category total and applying a percentage has no step a reviewer can disagree with, because the percentage was chosen to make the answer reasonable. A bottom-up chain has five or six terms, each of which someone can challenge individually, which is the point.

### How do I check a citation the model gave me?

Open it and search the page for the number. If the URL does not resolve, or resolves but does not contain the figure, delete the claim rather than looking for a replacement source that agrees with it. Searching for a source that confirms a number you already have is how a fabricated figure gets laundered into a real one.

### Where do bottom-up counts actually come from?

Official statistics and registers. Eurostat's structural business statistics publish enterprise counts by activity and size class across the EU, national statistics offices publish equivalents, and company registers such as Companies House in the UK hold filings that carry legal consequences for being wrong. Regulator lists and trade association directories cover narrower populations.

### Are paid market reports worth citing?

They move responsibility without reducing uncertainty. Most published category sizes are themselves modelled estimates whose method you cannot inspect, so citing one means inheriting assumptions you have never seen. If a report is the only source, cite it as an estimate with its publisher and year, and never use it as the base of a bottom-up chain.

### What is AI genuinely good at in market research?

Reading at volume. Public review sites, support forums, job adverts and conference talks contain the language a market actually uses about its problem, and a model will code hundreds of them against a taxonomy you wrote. That is qualitative work that was previously not being done, rather than work being done faster.

### Which Builders Camp bootcamp covers market sizing?

Business for Product Managers covers the business intuition underneath it: revenue, costs, unit economics and profit and loss thinking, across one week with 2 live sessions and 8 microlessons. Product Strategy covers the choice the sizing is meant to inform, across two weeks with 4 live sessions and 8 microlessons.

## Sources

- [Eurostat: structural business statistics](https://ec.europa.eu/eurostat/web/structural-business-statistics)
- [Companies House: search the UK register of company filings](https://find-and-update.company-information.service.gov.uk/)
- [Anthropic Docs: Models overview, published training and knowledge cutoffs](https://docs.claude.com/en/docs/about-claude/models/overview)
- [Builders Camp: Product Strategy bootcamp](https://builderscamp.com/bootcamps/product-strategy)
- [Builders Camp: Business for Product Managers bootcamp](https://builderscamp.com/bootcamps/business-for-product-managers)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
