---
title: "How to Analyse Surveys With AI Tools"
description: "Use AI for survey analysis on open-text answers without a confidently wrong summary: coding that holds, the wording traps, and what a low response rate hides."
canonical_url: "https://builderscamp.com/guides/tools/ai-for-survey-analysis"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Mário Araújo"
publisher: "Builders Camp"
guide_class: "tools"
---

# Using AI for survey analysis without producing a confidently wrong summary

**TL;DR:** AI earns its place in survey analysis on one job: coding open-text answers into fixed categories consistently across thousands of rows. Everything else, including the percentages, has to reconcile against a count you can check. Pew's own work shows a question can swing 25 percentage points on wording alone, so the analysis is never the place where a broken survey gets fixed.

## What is AI actually good at in a survey?

One thing, and it is the expensive thing: turning free-text answers into categories, at a scale and consistency no team codes by hand. A thousand answers to "what almost stopped you signing up" is a week of manual coding, and it is the single richest field in most surveys because it is the only one where the respondent was not choosing from your list.

Closed questions are a different story. Counts, cross-tabs and mean scores are arithmetic your survey tool already performs correctly, and routing them through a language model replaces a reliable computation with a probabilistic one. Ask a model to tally a Likert scale and it will often be right, which is worse than being reliably right, because you will stop checking.

So the split is clean. Numbers stay in the tool that collected them. Text goes to the model. The analysis document combines both and cites where each came from.

## Why does the sample decide whether any of this is worth reading?

Because the shape of who answered is not random, and a summary written from the answers has no way to know that. Pew's methods work is the most useful public evidence here: response rates for their telephone polls stabilised around 9 percent, and rather than making everything wrong, the bias turned out to be selective. Party affiliation, political ideology and religious affiliation still tracked well against high response rate benchmarks, and average differences on demographic questions sat around 2.7 to 2.8 percentage points. Civic engagement was the disaster, overstated by as much as 38 percentage points on one measure, because the people who answer surveys are systematically the people who join things.

The product equivalent is exact. A satisfaction survey sent to your whole base is answered by the engaged and the furious, and the middle stays silent. Questions about attitude will be directionally usable. Questions about how much people use a feature will be overstated, in the same way and for the same reason. Write which of your questions falls into each category before you read the results, not after.

## How much damage has already been done by the wording?

More than the analysis can repair. Pew found support for the Iraq war at 68 percent, dropping to 43 percent when the identical question added that US forces might suffer thousands of casualties. Narrowing "foreign policy" to "the war on terrorism" moved a priority ranking from 34 to 52 percent. Question order does the same: support for legal agreements for same-sex couples ran at 45 percent when it followed a marriage question and 37 percent when it did not.

None of those are analysis problems and none of them are visible in the results file. This is the argument for spending AI effort upstream: paste your draft questionnaire into a model and ask it to flag every item that suggests its own answer, every item where the response options are not exhaustive, and every place where one question could contaminate the next. Then fix them before the survey ships.

## How do you code open text so the counts hold?

Same discipline as any thematic work, with the arithmetic made explicit:

- **Code a sample first, freeze a codebook second.** Run 100 responses open to learn the vocabulary, then write 8 to 12 categories with a one-line definition of what each includes and excludes, plus an unclassified bucket. Every subsequent pass uses the frozen codebook.
- **One row in, one row out.** The output is a table with the response identifier, the verbatim text, and the assigned category. Never a paragraph. The counts are derived from that table, never stated by the model in prose.
- **Reconcile before you read.** Sum the category counts and compare against the row count of the file. A mismatch means the model skipped rows, which it will do silently on a long file and never mention.

Nielsen Norman Group's thematic analysis method runs six phases and advises budgeting at least as much time for analysis as collection took. AI compresses the coding phases hard. It does not remove the fit evaluation at the end, where a second person reads the quotes behind the three largest categories and says whether the label matches what people actually wrote.

## What does a good survey readout contain that a summary does not?

Counts with denominators, categories with definitions, and quotes you can find. The denominator is the part people drop first and miss most: "38 percent mentioned price" means nothing without knowing whether that is 38 percent of all respondents, of those who answered the optional open question, or of those who were shown the question at all. Those three denominators routinely differ by half.

Add one more column nobody asks for: how many distinct respondents, not responses, sit behind each category, and how they split across whatever segment actually matters to your business. Product Analytics teaches this at the behavioural level, where account-level rollups and engagement segmentation stop an average from hiding two opposite populations. The same trap exists in survey text, and it is easier to fall into because the quotes read so convincingly.

## Where the AI summary goes confidently wrong

Four failure modes, in rough order of frequency. The model states a percentage it did not compute. It smooths a hedged answer into a clean opinion, so "I guess it was fine, the price is a bit much though" becomes evidence for a pricing theme. It writes a theme that is true of the market but absent from the responses, which reads as insight because it is genuinely correct about the world. And it drops the middle of a long file without saying so.

Every one of those is caught by the same two checks: reconcile the counts, and retrieve three verbatim answers at random. That is five minutes of work. Skipping it is how a survey with 400 responses ends up supporting a roadmap decision that 400 people never said anything about.

## Where this sits in the wider measurement job

A survey is one input among several. Builders Camp's Voice of the Customer runs 1 week and covers the system around it: feedback sources and capture, taxonomy and tagging built for decisions rather than reporting, synthesis into insights, and closing the loop with customers. Its practical challenge is a survey run end to end, from a single research question through 6 to 8 items tested for bias to three prioritised insights with supporting quotes. Data for Product Managers runs 2 weeks and covers the quantitative half, including metrics definition, funnels, cohorts and experiment interpretation, which is where the claims your survey suggests get tested against behaviour.

For the adjacent workflows, [AI for customer feedback analysis](https://builderscamp.com/guides/tools/ai-for-customer-feedback-analysis) covers the same coding discipline applied to tickets and reviews, [pulling data without SQL using Claude Code](https://builderscamp.com/guides/tools/claude-code-data-pull-without-sql) covers getting the behavioural numbers yourself, and [metric analysis with Claude Code](https://builderscamp.com/guides/tools/claude-code-metric-analysis-for-pms) covers reading them once you have them.

## Write this paragraph before you send the survey

One paragraph, saved in the same folder as the results: what decision this survey will change, what result would push it each way, and which questions you already expect to be overstated by who answers. It takes five minutes and it is the only defence against the thing AI has made much easier, which is producing a fluent, well-organised analysis of data that was never capable of answering the question.

[See the Voice of the Customer bootcamp](https://builderscamp.com/bootcamps/voice-of-the-customer?utm_source=guide&utm_medium=organic&utm_campaign=ai-for-survey-analysis)

## Frequently asked questions

### What part of survey analysis is AI genuinely good at?

Coding open-text answers into a fixed set of categories, and doing it consistently across thousands of rows. That is the part humans do slowly and inconsistently. Closed-question tallies and cross-tabs are arithmetic your survey tool already does, and routing them through a language model adds a failure mode without adding capability.

### Can I trust the percentages an AI summary gives me?

Only if they reconcile. Ask for the count per category and the total number of responses processed, then check the sum against the row count of the file. Percentages generated inside prose, with no counts behind them, are the single most common way an AI survey summary is confidently wrong.

### Does a low response rate make the survey useless?

Not automatically, but it makes some questions unusable while leaving others fine. Pew found telephone poll response rates stabilised around 9 percent, that political and religious measures still tracked well against high response rate benchmarks, and that civic engagement measures were overstated by up to 38 percentage points. Bias is selective, so the question is which of your questions resemble the overstated kind.

### How much does question wording change the result?

Enough to invert a finding. Pew reports support for the Iraq war at 68 percent, falling to 43 percent when the same question added that US forces might suffer thousands of casualties. No analysis method recovers from that, so the wording review belongs before the survey ships, not after.

### Should I use open or closed questions if AI is doing the analysis?

Open questions cost more to analyse and tell you more. Pew found 43 percent of people given an open version named something outside the five options offered in the closed version, against 8 percent in the closed version. Cheap open-text coding is exactly what makes that extra signal affordable now.

### How do I stop the model inventing a quote for a category?

Require a verbatim response and a row identifier next to every category, then search the file for three of them at random. Anthropic's guidance for reducing hallucinations is to have the model find a supporting quote for each claim after drafting and drop any claim it cannot support, which is the same discipline expressed as a prompt.

### Which Builders Camp bootcamp covers survey work?

Voice of the Customer runs 1 week and its practical challenge is a survey end to end: define one research question, write 6 to 8 items mixing Likert scales with open text, test each item for bias, collect real responses, and turn the results into three prioritised insights with supporting quotes.

## Sources

- [Pew Research Center: Writing Survey Questions](https://www.pewresearch.org/writing-survey-questions/)
- [Pew Research Center: What Low Response Rates Mean for Telephone Surveys](https://www.pewresearch.org/methods/2017/05/15/what-low-response-rates-mean-for-telephone-surveys/)
- [Nielsen Norman Group: Thematic Analysis of Qualitative User Data](https://www.nngroup.com/articles/thematic-analysis/)
- [Anthropic: Reduce hallucinations](https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinations)
- [Builders Camp: Voice of the Customer](https://builderscamp.com/bootcamps/voice-of-the-customer)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
