---
title: "How to Generate Test Cases With AI"
description: "Use AI for QA test cases: expand acceptance criteria into positive, negative and boundary cases, and learn the eight edge cases a model reliably fails to see."
canonical_url: "https://builderscamp.com/guides/tools/ai-for-qa-test-cases"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Tiago Pedro da Costa"
publisher: "Builders Camp"
guide_class: "tools"
---

# AI for QA test cases: expanding acceptance criteria, and what the model will miss

**TL;DR:** A model turns one acceptance criterion into twenty test cases in seconds, and every one of them tests what the criterion says rather than what the system does. Generate the expansion, then add the eight categories it structurally cannot see, starting with pre-existing data and two users acting at once. Builders Camp covers the delivery discipline around this in Project Management for Product, a 1 week bootcamp on planning, dependencies, risk and cadence.

## What can a model do with acceptance criteria?

Expand them, exhaustively and fast. Give it a criterion written as a condition, an action and an expected result, and it will return the positive case, the negative cases, the boundary values on either side of every threshold, and the combinations across inputs that a person writing by hand stops generating at about item nine.

That is a real contribution, because the combinations are where defects hide and the combinations are exactly what human attention degrades on. Three inputs with four states each is 64 combinations; nobody writes 64 rows by hand and nobody should, but starting from 64 and deleting the meaningless ones is a better process than starting from zero and remembering.

The limitation sits in the same sentence. The model is testing the criteria. It has no access to the system, so it cannot test what the system does, only what the document says it should do. Everything useful about this technique and everything dangerous about it follows from that one fact.

## How do you write criteria that expand well?

Specific thresholds, named error behaviour, and no adjectives.

"The form should handle invalid input gracefully" expands into nothing useful, because there is no value to sit on either side of and no stated result to assert. "If the email field does not contain an at sign, the form does not submit and shows the message Enter a valid email address below the field" expands into a dozen cases with obvious pass conditions. The [acceptance criteria](https://builderscamp.com/guides/glossary/acceptance-criteria) you would want anyway are the same ones that make generation work, which is a convenient alignment: the technique rewards the discipline it depends on.

This is also the fastest ambiguity detector available to a product manager. Generate the expansion and read it. Every case you cannot confidently mark pass or fail is a place the criteria are underspecified, found before an engineer has spent a day implementing your guess. That makes the exercise worth running even on a feature whose tests someone else will own. [How to write acceptance criteria with AI](https://builderscamp.com/guides/templates/how-to-write-acceptance-criteria-with-ai) covers the drafting side of the same loop.

## What does the model structurally miss?

Everything not implied by the text you gave it. In practice that clusters into eight categories, and they recur across products because they come from the world rather than from the spec.

Data that predates the feature is the first and the most damaging: existing rows with null values in a column your feature assumes, users who signed up before the field existed, a half-finished migration. Second, two actors at once: two people editing the same record, a webhook arriving while a user is mid-save. Third, permission boundaries: the same action attempted by an admin, a member, a suspended account, and a user who was removed from the workspace ten seconds ago.

Then money and rounding, which behaves differently from every other number and does so silently. Then time: time zones, daylight saving transitions, a date entered in one locale and displayed in another. Then third-party calls, which do not just succeed or fail but also time out, half-succeed, and succeed twice. Then repetition of an irreversible action, the double-clicked submit that charges a card twice. And finally the silent case, where the system neither errors nor completes and the user is left looking at an unchanged screen with no information at all.

The formal testing literature has names and techniques for most of this. The ISTQB glossary and its Foundation Level syllabus cover boundary value analysis, equivalence partitioning and state transition testing directly, and a model prompted with those technique names produces noticeably better expansions than one asked simply for test cases.

## How should a product manager actually use the output?

As candidates, never as a suite. This is the part teams get wrong when generation makes production free: the cost of a test is not writing it, it is maintaining it on every future change, and a generated set is unusually good at producing near-duplicates that all need updating when one label moves.

So the review step has three questions, and a case that fails any of them gets cut:

- **Would this catch a regression anybody would care about?** If the answer is no, the row is coverage theatre.
- **Does it assert a specific result, or does it assert that something happened?** Vague assertions pass while broken.
- **Can it be traced to a line in the criteria or a stated rule?** Cases with no source are the model filling gaps from pattern rather than from your system, which is the behaviour Anthropic's guidance on reducing hallucinations addresses by grounding the model in supplied documents.

Ownership matters here too. A PM is not writing the test suite, and a generated pile dropped on an engineering team reads as work created rather than work done. The useful handoff is the criteria, the expansion, and the list of cases you could not answer, which is a specification conversation rather than a QA instruction.

## Where does this fit in delivery?

Alongside the definition of done rather than after it. Project Management for Product teaches planning, dependency management, risk and delivery cadence as the mechanics that keep a release honest, and test case generation sits inside the planning half of that, not the end. Running the expansion at refinement time surfaces the ambiguity while the scope can still change; running it the week before release surfaces the same ambiguity when the only remaining options are slip or ship with a known gap.

The Product Delivery Specialist Track covers that cadence across several bootcamps, and the connection worth making explicit is to [definition of done](https://builderscamp.com/guides/glossary/definition-of-done): generated cases are only meaningful if the team has already agreed what finished means, otherwise the suite becomes the definition by accident.

## What about testing an experiment rather than a feature?

Different failure mode, and routinely skipped. An A/B test that is instrumented wrong does not fail loudly; it produces a confident number. Before the experiment runs, the checks that matter are assignment correctness, no user landing in both variants, events firing exactly once per action, and guardrail metrics recording for both arms.

A model will generate that checklist well, because the cases are mechanical and well documented in the testing literature. What it cannot tell you is whether the metric you chose will answer your question, which is the judgment A/B Testing for Product Managers spends its time on: hypotheses that are testable, primary versus secondary metrics, and when a mixed result means iterate rather than ship.

## The case you should write first

Before any of the expansion, write one case by hand: the most valuable single path through the feature, end to end, as a real user would do it with real data. Generated cases are strongest at breadth and weakest at sequence, because a criterion describes a moment and a user experiences a journey. If you only ever add one hand-written test to the generated set, make it that one, and run it against production data shapes rather than the clean fixtures the model assumed.

[See the Project Management for Product bootcamp](https://builderscamp.com/bootcamps/project-management-for-product?utm_source=guide&utm_medium=organic&utm_campaign=ai-for-qa-test-cases)

When the cases start failing, [Claude Code bug triage for PMs](https://builderscamp.com/guides/tools/claude-code-bug-triage-for-pms) covers turning a failure into a reproducible report an engineer can act on.

## Frequently asked questions

### Can AI write test cases from a user story?

From acceptance criteria, yes, and well. From a user story alone it will invent the criteria first and then test its invention, which produces a tidy suite for a feature nobody specified. Write the criteria, then generate. The quality of the output tracks the specificity of the input almost exactly.

### What does a model actually add over writing cases by hand?

Coverage of the combinations people skip. Given three inputs with four states each, the mechanical expansion is 64 rows and a human will write the nine they can imagine. The model produces the full grid, you delete the meaningless rows, and the deleting is faster and more reliable than the remembering.

### What edge cases does a model reliably miss?

The ones not implied by the text it was given: data that predates the feature, two users acting at once, permission boundaries across roles, money rounding and currency, time zones and daylight saving, third-party calls that time out or half-succeed, repeated submission of an irreversible action, and the silent case where nothing happens at all. None of those appear in acceptance criteria, so none appear in the expansion.

### Is a product manager supposed to write test cases?

Not the suite. A PM owns whether the acceptance criteria describe the right behaviour, and generated cases are a fast way to find out: if the expansion contains a case you cannot answer, the criteria are ambiguous and that is a specification defect found before any code was written.

### Should generated cases go straight into the test suite?

No. Every test has a maintenance cost paid on every future change, and a generated suite is unusually good at producing near-duplicates that each need updating. Treat the output as candidates and keep the ones that would actually catch a regression you care about.

### How do I stop the model inventing system behaviour?

Give it the criteria, the relevant interface description and any error contract, and require every case to cite the line it came from. Anthropic's guidance on reducing hallucinations recommends exactly this pattern of grounding in supplied documents plus permission to say it does not know. Cases with no citation get reviewed, not merged.

### Does any of this apply to testing an experiment?

Yes, and it is frequently skipped. Before an A/B test runs, the instrumentation needs its own checks: correct assignment, no user seeing both variants, events firing once, guardrail metrics recording. A broken experiment produces a confident result, which is worse than no result.

## Sources

- [ISTQB: Glossary of testing terms](https://glossary.istqb.org/)
- [ISTQB: Certified Tester Foundation Level](https://www.istqb.org/certifications/certified-tester-foundation-level/)
- [Anthropic: Reduce hallucinations](https://docs.anthropic.com/en/docs/test-and-evaluate/strengthen-guardrails/reduce-hallucinations)
- [Builders Camp: Project Management for Product bootcamp](https://builderscamp.com/bootcamps/project-management-for-product)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
