Builders Camp

Glossary

What Is Few-Shot Prompting?

Few-shot prompting means putting a few worked examples, typically 3 to 5, into the prompt before the real input, so the model copies a demonstrated pattern instead of guessing at one. Start zero-shot, and switch to few-shot when the output drifts in format or applies your labels inconsistently.

What does few-shot prompting mean?

Few-shot prompting is a technique where you include a small set of solved examples in the prompt, then give the model a new input to solve the same way. IBM defines it as "the process of providing an AI model with a few examples of a task to guide its performance," and contrasts it with zero-shot prompting, "which requires no examples" (IBM). Anthropic's prompting documentation is more specific about the payoff: "A few well-crafted examples (known as few-shot or multishot prompting) improve accuracy and consistency," and it recommends 3 to 5 examples for best results (Anthropic prompting best practices).

The mechanism is that examples carry information instructions cannot. "Write in a friendly tone" leaves the model to guess what friendly means to you. Two replies you already consider friendly settle it.

Where does the term few-shot come from?

The 2020 GPT-3 paper put the term in its title: "Language Models are Few-Shot Learners." Tom Brown and colleagues trained a model with 175 billion parameters and tested it by writing the task and a handful of demonstrations straight into the prompt. In their words, "GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model" (arXiv:2005.14165).

That sentence is why the technique matters to product teams. Before it, teaching a model a new task meant collecting a labelled dataset and retraining. After it, the same result could often come from a paragraph of examples that a PM can write in ten minutes.

Zero-shot or few-shot: which should you use?

Start with zero-shot prompting, a clear instruction and no examples, and move to few-shot only when the output tells you to. The table is the decision in practice.

What you see in the output Pick Why
Correct and in the right shape Zero-shot Examples would add length and cost for no gain
Correct content, wrong or inconsistent format Few-shot A shown format is copied more reliably than a described one
Your labels applied differently from how your team uses them Few-shot Examples define the boundary between labels
Tone or house style keeps missing Few-shot Style is easier to show than to describe
Wrong because the task needs several reasoning steps Neither alone Add chain of thought prompting or split the task
Wrong because the model lacks facts about your product Neither alone Add the facts as context; examples cannot supply missing knowledge

The last two rows are where teams lose time. Adding examples to a prompt that fails for lack of context produces confident, well-formatted wrong answers.

How do you turn a zero-shot prompt into a few-shot one?

Take a common PM task: tagging incoming feedback for a fictional project management app so the weekly review can count themes. The zero-shot version looks reasonable.

Tag this customer message with one label: Bug, Feature request,
Billing, or How-to. Reply with the label only.

Message: """I can't export my board to PDF anymore."""

Run it on a week of real messages and the problem shows up in the ambiguous ones. "I can't export to PDF" is a Bug if export used to work and a Feature request if it never existed. "Why was I charged for 12 seats?" is Billing to finance and a How-to to support. The model picks one each time, and not the same one each time.

The few-shot version keeps the instruction and adds labelled examples, each wrapped in delimiters so the model can tell examples from the real input. Anthropic recommends wrapping examples in tags for exactly this reason.

Tag the customer message with one label: Bug, Feature request,
Billing, or How-to. Reply with the label only.

<examples>
<example>Message: "Export to PDF worked last week, now it spins forever."
Label: Bug</example>
<example>Message: "Would love to export boards as PDF for clients."
Label: Feature request</example>
<example>Message: "Why was I charged for 12 seats when we have 9 users?"
Label: Billing</example>
<example>Message: "How do I add a guest who can only view one board?"
Label: How-to</example>
</examples>

Message: """I can't export my board to PDF anymore."""

Four examples, one per label, and the first two sit on either side of the exact boundary that caused trouble. The word "anymore" in the new message now points at Bug, because the examples showed that "worked before" is what separates the two labels.

How do you choose the examples?

Anthropic's guidance lists three properties: examples should be relevant to the real use case, diverse ("Cover edge cases and vary enough that Claude doesn't pick up unintended patterns"), and structured so they are clearly separated from the instructions (Anthropic prompting best practices).

Two failure modes follow from getting the set wrong. If three of your four examples carry the same label, the model leans toward that label on anything uncertain. If every example follows the same order, the model can read the order as part of the task. The fix for both is boring: one example per label where you can, and a shuffled order.

Pull examples from real inputs, not invented ones. A real message has the typos, half-sentences and pasted screenshots text that your production inputs have, and an invented one is usually cleaner than anything the model will see.

How does the example-thread version work in the API?

In a chat window, examples go inside one message. In an API feature, there is a second option: write the examples as earlier turns of the conversation. Anthropic's documentation says "Earlier conversational turns don't necessarily need to actually originate from Claude. You can use synthetic assistant messages" (Anthropic: Working with the Messages API).

For the tagging feature, the thread would be a system prompt with the instruction, then four invented user turns each holding one example message, each followed by an invented assistant turn holding only the label. The real message comes last. Because the model has already "answered" four times with a bare label, it answers the fifth the same way, which is often enough to stop it adding explanations you did not ask for.

When does few-shot prompting stop being the right tool?

Examples are paid for on every request. Four short examples cost little, but ten long example documents are sent, and billed, again for every item you process. At high volume that is a reason to test whether a sharper instruction gets the same result with fewer examples.

Examples also teach surface features. If every Bug example mentions an error message, the model may start treating "error" as the signal instead of the broken behaviour. When the set keeps needing patches like that, the task usually wants a clearer label definition in the prompt, or a structured template, more than it wants a fifth example. The prompt template for product managers shows where examples sit alongside role, context and format.

Where does the AI Prompting for Product bootcamp fit?

AI Prompting for Product is a 1 week Builders Camp bootcamp with 2 live sessions and 19 self-paced microlessons, directed by Andre Albuquerque. Its public syllabus lists prompt structure and evaluation and iteration loops among its topics, including building quick rubrics and test cases to measure prompt quality.

See the AI Prompting for Product bootcamp

For the wider foundation, start with prompt engineering.

Before you add a fifth example to a prompt that keeps failing, pull the ten inputs it got wrong and label them yourself first. If you and a colleague disagree on three of them, the problem is the label definitions, and no number of examples will fix a boundary your own team has not agreed on.

Bootcamps referred in this Guide

Frequently asked questions

How many examples count as few-shot prompting?

Two or more. One example is usually called one-shot and none is zero-shot. Anthropic's prompting guidance recommends 3 to 5 examples for best results, which is enough to show a pattern without spending the context on near-duplicates.

Does few-shot prompting always beat zero-shot prompting?

No. On a simple task the model already handles well, examples add length and cost without changing the output. Few-shot earns its place when zero-shot output drifts in format or applies your labels inconsistently.

What makes a good example in a few-shot prompt?

An example that looks like your real inputs, including the messy ones. A set of clean, similar examples teaches a narrower pattern than the task needs, while one deliberately ambiguous example that shows how you want the tie broken often does more than three easy ones.

Can biased examples make few-shot prompting worse?

Yes. If most examples carry the same label, the model leans toward that label. If the examples always appear in the same order, it can pick up the order as if it were part of the task. Balance the labels and vary the order.

Is few-shot prompting a form of fine-tuning?

No. Fine-tuning changes a model's weights with a training dataset. Few-shot prompting only changes what goes into one request, and the original GPT-3 paper described it as working without any gradient updates or fine-tuning. The next request without the examples gets default behaviour.

What is the difference between few-shot prompting and a fake thread?

They teach the same way. A fake thread, or example thread, delivers the examples as invented user and assistant turns in the conversation history instead of pasting them into one message. It suits API features where every request should produce the same format.

When should a PM stop pasting examples into chats?

When the same examples go into more than a handful of chats. That repetition means the examples belong in a saved system prompt or a reusable template, so every run uses the same set and you can improve it in one place.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque

Last updated 2026-09-27

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the AI Prompting for Product bootcamp