Learn / Prompting that works / Examples and few-shot prompting

Lesson 2 of 5 6 min

Examples and few-shot prompting

When showing examples beats describing rules, how many to use, and the trap of examples that are too similar to each other.

Zero-shot vs few-shot

Zero-shot: just instructions, no examples. Works well for tasks the model has clearly seen a lot of during training and where the format is standard - summarizing a paragraph, translating a sentence, answering a factual question.

Few-shot: instructions plus a handful of example input/output pairs. Earns its extra tokens when:

  • the output format is specific and non-standard (a particular JSON shape, a custom tagging scheme, a house style),
  • the task is genuinely ambiguous from a description alone, and an example resolves the ambiguity faster than more words would,
  • you’ve tried zero-shot and the model’s format drifts across runs.

A concrete comparison

Zero-shot, describing the format in words:

Classify the sentiment of this review as positive, negative, or mixed.
Return only the label.

Few-shot, showing it:

Classify the sentiment as positive, negative, or mixed. Return only the label.

Review: "Fast shipping, exactly as described."
Label: positive

Review: "Works fine but the app crashes weekly."
Label: mixed

Review: "Broke on day two, refund denied."
Label: negative

Review: "Comfortable chair, arrived a week late though."
Label:

The few-shot version doesn’t just state the three labels - it shows the model how close a call “mixed” is meant to catch (a genuine pro and con, not just any imperfect review), which is exactly the kind of nuance that’s hard to fully specify in a one-line instruction.

How many examples, and which ones

1-5 examples is the typical useful range; more than that usually helps less per token spent, and adds latency and cost (lesson 5) for diminishing returns. Choose them deliberately:

  • Vary surface features that don’t matter (length, phrasing, topic) so the model can’t latch onto an accidental pattern instead of the real rule.
  • Include one edge case, not just three easy, obviously-typical ones - it’s the difference between “the model handles the 80% case” and “the model handles the case that actually causes tickets.”
  • Keep the format of the examples identical to the format you want back - if your examples end with Label: followed by the answer, the model will continue that exact pattern, which is the whole point.

Key takeaways

  • Few-shot prompting (showing 1-5 example input/output pairs) works especially well for format and tone, where 'show, don't tell' beats a written description.
  • Zero-shot (no examples) is fine for simple, well-known tasks; few-shot earns its extra tokens when the desired output shape is unusual or hard to describe precisely.
  • Vary your examples - if every example looks the same, the model will overfit to surface patterns instead of learning the actual rule you meant.
  • Include at least one edge case among your examples, not just the easy, typical case - it teaches the model how to handle the boundary, not just the middle.

Quick check

3 questions - see how much stuck.

1. When does few-shot prompting earn its cost most clearly?
2. What goes wrong if all your few-shot examples look very similar to each other?
3. Why deliberately include an edge case in your few-shot examples?