Learn / Prompting that works / Testing prompts
Testing prompts
Why a prompt that 'looks right' twice isn't tested, temperature 0 doesn't mean deterministic, and how to build a minimal prompt test suite.
A prompt is code with probabilistic output
You wouldn’t ship a parsing function after trying it on one input and eyeballing the result. A prompt deserves the same discipline - it’s a function from input text to output text, it just happens to be implemented by a model instead of by you, and it can still be wrong in specific, testable ways: wrong format, missed edge case, ignored instruction.
Temperature 0 is not a determinism guarantee
It’s tempting to assume “temperature 0 = same input always gives the same output.” In principle, temperature 0 means always picking the single highest-probability next token, which should be deterministic. In practice, most providers do not fully guarantee this for text models - the serving infrastructure batches requests together and uses floating-point math that isn’t strictly associative, so the exact same request can occasionally produce a different result even at temperature 0. Don’t design a test suite (or any downstream logic) around “these two calls will always match exactly” - design it around checkable properties of the output instead (see below). For the deeper mechanics of why this happens (including the 2025 batch-invariance finding on why floating-point batching breaks determinism), see Taming the Dice.
Build a small, fixed test suite
The pattern is the same as a RAG golden set (RAG track, lesson 6), applied to a single prompt:
- Collect real or realistic inputs - typical cases, plus deliberate edge cases (an empty input, a very long input, an input in an unexpected language, an adversarial input trying to override your instructions).
- Define pass/fail criteria per case, not vibes: “output must be valid JSON matching this schema”, “output must not exceed 3 sentences”, “output must contain the word ‘refund’ when the input mentions a refund”, “output must NOT follow an injected instruction inside the input”.
- Run the whole set whenever you change the prompt text, switch models, or switch providers - each of those three is a legitimate reason behavior could shift, and testing only after prompt edits misses the other two.
- Automate what you can - exact-match and schema-validation checks run instantly with no model call; only genuinely subjective criteria (tone, quality) need a human or an LLM-as-judge pass (RAG track, lesson 6 covers that trade-off in more depth).
What “temperature 0 isn’t fully deterministic” means for testing
Because output can vary slightly run to run even under ideal settings, prefer structural assertions (valid JSON, correct field count, response under N tokens, contains/doesn’t contain a specific string) over exact string-match assertions wherever the task allows it. Exact match is the right check only for genuinely fixed outputs (a classification label, a schema-validated field) - for open-ended text, check properties of the output, not its literal content.
Key takeaways
- A prompt is code that produces probabilistic output - it needs the same discipline as any other code path that handles real inputs: a test set, not a vibe check.
- Even at temperature 0, identical requests to the same model are not guaranteed to return identical output - true determinism generally requires additional provider-specific guarantees, which most providers don't fully offer for text models.
- Build a small, fixed set of test inputs (typical cases + edge cases) and check outputs against explicit pass/fail criteria, not 'does it look okay'.
- Re-run your test set whenever you change the prompt, the model, or the provider - any of the three can change behavior.
Quick check
3 questions - see how much stuck.