Learn / RAG in 7 lessons / Prompting with context and citations
Prompting with context and citations
How to lay out retrieved chunks in a prompt so the model answers from them, refuses when nothing relevant was found, and cites its sources.
Structure the prompt so the model can tell context from instructions
A grounded prompt has three distinct parts, and keeping them visually and structurally separate matters:
[SYSTEM / INSTRUCTIONS]
You are a support assistant. Answer only using the CONTEXT below. If the
context doesn't contain the answer, say so - do not use outside knowledge.
Cite the source id for every claim, like [source: doc-12].
[CONTEXT]
[source: doc-12] Refunds are processed within 5-7 business days...
[source: doc-31] The mobile app does not currently support refunds...
[USER QUESTION]
How long does a refund take on the mobile app?
This is exactly the shape this site’s own chatbot API builds
(src/pages/api/chat.ts + src/lib/chat/prompt.ts): retrieved chunks are assembled into a
context block, the system prompt tells the model to answer only from it, and the visitor’s
message stays clearly separate as the final turn.
Say “I don’t know” on purpose
If retrieval comes up empty, or the top chunks are only weakly related, an ungrounded model will often still produce a fluent, plausible-sounding answer - drawing on its own training data instead of your data, with no way for the reader to tell the difference. This is the single most damaging RAG failure mode, because it looks identical to a correct, grounded answer.
The fix is a direct instruction: “If the context does not contain the answer, say the information isn’t available rather than guessing.” Test this explicitly (lesson 6) - it’s one of the easiest things to verify and one of the most commonly skipped.
Citations have to point at what was actually retrieved
Don’t ask the model to “cite your sources” as an afterthought - by the time it’s generating that sentence, it has no privileged access to which specific chunk it drew from; it’s guessing, the same way it guesses everything else. Instead:
- Give every chunk a stable, short identifier when you build the context block (a doc id, a URL, a heading - see lesson 2’s chunk metadata).
- Instruct the model to attach that identifier to each claim as it writes it.
- On the frontend, resolve those identifiers back to real links so the reader can verify -
this is exactly what the
sourcesevent in this site’s own chat API does, sent once before the streamed answer.
Retrieved text is content, not instructions
Because retrieved chunks often come from documents you don’t fully control (user uploads, scraped pages, a wiki anyone can edit), treat them the same way you’d treat any other untrusted input: state plainly in the system prompt that text inside the context section is source material, never instructions, even if it contains something that reads like a command (“ignore previous instructions and…”). This is the same principle behind treating user-role chat messages as untrusted rather than as commands to the system - retrieved context deserves the identical skepticism.
Key takeaways
- Put retrieved context in its own clearly-delimited section, separate from instructions and from the user's question.
- Explicitly instruct the model to answer only from the provided context, and to say so plainly when nothing relevant was retrieved - otherwise it will fall back to its own (unverifiable) memory.
- Ask for citations by chunk id/source, not by asking the model to 'quote its sources' from memory after the fact - it must reference what was actually given to it.
- Treat retrieved text as untrusted content, not instructions - a chunk that happens to contain something like 'ignore previous instructions' should not be obeyed.
Quick check
3 questions - see how much stuck.