Learn / RAG in 7 lessons / Chunking
Chunking
Why documents need to be split before they're indexed, the trade-offs between chunk sizes, and a runnable toy chunker.
Why chunk at all
A retriever compares a question’s embedding against the embeddings of your stored text and returns the closest matches. If you embed a 40-page document as one vector, that vector is an average of everything the document talks about - a specific question about page 30 matches it about as poorly as a question that has nothing to do with the document at all.
Chunking splits source documents into smaller, semantically coherent pieces before indexing, so each stored vector represents one focused idea and can be matched precisely.
The size trade-off
- Chunks too small (a single sentence): each one embeds precisely, but loses surrounding context - “it improved by 40%” means nothing without the sentence before it that says what “it” is.
- Chunks too big (a full section or more): context is preserved, but the embedding blurs multiple ideas together, and you waste prompt tokens shipping irrelevant text alongside the one relevant sentence the model actually needed.
A common starting point for prose is 500-1500 characters (roughly 100-300 words), with 10-15%
overlap between consecutive chunks - close to the CHUNK_SIZE/CHUNK_OVERLAP values this very
site’s own chat-index builder uses (src/lib/chat/chunker.ts). Tune from there based on how
specific your real questions are.
Split on structure first, characters last
Naively cutting every N characters will slice a sentence in half. A better chunker respects structure:
- Split on headings or paragraph breaks first.
- If a resulting piece is still too big, split on sentence boundaries.
- Only fall back to a hard character cut if a single sentence is itself longer than the limit.
Here’s a minimal version of that idea - paragraph-aware, with overlap, using only the Python standard library. Click Run to execute it right here in your browser (via Pyodide - no server involved).
:::pyrun
def chunk_text(text, max_chars=200, overlap=40):
paragraphs = [p.strip() for p in text.split("\n\n") if p.strip()]
chunks = []
buffer = ""
def flush():
if buffer.strip():
chunks.append(buffer.strip())
for para in paragraphs:
if len(buffer) + len(para) + 1 <= max_chars:
buffer = f"{buffer} {para}".strip()
continue
# Current buffer is full - flush it, then start the next one with
# a small tail of overlap so a boundary fact isn't stranded.
flush()
tail = buffer[-overlap:] if buffer else ""
buffer = f"{tail} {para}".strip()
flush()
return chunks
document = """RAG systems retrieve text before generating an answer.
Retrieval quality depends heavily on how the source text was chunked.
Chunks that are too small lose context. Chunks that are too big dilute relevance.
A good default is a few hundred characters with light overlap between chunks."""
for i, c in enumerate(chunk_text(document)):
print(f"[chunk {i}] ({len(c)} chars): {c}")
Run it, then try shrinking max_chars to 80 and re-running - watch sentences start getting
split mid-thought, which is exactly the failure lesson 4 (retrieval quality) revisits from the
retrieval side.
Key takeaways
- Chunk size is a trade-off: too small loses context, too big dilutes relevance and wastes prompt tokens on a partial match.
- Split on natural boundaries (paragraphs, headings, sentences) before falling back to a hard character limit.
- A small overlap between consecutive chunks stops a fact from being cut exactly at a chunk boundary.
- Chunk size should roughly match how specific your questions are - broad questions want bigger chunks, precise lookups want smaller ones.
Quick check
3 questions - see how much stuck.