Learn / RAG in 7 lessons / Production failure modes

Lesson 7 of 7 8 min

Production failure modes

The RAG failures that don't show up in a demo: silent misses, stale embeddings, citation-answer mismatch, and the retrieval trace that catches them.

Demos fail loudly; production fails quietly

A demo failure is obvious: wrong answer, right in front of you, in a session you’re actively watching. Production failure is different - the system keeps returning fluent, confident, correctly-formatted answers, and the only sign something’s wrong is a slow drift in user trust or a support ticket weeks later. That gap is why RAG needs deliberate detection mechanisms, not just “it looked fine when we shipped it.”

The failure modes below are the ones that actually show up once a RAG system has real traffic and real, changing data - covered in first-hand detail in Why Your RAG Pipeline Is Confidently Wrong, which this lesson summarizes the practical checklist from.

Silent retrieval misses

Retrieval returns something almost every time - there’s rarely a “no results” signal, just a ranked list, even when nothing in it is actually relevant. If nothing stops the model from answering anyway, you get a confident answer built on irrelevant context. Fix: explicitly check a relevance/score threshold before calling the model at all, and have the model refuse when context is weak (lesson 5).

Semantic near-misses

A chunk that’s topically adjacent but factually wrong for the specific question - e.g. the refund policy for annual plans retrieved when the question was about monthly plans - often scores nearly as high as the correct chunk, because they’re semantically similar. This is exactly what a golden set’s deliberate near-miss test cases (lesson 6) are built to catch.

Stale embeddings

If your indexing pipeline doesn’t re-run when source documents change - or when you upgrade to a newer embedding model - the retriever keeps confidently serving chunks that no longer match reality. Two concrete guards:

  • Re-index on a schedule or on write (whichever fits your data’s update frequency), and alert if the index’s last-updated timestamp gets too old.
  • Track which embedding model version produced each stored vector, and never mix vectors from two different models in the same similarity search - they don’t live in comparable spaces.

Citation-answer mismatch

The model cites [source: doc-12] but the actual claim in its answer isn’t what doc-12 says - it drifted while generating, or blended two sources into one sentence. Catch this with faithfulness evaluation (lesson 6): does each cited claim actually appear in its cited source, checked programmatically or by an LLM judge, not just “did it produce a citation at all.”

The fix that makes the rest debuggable: a retrieval trace

None of the above is fixable after the fact without visibility into what actually happened for a specific bad answer. Log, per request: the query, every candidate chunk considered and its score, which chunks made the final cut into the prompt, and the generated answer. When a user reports a wrong answer, this trace turns “something was wrong” into “retrieval scored the correct chunk 4th, below the top-3 cutoff” - a specific, fixable bug instead of a mystery.

Key takeaways

  • A RAG pipeline can fail silently - it returns a confident, well-formatted answer with no visible error, which is why these bugs survive so long in production.
  • Re-index on a schedule (or on write) and track embedding-model versions - a stale index quietly serving outdated chunks is one of the most common real-world causes.
  • Log a retrieval trace per request (which chunks were retrieved, their scores, what made it into the final prompt) - without it you cannot debug a bad answer after the fact.
  • See 'Why Your RAG Pipeline Is Confidently Wrong' for a deeper, first-hand production account of these failure modes.

Quick check

3 questions - see how much stuck.

1. Why are RAG failures often harder to catch than a normal application bug?
2. What's a 'stale embedding' problem?
3. What does a per-request retrieval trace let you do that you can't do without one?