When a production RAG system gives wrong answers, check retrieval before you blame the model. Fix how you find and rank context first.
You built a RAG (Retrieval-Augmented Generation) demo. On your machine, with a handful of PDFs, it looked convincing and the answers felt coherent.
Then you pushed it to production.
Where the demo breaks
Users report that the LLM is “hallucinating” when the real issue is retrieval. Obvious answers go missing even though they’re in the docs. Irrelevant chunks surface because they’re semantically close but useless.
A retrieval design problem
If your RAG system feels unreliable, look at retrieval design first. Most production RAG systems fail because they lean too hard on vector search and mistake a strong demo for a dependable system.
For the broader catalogue of ways this goes wrong, see 7 mistakes you’re making with your production RAG stack. One pattern keeps showing up in my own work: a demo proves possibility, but production demands precision.
The “vector noise” trap
In a demo, semantic resemblance often feels good enough. In production, “good enough” is where failures begin.
Embeddings are useful. They map text into vectors so you can retrieve by meaning instead of exact wording. Semantic similarity still differs from retrieval accuracy.
Strong on concepts, weak on specifics
Vector search is good at finding related concepts and weak at specificity.
If a user searches for “Project-X-99 deployment logs,” vector search might return “Project-A deployment” or “logging best practices,” because they’re semantically close. It can miss the exact identifier “X-99,” which carries little semantic weight in a high-dimensional space.
The model inherits the drift
Once retrieval drifts, the LLM inherits the drift. It can’t reason its way out of missing or irrelevant context. You pay for tokens that produce confident but unhelpful answers, and users lose trust for a reason that sits one layer below the model.
The fix: hybrid search (vector plus BM25)
Combine semantic retrieval with lexical precision. That combination is hybrid search.
What is BM25?
BM25 (Best Matching 25, from the Okapi probabilistic retrieval framework) is the standard lexical ranking method behind classic search. It ignores meaning and rewards exact terms by how important they are in a document and across the collection, while accounting for document length and diminishing returns from repeated terms.
Why you need both
- Vector search handles synonyms, multilingual queries and conceptual matches.
- BM25 handles exact matches: IDs, SKUs, product codes and technical jargon.
User questions are rarely pure meaning or pure keyword. They’re usually a mix. If you’re still deciding how to structure retrieval in the first place, vector search vs graph search is a useful companion read.
| Query type | Vector search | BM25 | Hybrid |
|---|---|---|---|
| ”How do I cancel my plan?” | Strong | Decent | Strong |
| ”Error E-4012 on checkout” | Often misses the code | Strong | Strong |
| ”SKU HB-220 return policy” | Often misses the SKU | Strong on the SKU | Strong |
| ”Ways to speed up the store” | Strong | Weak on synonyms | Strong |
Reciprocal Rank Fusion (RRF)
Running two retrieval strategies creates a new question: how do you combine them?
A practical answer is Reciprocal Rank Fusion (RRF), introduced by Cormack, Clarke and Büttcher in 2009. It’s simple and reliable, and it doesn’t pretend that scores from different retrievers are directly comparable.
How it works
- Score by rank. For every document returned by either method, compute a rank-based score.
- The formula.
score = 1 / (k + rank), whererankis the 1-based position in each result list (so the code below adds1to a 0-based index). The constantk(60 in the original paper) keeps top-ranked items from dominating too much. - Sum it up. If a document appears in both the vector and BM25 lists, its scores add together.
- Sort. The documents with the highest combined scores go to the LLM.
Here’s the minimal PHP version I drop into a Laravel service:
function reciprocalRankFusion(array $resultSets, int $k = 60): array
{
$scores = [];
foreach ($resultSets as $results) {
foreach ($results as $rank => $docId) {
$scores[$docId] = ($scores[$docId] ?? 0.0) + 1 / ($rank + 1 + $k);
}
}
arsort($scores);
return $scores;
}
A document that is semantically relevant and lexically precise now rises to the top for a reason.
The second pass: re-rankers
Hybrid search is a strong foundation, but production RAG usually needs one more layer of judgment: a re-ranker.
Judge the pair
A re-ranker like Cohere Rerank or BGE-Reranker is a cross-encoder that reads the query and the document together. Relevance is relational: the question is whether this document answers this question, beyond what it happens to contain.
- Retrieve the top 50 results with hybrid search.
- Pass those 50 through the re-ranker.
- Send only the top 5 to your LLM.
This cuts context stuffing and improves what reaches the model. It’s one of the clearest differences between a demo and a production system that behaves consistently. The extra hop adds latency, which is where a Redis and semantic caching layer earns its keep.
Demo pipeline
Production pipeline
Your production RAG checklist
A RAG system can impress in a demo and still be weak in production. Once real users, messy documents and ambiguous queries arrive, weak retrieval turns the LLM into expensive guesswork.
The fixes I rely on
- Stop relying on vector-only search. Add a BM25 layer.
- Use RRF. Fuse lexical and semantic results without fiddly score calibration.
- Tune chunking on purpose. Chunks that are too small lose context, and chunks that are too large add noise. For technical docs, 512 to 1,024 tokens with 10 to 15% overlap usually works well for me.
- Add a re-ranker. Refine the final candidates before anything reaches the LLM.
- Evaluate with RAGAS. Track faithfulness, response relevancy, context precision and context recall instead of trusting intuition.
Key takeaways
- Retrieval first. Most “hallucinations” in production RAG start with bad context.
- Hybrid beats vector-only. BM25 catches the IDs and codes embeddings miss.
- Fuse with RRF. Rank-based fusion avoids comparing incompatible scores.
- Re-rank before you generate. Fifty candidates in, five good ones out.
- Measure it. RAGAS scores beat gut feel.
Building a RAG demo is easy. Building reliable RAG takes a real understanding of retrieval, ranking and context design. If you’re still choosing your foundation, picking the right RAG stack walks through the vector database trade-offs. If your production RAG keeps missing, here’s how I help teams fix retrieval, ranking and evals, or get in touch with your war story.
Where does your own system still behave like a demo when it should be behaving like production?