Skip to content
ansezz.
← Back to blog
AI Mar 8, 2026 7 min read 1,310 words

Why your RAG implementation is failing in production

Vector-only retrieval quietly breaks production RAG. Hybrid search, BM25, rank fusion, re-rankers and evals: the fixes that make it reliable.

Anass Ez-zouaine

Backend · Architect · AI

▸ Share

Pop-art illustration of a RAG pipeline breaking under production load
▸ On this page (5)

When a production RAG system gives wrong answers, check retrieval before you blame the model. Fix how you find and rank context first.

You built a RAG (Retrieval-Augmented Generation) demo. On your machine, with a handful of PDFs, it looked convincing and the answers felt coherent.

Then you pushed it to production.

Where the demo breaks

Users report that the LLM is “hallucinating” when the real issue is retrieval. Obvious answers go missing even though they’re in the docs. Irrelevant chunks surface because they’re semantically close but useless.

A retrieval design problem

If your RAG system feels unreliable, look at retrieval design first. Most production RAG systems fail because they lean too hard on vector search and mistake a strong demo for a dependable system.

For the broader catalogue of ways this goes wrong, see 7 mistakes you’re making with your production RAG stack. One pattern keeps showing up in my own work: a demo proves possibility, but production demands precision.

The “vector noise” trap

In a demo, semantic resemblance often feels good enough. In production, “good enough” is where failures begin.

Embeddings are useful. They map text into vectors so you can retrieve by meaning instead of exact wording. Semantic similarity still differs from retrieval accuracy.

Strong on concepts, weak on specifics

Vector search is good at finding related concepts and weak at specificity.

If a user searches for “Project-X-99 deployment logs,” vector search might return “Project-A deployment” or “logging best practices,” because they’re semantically close. It can miss the exact identifier “X-99,” which carries little semantic weight in a high-dimensional space.

The model inherits the drift

Once retrieval drifts, the LLM inherits the drift. It can’t reason its way out of missing or irrelevant context. You pay for tokens that produce confident but unhelpful answers, and users lose trust for a reason that sits one layer below the model.

Comic panel of a sweating grey robot at a desk buried in paper scraps that a grinning yellow dog-like robot shovels onto it, while a purple robot holds up a ticket and waits
Near-miss chunks bury the model.

The fix: hybrid search (vector plus BM25)

Combine semantic retrieval with lexical precision. That combination is hybrid search.

What is BM25?

BM25 (Best Matching 25, from the Okapi probabilistic retrieval framework) is the standard lexical ranking method behind classic search. It ignores meaning and rewards exact terms by how important they are in a document and across the collection, while accounting for document length and diminishing returns from repeated terms.

Why you need both

  • Vector search handles synonyms, multilingual queries and conceptual matches.
  • BM25 handles exact matches: IDs, SKUs, product codes and technical jargon.

User questions are rarely pure meaning or pure keyword. They’re usually a mix. If you’re still deciding how to structure retrieval in the first place, vector search vs graph search is a useful companion read.

Query typeVector searchBM25Hybrid
”How do I cancel my plan?”StrongDecentStrong
”Error E-4012 on checkout”Often misses the codeStrongStrong
”SKU HB-220 return policy”Often misses the SKUStrong on the SKUStrong
”Ways to speed up the store”StrongWeak on synonymsStrong
Embeddings capture what a question means. Keywords catch exactly what it names.

Reciprocal Rank Fusion (RRF)

Running two retrieval strategies creates a new question: how do you combine them?

A practical answer is Reciprocal Rank Fusion (RRF), introduced by Cormack, Clarke and Büttcher in 2009. It’s simple and reliable, and it doesn’t pretend that scores from different retrievers are directly comparable.

How it works

  1. Score by rank. For every document returned by either method, compute a rank-based score.
  2. The formula. score = 1 / (k + rank), where rank is the 1-based position in each result list (so the code below adds 1 to a 0-based index). The constant k (60 in the original paper) keeps top-ranked items from dominating too much.
  3. Sum it up. If a document appears in both the vector and BM25 lists, its scores add together.
  4. Sort. The documents with the highest combined scores go to the LLM.

Here’s the minimal PHP version I drop into a Laravel service:

function reciprocalRankFusion(array $resultSets, int $k = 60): array
{
    $scores = [];

    foreach ($resultSets as $results) {
        foreach ($results as $rank => $docId) {
            $scores[$docId] = ($scores[$docId] ?? 0.0) + 1 / ($rank + 1 + $k);
        }
    }

    arsort($scores);

    return $scores;
}

A document that is semantically relevant and lexically precise now rises to the top for a reason.

The second pass: re-rankers

Hybrid search is a strong foundation, but production RAG usually needs one more layer of judgment: a re-ranker.

Judge the pair

A re-ranker like Cohere Rerank or BGE-Reranker is a cross-encoder that reads the query and the document together. Relevance is relational: the question is whether this document answers this question, beyond what it happens to contain.

  1. Retrieve the top 50 results with hybrid search.
  2. Pass those 50 through the re-ranker.
  3. Send only the top 5 to your LLM.

This cuts context stuffing and improves what reaches the model. It’s one of the clearest differences between a demo and a production system that behaves consistently. The extra hop adds latency, which is where a Redis and semantic caching layer earns its keep.

Demo pipeline

Embed the query, take the top 5 vector matches and send them straight to the LLM. Exact codes get missed and near-miss chunks fill the context.

Production pipeline

Run vector and BM25 search, fuse the lists with RRF, re-rank the top 50 with a cross-encoder and send the best 5. Then score the answers every week.
Comic panel of a robot with a compass and a robot with a magnifying glass pouring result cards into a pink funnel, while a judge robot in a white wig holds up the best cards
Fuse both lists, then let a re-ranker pick.

Your production RAG checklist

A RAG system can impress in a demo and still be weak in production. Once real users, messy documents and ambiguous queries arrive, weak retrieval turns the LLM into expensive guesswork.

The fixes I rely on

  • Stop relying on vector-only search. Add a BM25 layer.
  • Use RRF. Fuse lexical and semantic results without fiddly score calibration.
  • Tune chunking on purpose. Chunks that are too small lose context, and chunks that are too large add noise. For technical docs, 512 to 1,024 tokens with 10 to 15% overlap usually works well for me.
  • Add a re-ranker. Refine the final candidates before anything reaches the LLM.
  • Evaluate with RAGAS. Track faithfulness, response relevancy, context precision and context recall instead of trusting intuition.

Key takeaways

  • Retrieval first. Most “hallucinations” in production RAG start with bad context.
  • Hybrid beats vector-only. BM25 catches the IDs and codes embeddings miss.
  • Fuse with RRF. Rank-based fusion avoids comparing incompatible scores.
  • Re-rank before you generate. Fifty candidates in, five good ones out.
  • Measure it. RAGAS scores beat gut feel.

Building a RAG demo is easy. Building reliable RAG takes a real understanding of retrieval, ranking and context design. If you’re still choosing your foundation, picking the right RAG stack walks through the vector database trade-offs. If your production RAG keeps missing, here’s how I help teams fix retrieval, ranking and evals, or get in touch with your war story.

Where does your own system still behave like a demo when it should be behaving like production?

▸ Made it to the end? Send it around.

▸ Share

▸ Comments