Back to blogEngineering

RAGThatActuallyWorks:LessonsFromProductionRetrievalSystems

Marcus Lindqvist· Staff ML Engineer· October 22, 2025· 9 min read

Chunking is a search problem, not a formatting problem

Most RAG systems that fail in production were chunked to make the code simple, not to make retrieval accurate. Splitting documents every 500 tokens regardless of structure means a single answer often gets sliced across two chunks, and the retriever has to get lucky twice. We chunk along semantic boundaries instead: sections, subsections, and list items, with a modest overlap of 10 to 15% so that a sentence referencing prior context is not orphaned.

For structured documents like contracts or technical manuals, we keep a parent-child relationship: a small chunk gets embedded and searched, but retrieval returns the full parent section for context. This alone has fixed more "the answer is almost right but missing a caveat" bugs than any change to the model or the prompt.

Hybrid search beats pure vector search

Pure embedding search is bad at exact matches: product SKUs, error codes, legal clause numbers, and acronyms often sit far apart in embedding space from the query that references them. We run a hybrid pipeline combining a keyword-based sparse retriever (BM25 in most stacks) with dense vector search, and merge the results with a weighted reciprocal rank fusion rather than picking one or the other.

In one internal benchmark against a support knowledge base, adding sparse retrieval on top of dense-only search lifted recall at the top five results from roughly 68% to 84% specifically on queries containing exact identifiers. Dense search alone remains stronger on paraphrased, conceptual questions, which is exactly why the two need to run together rather than as a choice.

Building an eval harness before you ship

A RAG system without a labeled evaluation set is a demo, not a product. We build a set of 50 to 150 question-answer pairs grounded in the source documents, each labeled with the specific passage that should have been retrieved. This lets us separate retrieval failures from generation failures: if the right passage was retrieved but the answer was wrong, that is a prompting problem; if the passage was never retrieved, that is a search problem, and no amount of prompt tuning will fix it.

Hallucination control in high-stakes domains

In domains like healthcare intake or financial compliance, an ungrounded answer is worse than no answer. We require the model to cite the specific chunk it used for each claim, and we run a lightweight verification pass that checks whether the generated answer is actually entailed by the cited text, rejecting or flagging answers that are not. This adds latency, typically 300 to 600 milliseconds, which is worth it anywhere a wrong answer carries real cost.

We also explicitly instruct the model to say it does not know when retrieval returns low-confidence or no matches, and we test that behavior directly in the eval set with questions that have no answer in the corpus. A system that never says "I don't know" almost always has weaker guardrails than it looks like it does.

What we monitor once it is live

Retrieval quality drifts as source documents change, so production RAG needs ongoing monitoring, not a one-time launch eval. We track retrieval hit rate against a rotating sample of real queries, log every citation the model produces for periodic human spot checks, and alert when average confidence scores drop, which is often the earliest sign that a document set has changed shape faster than the index has been refreshed.

RAGRetrievalLLM EngineeringSearch

Wanthelpshippingsomethinglikethis?