Skip to main content
    Engineering

    Retrieval-Augmented Generation in Production: What Actually Breaks

    The demo always works. Production is a different sport.

    By gAIcko Editorial TeamPublished Updated

    Written, fact-checked and maintained by the gAIcko Editorial Team. Corrections: admin@gaicko.com.

    What does it take to run RAG in production?

    Production retrieval-augmented generation needs disciplined chunking, a refresh pipeline that keeps the index current, an evaluation set scored on every change, citation of retrieved sources, and monitoring for retrieval misses. Model choice matters far less than retrieval quality.

    The short version

    Retrieval-augmented generation is easy to demo and hard to operate. In production, failures rarely come from the language model — they come from stale content, poor chunking, missing permission filters, and the absence of evaluation. Running RAG well is a content-operations problem with a machine learning component attached.

    What actually breaks

    1. Stale and duplicated source content

    The single most common cause of wrong answers. Three versions of a policy exist; retrieval finds the 2023 one. Fix at the source: a canonical document register, an owner per document, an expiry date, and an ingestion pipeline that removes superseded versions rather than adding to them.

    2. Chunking that destroys meaning

    Fixed 500-token splits cut tables in half and separate headings from the clauses they govern. Chunk on document structure — sections, headings, table boundaries — and attach the heading path as metadata so the model sees context, not just text.

    3. Retrieval that only does semantics

    Pure vector search misses exact identifiers: product codes, error numbers, clause references. Hybrid retrieval — BM25 keyword plus dense vectors, then a reranker over the top 50 — is the reliable default. The reranker usually buys more accuracy than upgrading the generation model.

    4. Permissions leaking

    If the index does not carry access control metadata and the query is not filtered by the caller's entitlements, RAG becomes a data-exfiltration surface. Filter before retrieval, not after generation. Never rely on prompt instructions for access control.

    5. No evaluation harness

    Teams ship without a test set, then argue about quality anecdotally. Build a set of 150–300 real questions with accepted answers and source citations. Score retrieval (was the right chunk in the top k?) separately from generation (was the answer faithful to the chunk?). Most quality problems are retrieval problems, and separating them tells you where to spend.

    6. Hallucination without abstention

    A system that never says "I don't know" will invent. Require citations, and configure the model to abstain when retrieval confidence is below a threshold. Abstention rate is a healthy metric, not a failure.

    7. Cost and latency drift

    Context windows grow as teams add "just one more" retrieved chunk. Cap context, measure tokens per answer, and cache aggressively — embeddings, reranker results and frequent answers.

    A reference architecture

    1. Ingestion: connectors → parsing (layout-aware for PDFs) → structural chunking → metadata enrichment (owner, date, ACL, heading path) → embedding → index.
    2. Query path: query rewrite → hybrid retrieval with ACL filter → rerank top 50 to top 5 → prompt assembly with citations → generation → citation validation.
    3. Feedback: thumbs and free-text on every answer, sampled human review, failures routed back to the content owner.
    4. Observability: per-query logging of retrieved IDs, scores, tokens, latency and cost.

    Metrics worth tracking

    • Recall@k — did retrieval surface a correct source in the top k?
    • Faithfulness — is every claim supported by a retrieved chunk?
    • Answer rate vs abstention rate — the honesty dial.
    • p95 latency and cost per answer.
    • Content freshness — percentage of indexed documents past their review date.

    A worked example

    An engineering support desk indexed 40,000 pages of manuals. Initial accuracy on a 200-question test set was 61%. Three changes moved it to 88%: structural chunking with heading paths (+11 points), hybrid retrieval plus a reranker (+12 points), and retiring 6,000 superseded pages (+4 points). The generation model was never changed. That distribution is typical.

    Operating model

    RAG needs an owner in the business, not only in engineering. Assign a content steward responsible for freshness and coverage, review the failure queue weekly, and treat repeated unanswerable questions as a documentation backlog rather than a model shortcoming.

    When not to use RAG

    If the answer requires calculation over structured data, query the database — text-to-SQL with validation beats retrieving spreadsheets. If the corpus is small and stable, putting it directly in the context window may be cheaper and more accurate. If the questions require multi-hop reasoning over many documents, expect to add an agentic retrieval loop and budget accordingly.

    Frequently asked questions

    What does it take to run RAG in production?

    Layout-aware ingestion, structural chunking, hybrid retrieval with reranking, permission filtering before retrieval, an evaluation set that scores retrieval and generation separately, citation enforcement, and a named content owner responsible for freshness.

    Why does RAG give wrong answers?

    Most often because of stale or duplicated source documents, chunking that separates context from content, or retrieval that misses exact identifiers. The generation model is rarely the primary cause.

    Is vector search enough for RAG?

    No. Pure semantic search misses exact codes and references. Hybrid keyword plus vector retrieval with a reranker over the top results is the reliable production default.

    How do you stop RAG from leaking confidential documents?

    Store access control metadata in the index and filter by the caller's entitlements before retrieval. Never rely on prompt instructions or post-generation filtering for access control.

    How do you evaluate a RAG system?

    Maintain 150–300 real questions with accepted answers and sources. Score recall@k for retrieval and faithfulness for generation separately, and track abstention rate, p95 latency and cost per answer.

    When should you not use RAG?

    When the answer requires calculation over structured data, when the corpus is small enough to fit in context, or when questions need deep multi-hop reasoning that a single retrieval pass cannot support.

    Sources and further reading

    • Pinecone / Elastic engineering documentation on hybrid retrieval and reranking

    Revision history

    • — Published in full with worked examples, FAQs and sources.