CreativeCape

RAG That Survives Production

Chunking, retrieval quality, evals and the boring infrastructure that separates a RAG demo from a system that holds up under real questions.

August 24, 2026·5 min read

A retrieval-augmented generation demo takes an afternoon. Load documents, chunk them, embed, query, feed the results to a model. It answers questions and everyone is impressed.

Then real users arrive and ask questions your test set never contained, and the gap between the demo and a system people trust becomes obvious.

Why demos break under real questions

Demo questions are well-formed and answerable from a single passage. Real questions are not:

Multi-hop questions. "How does the refund policy differ for annual plans?" needs the refund policy and the plan definitions. Top-k retrieval on one embedding returns passages about refunds, or about plans, rarely both.

Questions the corpus cannot answer. The most damaging failure. Nothing relevant exists, retrieval returns the five least-irrelevant chunks anyway, and the model — handed context and asked to answer — produces something confident and wrong.

Vocabulary mismatch. Users say "cancel my subscription"; the documentation says "terminate your agreement". Semantically close, but if the exact product term matters, pure vector search can miss it.

Stale content. The corpus has both the 2024 and 2026 pricing page. Retrieval has no opinion about which is current.

Chunking that respects meaning

Chunking is where most quality is won or lost, and the default — split every 512 tokens — is close to the worst option, because it cuts mid-sentence and separates headings from the text beneath them.

Better rules:

Split on structure, not length. Markdown headings, sections, list boundaries. A chunk should be a coherent unit someone could read standalone.

Keep the heading path in the chunk. Prefix each chunk with its breadcrumb:


Billing > Refunds > Annual plans

Annual plans may be refunded pro-rata within 30 days…

This is cheap and disproportionately effective — it gives both the embedding and the model the context that the document structure implied.

Overlap modestly. 10–15% stops answers being cut in half. More wastes context and returns near-duplicates.

Keep tables and code intact. Splitting a table across chunks destroys it. Oversized chunks are better than fragmented ones here.

Store metadata: source, section, last-updated, product area. You need it for filtering and for citations.

Retrieval quality beats model choice

Teams reach for a bigger model when answers are poor. Usually the model was fine and was handed the wrong context.

Measure retrieval separately from generation. Build a set of question → correct-chunk pairs and measure recall@k: how often is the right chunk in the top k? If recall@5 is 60%, no model fixes the missing 40%. You are asking it to answer from material it was never shown.

Diagnose in that order: is the chunk retrievable at all → is it in the top k → did the model use it correctly. Skipping to the last step is why so much RAG tuning is wasted effort.

Hybrid search and reranking

Two upgrades with the best return on effort:

Hybrid search. Combine vector similarity with keyword search (BM25). Vectors handle paraphrase; keywords handle exact terms — error codes, product names, part numbers. Vector search is notably weak on rare exact tokens, which is exactly what technical users type.

Reranking. Retrieve 20–50 candidates cheaply, then rerank with a cross-encoder that scores each against the query properly, and pass the top 5 to the model. Precision improves substantially for one extra call.

Together: hybrid retrieve wide, rerank precisely, generate from a small high-quality set.

Evals you can run in CI

Without evaluation you are tuning by vibes, and every change is a coin flip.

A workable minimum:

  1. A question set — 50–200 real user questions (from logs once you have them), with expected source chunks and acceptable answers.

  2. Retrieval metrics — recall@k and MRR. Fast, deterministic, cheap. Run on every change.

  3. Answer grading — an LLM-as-judge scoring faithfulness (is it supported by the context?) and relevance. Noisier, but catches regressions.

  4. An adversarial set — questions the corpus genuinely cannot answer. The correct behaviour is "I don't know." Score refusals as successes. This is the set that prevents confident nonsense.

Run 1–3 in CI. Nobody re-tunes chunking by hand at 2am to find they made recall worse.

Caching and cost

Cache embeddings keyed by content hash. Re-embedding an unchanged corpus on every deploy is pure waste.

Cache full answers for repeated questions — support corpora have a long head of identical queries.

Use prompt caching where your provider supports it. A long system prompt plus retrieved context is mostly stable within a session.

Retrieve fewer, better chunks. Ten mediocre chunks cost more than five good ones and produce worse answers, because relevant material gets diluted.

Failure modes and guardrails

Always cite sources. Return the documents used, linked. It lets users verify, and it makes hallucination visible instead of invisible.

Allow "I don't know". Instruct the model explicitly to refuse when the context is insufficient, and set a retrieval score threshold below which you do not call the model at all.

Handle conflicting sources. When two chunks disagree, prefer the more recent and say so.

Log everything — query, retrieved chunks with scores, final answer. Without this you cannot diagnose a complaint, and "the AI said something wrong" is unactionable.

Watch for corpus drift. Content changes; embeddings go stale. Re-index on publish, not on a quarterly cron.

The pattern across all of these: a production RAG system is mostly a retrieval and data-quality system. The model is the part you change last.


→ AI Development

Tagged with
#rag#llm#embeddings#vector database#ai engineering#evals

Found this useful? Share it.

Keep Reading

Related articles

Booking Q2 2026 Projects

Ready to Build Something Great?

From idea to launch — let our senior engineers build, ship and scale your next product. No commitment, just a conversation.

Senior Engineers
On-Time Delivery
Enterprise-Grade
Free Consultation

Free 30-min discovery call

Talk to a senior engineer — not a salesperson.

We'll review your goals, suggest the leanest path forward, and send a clear proposal within 24 hours.

24h

Response Time

100+

Projects Delivered

No commitment · No automated bots · Fully transparent