All Tabs

RAG Starts With the Documents

airagretrieval

I used to start RAG work by tuning the prompt. Now I look at the documents, the chunking, and the retrieval results first, because a clever prompt cannot recover facts we never fetched.

The first question is embarrassingly basic: do the source documents actually contain the answer? If the information is missing, stale, or buried in a format the loader ignores, the model is being blamed for a data problem.

Chunking changes the result more than I expected. Chunks that are too small lose the surrounding explanation, while chunks that are too large bring back a whole page and make the useful sentence harder to find.

I also want to inspect what retrieval returned before I judge the answer. A low-quality context window can look like a prompt failure when the real issue is the query, the metadata filter, the embedding model, or a document that should never have been indexed.

That gives me a better debugging order: inspect the source, inspect the chunks, inspect the retrieved context, and only then change the prompt. It is slower than tweaking instructions, but it gives each change a reason.

RAG is not a magic memory layer. It is a document pipeline with a language model at the end, and the boring parts decide how much truth reaches the answer.

Thanks for reading.Read more tabs →