Understanding Retrieval-Augmented Generation (RAG)
A language model does not automatically know the contents of your private documents or the latest information in your application. Retrieval-Augmented Generation (RAG) addresses this by finding relevant source material and including it in a model's prompt.
The basic flow
- Split source documents into smaller chunks, keeping useful context such as headings and document IDs.
- Convert each chunk into an embedding, a numeric representation that captures semantic relationships.
- Store embeddings and their text in a searchable index.
- Embed a user's question and retrieve the most relevant chunks.
- Ask the language model to answer using those chunks, and return source references where possible.
Retrieval can use vector similarity, keyword search, or a hybrid of both. The best choice depends on the data: exact identifiers often benefit from keyword search, while paraphrased questions can benefit from semantic retrieval.
Keep the prompt grounded
Tell the model which retrieved text is evidence and what to do when that evidence is insufficient. For example:
The final rule matters because retrieved documents are untrusted input. A document might contain text that attempts to override the system prompt. Do not let retrieved content change permissions, invoke tools, or bypass application policy.
Improve retrieval before changing models
If answers are poor, inspect the retrieved chunks first. Common issues include chunks that are too large or too small, missing document metadata, stale indexes, and a retrieval query that does not reflect the user's intent. Measure retrieval quality separately from answer quality.
An evaluation set should contain representative questions, expected source documents, and examples where the correct response is to abstain. Track whether the right evidence was retrieved and whether the final response is supported by it.
Production considerations
- Apply access control before returning chunks; users must not retrieve documents they cannot read.
- Rebuild or update the index when source documents change.
- Limit how much retrieved text enters a prompt to control cost and latency.
- Record citations so answers can be checked against their sources.
- Avoid sending sensitive documents to a model provider unless the data handling terms permit it.
Conclusion
RAG is a pipeline, not a single model setting. Reliable results depend on good source preparation, appropriate retrieval, grounded prompts, access control, and evaluation of both evidence and answers.
Other posts that might interest you...
آشنایی با تولید تقویتشده با بازیابی (RAG)
یاد بگیرید بازیابی، embedding و پرامپتهای مستند چگونه در یک سیستم RAG با هم کار میکنند و پاسخها را چگونه ارزیابی کنید.
شروع یادگیری ماشین با Scikit-learn
با مفاهیم اولیه یادگیری ماشین و کتابخانه Scikit-learn آشنا شوید و اولین مدل خود را آموزش دهید.
Getting Started with Machine Learning using Scikit-learn
A beginner-friendly introduction to machine learning with Python and Scikit-learn, covering datasets, training models, and evaluation.