Ask a general AI assistant about your company's leave policy, your product's warranty terms or last quarter's pricing changes, and it cannot know. It was trained on public data up to a cut-off date, not on your internal documents. Worse, it may produce a confident, plausible answer anyway. Retrieval augmented generation, usually shortened to RAG, is the most common way to fix this. Instead of relying on what the model memorised, the system first looks up relevant passages in your own content and then asks the model to answer using only those. This article explains how it works, where it fails, and how to judge whether it is working.
The problem RAG solves
Large language models (LLMs) have three limitations that matter for business use:
- No private knowledge. They have never seen your contracts, manuals or policies.
- Stale knowledge. Their training stops at a point in time; your documents change weekly.
- Hallucination. When they do not know, they can generate fluent text that is simply wrong.
You could paste every document into each question, but that quickly exceeds the amount of text a model can accept at once (its context window) and becomes slow and expensive. RAG sends only the few passages that matter.
How retrieval augmented generation works, step by step
Preparation (done ahead of time)
- Collect the source content: policies, help articles, manuals, product sheets, past tickets.
- Split each document into chunks of a few paragraphs, keeping headings and source information with each chunk.
- Embed each chunk. An embedding model converts text into a list of numbers (a vector) that captures its meaning, so passages about similar topics end up with similar vectors even if they use different words.
- Store the vectors in a searchable index, often a vector database or a vector extension of an ordinary database, such as pgvector for PostgreSQL.
At question time
- The user's question is embedded the same way.
- The system retrieves the chunks whose vectors are most similar, often combined with traditional keyword search to catch exact terms like product codes.
- Optionally, a reranker model re-scores those candidates for relevance.
- The best chunks are inserted into a prompt along with instructions such as: "Answer using only the sources below. If they do not contain the answer, say so. Cite the source for each statement."
- The LLM generates the answer, ideally with links back to the source documents.
A concrete example
An employee asks the internal assistant: "Can I carry forward unused leave into next year?" Retrieval finds two chunks: a section of the HR policy on leave carry-forward limits, and a recent HR circular that changed the deadline for applying. The model receives both, writes a short answer combining them, and links each point to its document. If the policy did not cover the question, a well-instructed system replies that it could not find the answer and suggests contacting HR, rather than guessing.
RAG vs fine-tuning
Fine-tuning means further training a model on your own examples. It is often confused with RAG but solves a different problem.
| RAG | Fine-tuning | |
|---|---|---|
| Best for | Answering from specific, changing facts and documents | Teaching a style, format or specialised task behaviour |
| Updating knowledge | Re-index the changed document | Retrain the model |
| Citing sources | Natural: you know which chunks were used | Difficult: knowledge is blended into the model |
| Access control | Can filter retrieval by user permissions | Anything trained in may be revealed to anyone |
For "answer questions from our documents", RAG is usually the right starting point. The two can be combined when a task needs both.
Where RAG goes wrong
Most poor RAG answers are retrieval problems, not model problems. If the right passage is not retrieved, the model cannot use it. Common causes:
- Bad chunking. A table split in half, or a rule separated from its exception three paragraphs later.
- Poor source content. Outdated, duplicated or contradictory documents produce outdated or contradictory answers.
- Vocabulary gaps. Users say "WFH" while the policy says "remote working arrangement". Hybrid keyword-plus-vector search and synonym lists help.
- Scanned PDFs and complex layouts that were never converted to clean text.
- Questions needing aggregation, such as "how many contracts expire this quarter?" RAG retrieves passages; it does not count across hundreds of documents. Questions like that belong in a database query.
- The model ignoring instructions and adding outside knowledge. Clear prompts, citations and testing reduce this but do not remove it entirely.
Security and permissions
If the HR assistant can retrieve salary documents, so can anyone asking the right question. Apply each user's access rights at retrieval time, filtering chunks by the same permissions as the original documents. Be aware that retrieved documents can themselves contain malicious instructions (a form of indirect prompt injection), so limit what actions the system can take based on retrieved text. If using an external AI provider, check how it handles and retains the data you send.
How to test a RAG system
Build an evaluation set of real questions with known correct answers and source documents. For each, check separately:
- Retrieval: was the right passage among those retrieved?
- Faithfulness: does the answer stick to what the sources say?
- Correctness and completeness: does it actually answer the question?
- Refusal: for questions the documents do not cover, does it say so?
Re-run the set whenever you change chunking, models or prompts. Separating retrieval from generation tells you which part to fix. Our AI and machine learning development team builds and evaluates RAG systems, and our database management team can set up vector search inside the databases you already run.
Key takeaways
- Retrieval augmented generation lets an AI answer from your own, current documents, with citations.
- Answer quality depends mostly on retrieval quality and the state of your source content.
- Use RAG for knowledge, fine-tuning for behaviour, and databases for counting and aggregation.
- Enforce document permissions at retrieval time and test with a fixed set of real questions.