ЗамыселЪ Глеба Киренкова

Concept by Gleb Kirenkov

Business ·

RAG is less about an AI's memory than its right to cite a source

When an employee asks an AI about company policy, the model should not guess from general knowledge. RAG retrieves relevant passages first and then uses them to form an answer. Yet a useful system starts with deciding which documents are trustworthy and who may see them, not with choosing a vector database.

What happens between question and answer

RAG stands for retrieval-augmented generation. In a typical flow, the question triggers a search across permitted sources. The system selects passages, passes them to a model with the question and asks for an answer grounded in those passages. If nothing relevant is found, an honest “I could not find this” may be the right response.

This is not the same as retraining a model on every company document. A changed document can be updated in the retrieval layer without retraining the model, although the index still needs to be refreshed. Nor does a relevant passage guarantee a correct answer: the model may misread a table, miss an exception or combine incompatible versions of a rule.

Give knowledge an owner before indexing it

Imagine the question, “Can we promise delivery tomorrow?” One file describes the usual lead time, another an exception for a particular warehouse, and a third is obsolete. Loading all three without status information may produce a convincing but wrong answer. Before building the search system, assign each source an owner, an effective date, a version, a business scope and a retirement rule.

Knowledge also needs a correction path. When an employee finds a wrong answer, who changes the source, who approves the change and when is the index updated? Without those decisions, RAG merely makes a messy document collection faster to consult.

Enforce access before retrieval

A harmless question may retrieve a restricted contract or personal information. Telling the model not to reveal secrets after it has received the passage is too late: the restricted text has already entered the model's context. Retrieval must respect the questioner's permissions before any passage is passed on.

Review query logs, retention periods and the model provider's handling of transmitted data as well. The architecture depends on what the documents contain and how they may be processed, not just on a polished demo answer.

Test retrieval and answers separately

Build a small set of real employee questions: routine ones, disputed ones, questions involving an obsolete rule and questions with no answer in the collection. For each, record the acceptable source and expected behaviour. First check whether retrieval finds the right passage without exposing restricted material. Then check whether the answer follows the source, includes its qualifications and declines to guess when needed.

This separates a missing document from an access, retrieval or answer-generation problem. Prompt changes cannot repair a false knowledge base. Keeping the source version with the answer also helps explain it after a policy has changed.

When RAG is more than you need

For a few recurring questions about a short, stable policy, a well-maintained page with search and an editor may be enough. If the user needs to change an order or send an email, RAG alone cannot do the job: the action needs an integration, permission and confirmation. RAG provides evidence for an answer; it does not automate an entire workflow.

Begin with one testable question

Start with one collection and one kind of question whose accuracy and access boundaries can be tested. The result will show whether you need RAG, a better search interface or a fuller AI assistant. In an AI implementation, I would begin with precisely that decision: which information is allowed, who owns it and what the system must say when no reliable answer exists.

AI for a specific process

If you are considering AI implementation, we can start with the process, available data, automation boundaries and a way to assess the result.

Discuss AI implementation