Rules, Retrieval, or an Agent: Pick the Smallest System That Can Fail Safely
A closed answer belongs in a lookup. Documents a rule cannot list belong in retrieval, and you accept the vector and misinformation risks. An agent is only for tool choice, and then excessive agency is the limit.

Someone on the team asks the same question every Monday. Opening hours. The refund rule. Which status means the job is done. The tempting build is a chat box on the shared folder. Most of those questions do not need a model that picks its own steps. They need the smallest system that can be wrong in a way you can see.
A pipeline that always runs the same steps, versus an agent that chooses steps, is a different cut. That comparison lives in deterministic workflows versus large language model (LLM) agents. This page is the knowledge question. Does the answer live in a closed table, in documents a rule cannot list, or in a task that must choose a tool?
Three systems, three ways to be wrong
A lookup or a rule returns the stored answer when the input matches. If the hours are 9 to 5, the reply is 9 to 5. It does not add a holiday you forgot to store. It does not soften the sentence. Two runs on the same input match. When they do not, you have a bug in the table, and you can diff it.
Retrieval is for an answer that is real and still not a row. A policy, a runbook, a set of notes the business will not flatten into columns. You hand the model passages from a search over vectors. You accept that the wrong passage can come back, or a passage that should never have been searchable.
An agent is for a task that must choose a tool. Read this, then maybe write that, then maybe notify someone. The failure is no longer a bad sentence. The failure is an action. If the task does not need that choice, the agent is extra machinery with a larger blast radius.
flowchart TD
Q["Knowledge question"] --> C{"Closed answer, identical every time?"}
C -->|Yes| R["Lookup or rule"]
C -->|No| D{"Answer lives in documents a rule cannot list?"}
D -->|Yes| V["Retrieval: vector and misinformation risks"]
D -->|No| T{"Task must choose a tool?"}
T -->|Yes| A["Agent: least privilege, human review"]
T -->|No| R
Read the diagram as a gate. Stop at the first box that can fail safely. A clever system you do not need is still the one that will wake you up. Fluency in a demo is not a reason to skip a box.
Use a rule when the answer is closed
Closed means you can list the legal answers before the question shows up. Identical means the business treats a paraphrase as a defect. Price, status, a code, a yes or no, a window of time. If two customers can compare replies, the replies have to match.
- Closed set: a person could keep the answer in a spreadsheet column. The model is a worse spreadsheet. It will not tell you which row it ignored.
- Identical output: the same input must produce the same answer. A rule does that. A generation step is the wrong tool when a paraphrase counts as a defect.
- No document hunt: the answer does not depend on a paragraph you would have to search. Retrieval would add a failure mode and no new fact.
The failure mode of a rule is dull, and that is the feature. The row goes stale because nobody updated it. You can see the stale row. You can name the person who owns the table. You do not need a log of which chunk scored highest. If the answer changes often, the work is still a row update, not a new architecture.
People skip the rule because the question is phrased in sentences. Phrasing is not the test. The test is whether the answer set is closed. "Are you open Saturday?" is a sentence. The answer is a cell. Put the cell in the reply and stop.
Use retrieval when the rule cannot list the documents
Some answers are a few sentences inside a file, and the file moves. "What does the refund note say about a custom order?" might be half a page. You will not encode every sentence as a branch. You will miss one. The miss stays quiet until a customer quotes you back.
That is retrieval. Sourced from the Open Worldwide Application Security Project (OWASP) Top 10 for Large Language Model Applications (2025): retrieval-augmented generation (RAG) combines a pretrained model with external knowledge via vectors and embeddings. The model does not contain your folder. A search returns passages, and the model speaks from those passages. See the OWASP Top 10 for LLM Applications and the LLM08:2025 Vector and Embedding Weaknesses entry.
You are accepting two named risks, on purpose, because the documents will not fit a table. Write them down next to the design so a later editor cannot pretend the bot is "just search."
- LLM08:2025 Vector and Embedding Weaknesses: weaknesses in how vectors are generated, stored, or retrieved can inject harmful content, manipulate outputs, or access sensitive information. Inadequate access controls can let the model retrieve personal or proprietary data. OWASP also notes that embeddings from one group can be retrieved for a query from another group. That multi-tenant case is a file-permission problem, handled in the next guide, not a prompt tweak.
- LLM09:2025 Misinformation: OWASP lists this separately. A retrieved passage can be stale, wrong, or poisoned, and the model can still deliver it in a confident sentence. If the business needs the sentence to be identical every time, you are in the wrong box. Go back to the rule.
Data poisoning can be intentional or unintentional. OWASP names insiders, prompts, and unverified providers. A folder you did not read is a provider you did not check. The last person who dropped a file in the drive is now part of the answer path.
Classify files before the first one is embedded. That gate is what an automation is allowed to read. Building the index first and writing the policy later is how secrets and contracts that belong to other parties end up in a prompt.
Google uses the same words for a different system
Search results for the acronym will also show Google. In Optimizing your website for generative AI features, Google defines retrieval-augmented generation as grounding AI Overviews in the Search index. That is Google Search. Your knowledge bot retrieves from a store you filled. Different index, different readers, different failure. Advice about being cited in an AI Overview does not decide who may query your vectors. Your bot answers are not a ranking project.
Keep the names apart in the design doc. Call the internal one "our retrieval store" and leave Google Search RAG for the search team. Mixing them makes people import the wrong threat. A public overview grounded in the open web is not a staff bot grounded in payroll notes.
An agent only when the task must choose a tool
If the question is "what does the policy say?", retrieval plus a draft is enough. If the question is "look up the order, check the policy, and decide whether to email the customer," you asked for tool choice. That is an agent, or a workflow that hides the choice inside one node and hopes nobody notices.
OWASP LLM06:2025 is Excessive Agency. The constraint is least privilege and human review. The agent holds only the tools this job needs. A knowledge job needs read. It does not need send, pay, or delete. Those are actions. A person approves them, or a rule that does not ask the model approves them.
The pattern for a person on the irreversible step is human review on LLM jobs. If you cannot name the reviewer, you do not have a gate. You have a hope that the model will be polite.
Prompt injection sits beside tool choice. OWASP LLM01:2025 Prompt Injection means user prompts alter the model behavior. A document in the retrieval store is text you did not write, and the model will read it. If the agent can call tools, a poisoned chunk is an instruction with hands. A lookup cannot be talked into issuing a refund. That is another reason to stay in the smaller box.
OWASP LLM02:2025 is Sensitive Information Disclosure. Retrieval can hand the model a secret, a contract clause, or a personal record, and the model can repeat it. Least privilege on tools does not fix a store that should not have held the file. Fix the store first.
Run this on one real question
Pick a question the team actually asked this month. Do not invent a platform. Walk the gate in order and write the answer in one paragraph someone else can disagree with.
- Write the question: one sentence a colleague would send in chat. If you need three sentences, you have three jobs. Split them before you pick a system.
- List the legal answers: if the list is short and stable, stop. Build the lookup. Put the owner of the row in the same note.
- Name the files: if the answers are paragraphs, list the files. If you cannot name who may retrieve each file, stop. Do not embed.
- List the tools: if the model must choose a tool, write the list. If send, pay, or delete is on it, the design is already too big. Split the knowledge step from the action.
- Test before production: a paragraph in a doc is not a test. The golden set and the failure checks live in the evaluation guide linked below.
Before the workflow takes live traffic, run evaluate the AI workflow before production. The doors (webhooks, keys, prompts, connectors) belong in the automation threat model. This page only picks the system. It does not replace either check.
Hypothetical: the hours bot that grew tools
Hypothetical. A studio wants a box on the site that answers "when are you open?" The hours are one row. A rule would answer it. Someone embeds the shared drive because the model might need context. The drive holds a draft refund note and a vendor contract. A visitor asks about a late delivery. The bot quotes the draft, which was never the policy.
A week later the same bot is given a mail tool so it can "close the loop." A visitor pastes instructions into the chat. If the tool can send, LLM06 is no longer a line in a PDF. It is an outbox. The repair is to shrink. Hours stay a rule. The one policy folder stays retrieval, with the LLM08 controls, and only after the file list is classified. Mail stays with a person.
Nothing in that story is a measured outcome. It is the shape of the mistake: the answer was closed, the store was not, and a tool arrived because the demo felt thin.
What you still have to operate
A rule rots when the row is old. Retrieval rots when the file is old, poisoned, or visible to the wrong query. An agent rots when someone adds a tool and does not add a review. None of these heal because the first week looked smooth. The owner should be able to point at the box you picked and say why the larger box was refused.
If the question is really a page on the site, fix the page as a website so the answer is public and identical, and the bot is unnecessary. If you want the gate run against a real folder and a real tool list, start a conversation with the question, the file names, and the tools. Leave the architecture slogan at the door.
Frequently asked questions
When is a lookup or a rule enough?
When the legal answers are a closed set and the same input must produce the same answer every time. Hours, a status code, a yes or no, a price you already store in a row. A large language model (LLM) will paraphrase. A rule will not. If a colleague could maintain the answer in a spreadsheet, start there.
What is retrieval in this decision?
Retrieval-augmented generation (RAG), in the Open Worldwide Application Security Project (OWASP) Top 10 for Large Language Model Applications, combines a pretrained model with external knowledge via vectors and embeddings. Use it when the answer lives in documents a rule cannot enumerate. You then accept OWASP LLM08:2025 Vector and Embedding Weaknesses and LLM09:2025 Misinformation. The model does not know your folder. It speaks passages a search returned.
Is Google Search retrieval the same as our internal knowledge bot?
No. In Optimizing your website for generative AI features, Google defines retrieval-augmented generation (RAG) as grounding AI Overviews in the Search index. That store is Google Search. An internal bot retrieves from files you embedded. Same phrase, different system. Do not copy Search advice onto who may query your vector store.
When should the knowledge task be an agent?
Only when the task needs a choice of tools. If the job ends at an answer, retrieval or a rule is the smaller system. An agent that can pick a tool falls under Open Worldwide Application Security Project (OWASP) LLM06:2025 Excessive Agency, in the Top 10 for large language model (LLM) applications. The constraint is least privilege and human review. Sending, paying, and deleting are actions, not knowledge.
What goes wrong if I embed the folder and add tools later?
You skipped the gate. Open Worldwide Application Security Project (OWASP) LLM08:2025 says weaknesses in how vectors are generated, stored, or retrieved can inject harmful content, manipulate outputs, or access sensitive information. OWASP LLM01:2025 Prompt Injection is when user prompts alter the large language model (LLM) behavior. A retrieved passage is text the model will read. If a tool can send or pay, a bad passage is no longer a wrong sentence. Classify files first, then decide whether any tool is allowed at all.