What an Automation Is Allowed to Read Before You Embed a Single File
Classify every file before it becomes a vector. Secrets, other parties contracts, and personal data you have no basis to pull into a prompt stay out. Permission-aware stores are the control, not a prompt line.

You are about to point a retrieval job at a folder. The first task is not the embedding model. The first task is a list: which files this automation is allowed to read, which it must ignore, and who is allowed to ask. If that list does not exist, the index will be the folder, and the folder was never a permission system.
The doors around webhooks, keys, and prompts are already in the automation threat model. This page is one door: the knowledge store. Whether you should use retrieval at all, instead of a rule or an agent, is rules, retrieval, or an agent. Arrive here only after that choice is retrieval.
Classify the file before it becomes a vector
The Open Worldwide Application Security Project (OWASP) describes this kind of retrieval as a pretrained model plus external knowledge, reached through vectors and embeddings. Once a file is in that store, every query path is a reader. "We will tell the model to be careful" is not a reader list. The class of the file decides embed or refuse, before any batch job runs.
| Class | Examples | Embed? | Why |
|---|---|---|---|
| Operating fact you already publish | Hours, a public price, a policy page that is already on the site | Yes, if this file is the source of truth | The answer is meant to be repeated. Prefer the page itself when the page is enough. |
| Internal runbook | A staff procedure that is not customer-facing | Only inside a partitioned staff store | Customers must not retrieve it. A shared index will let them try. |
| Secrets | Application programming interface (API) keys, passwords, tokens, recovery codes | Never | A prompt is not a vault. Retrieval can place the secret in a completion. |
| Data from another party | Contracts or statements that include their confidential terms | Never | You do not have a basis to pull their data into a prompt because a bot was convenient. |
| Personal data | Customer records you have no basis to retrieve into a prompt | Never | Retrieval is a new disclosure path. Collecting the record was not a decision to place it in a prompt. |
| Unverified dumps | Exports, drafts, files from a provider you have not checked | Not until a person validates the source | Poisoning can be unintentional. Unverified providers are on the OWASP list. |
The Never rows are the policy. A query that would be easier with the file is not a reason to embed it. If a contract has a public clause and a confidential schedule, split the file or leave the whole file out. Partial redaction you did not check is still the whole file.
What retrieval actually exposes
Sourced from the Open Worldwide Application Security Project (OWASP) Top 10 for Large Language Model Applications. Retrieval-augmented generation (RAG) combines a pretrained model with external knowledge via vectors and embeddings. Weaknesses in how vectors are generated, stored, or retrieved can inject harmful content, manipulate outputs, or access sensitive information. Inadequate access controls can let the model retrieve personal or proprietary data. The entry is LLM08:2025 Vector and Embedding Weaknesses, listed from the OWASP Top 10 for LLM Applications.
Read that as a storage problem, not a wording problem. The dangerous moment is retrieval, when the wrong chunk is selected and placed in the prompt. Generation then does what generation does: it talks. OWASP LLM02:2025 Sensitive Information Disclosure is the name for the secret, the personal record, or the proprietary paragraph that comes back out in the answer.
OWASP LLM09:2025 Misinformation sits next to disclosure. A file can be allowed and still be wrong. A draft policy embedded beside the live policy gives the model two stories. It may pick the draft. Validation of the source is part of permission. "Staff may read this folder" is not the same as "this file is the current rule."
The multi-tenant failure OWASP names
OWASP describes this failure for LLM08. It is their scenario, not a client of ours. Embeddings from one group can be retrieved for a query from another group. Group can mean two customers on one product, two brands in one agency, or two teams that were never supposed to see files from the other group. The index does not know your org chart. It knows vectors and whatever filter you actually enforced.
A shared collection with a metadata field you sometimes set is not partitioning. Partitioning means a query from group A cannot return group B chunks, including when the filter is missing, empty, or overwritten by the caller. Test the refusal with a query you know should miss. A successful demo on friendly questions does not prove the wall.
- One store, many groups: treat it as shared until a permission check runs on every retrieval, not only on the upload form.
- Filter supplied by the model: the asker, or the model, must not be the authority for which partition is in scope. The application sets the partition from the authenticated identity.
- Logs: OWASP asks for immutable logs of retrieval. If you cannot see which partition and which chunks came back, you cannot audit a leak after the fact.
Inadequate access control is the phrase OWASP uses. It covers a store with no tenant key, a key that is optional, and a key the prompt can talk the tool into dropping. Build the check outside the model.
Poisoned or sloppy sources
Data poisoning can be intentional or unintentional. OWASP names insiders, prompts, and unverified providers. An insider can drop a file that states a refund rule the business does not use. A prompt pasted into a note can sit in the corpus and later ride along with a legitimate chunk. A provider export can include columns you did not mean to index: emails, tokens, internal comments.
Accept data only from trusted sources means a person signs the folder, not that the vendor logo on the export is famous. Trusted here is operational. You know who wrote it, you know it is current, and you know it contains no Never-class fields. Audit for poisoning means you look again when the folder changes, not only on the day you launched.
If the folder is supposed to follow a written procedure, the procedure comes first. Standard operating procedures before automation is the sibling for that order. An index of a messy drive automates the mess.
A classification pass you can finish in one sitting
Do this before any embed script runs. One folder. One owner. One page of notes. If the folder is too large to classify, it is too large to embed.
- Name the askers: customers, staff, one team, or a single internal tool. Each asker group gets its own allow list. "Everyone logged in" is not a group.
- Walk the files: mark each file Read, Hold, or Never using the table. Hold means a person must remove secrets or other-party data before a second review.
- Split or exclude: Never files leave the job. Hold files do not enter the batch "just this once."
- Bind the partition: the store key for this asker is set by the application. Write down what a cross-group query must return: nothing from the other group.
- Log retrieval: who asked, which partition, which source ids came back. Immutable means a later edit cannot wipe the evidence of a bad read.
Then run the production checks in evaluate the workflow before production. Include a case where the right answer is refusal: a secret-shaped string, a file from the other group, a question the allow list does not cover. A bot that always answers is failing those cases.
Hypothetical: the drive that became the knowledge base
Hypothetical. A team points the embed job at "Company Drive" because that is where the answers live. The drive also holds a vendor contract with the other party pricing, a spreadsheet of customer emails exported for a mail merge, and a text file where someone pasted an API token while debugging. The bot is demoed on an hours question. It answers well. Nobody asks it for the token.
The classification pass would have marked three Nevers and left a thin Read set: the published service page and one staff runbook in a staff partition. The demo would have looked smaller. Smaller is the goal. OWASP multi-tenant note still applies if a second brand is later dropped into the same index to "save a project." New group, new partition, new refusal test.
This hypothetical is a shape, not a result. It does not claim a breach happened, a fine, or a recovery time. It claims the upload path was doing the job of a policy.
After the allow list exists
Reading is not acting. A file can be allowed for retrieval and still be the wrong input to a tool that sends, pays, or deletes. Keep those tools off the knowledge job. When a person must approve an action, use a review step. The model may quote an allowed passage. It may not treat the passage as permission to move money or mail.
That split is human review on LLM jobs. Who may change the allow list later is change control, not a quiet edit in the embed script. The next guide is prompt and workflow change control.
If the allowed source is really a public page, publish it on the website and let the page be the answer. If you want a second pair of eyes on a folder before anything is embedded, send the file list. The list, not a tour of the model.
Frequently asked questions
What should never be embedded for retrieval?
Secrets (keys, passwords, tokens), contracts that contain data from another party, and personal data you have no basis to retrieve into a prompt. Retrieval-augmented generation (RAG) puts those passages in front of a model. Open Worldwide Application Security Project (OWASP) LLM02:2025 is Sensitive Information Disclosure. If the chunk should not be said aloud, it should not be searchable.
What is a permission-aware vector store?
Open Worldwide Application Security Project (OWASP) LLM08:2025 Vector and Embedding Weaknesses names permission-aware vector stores and partitioning as a mitigation for large language model (LLM) retrieval. The store must refuse a query that the asker is not allowed to run, and it must keep embeddings from one group from answering a query from another group. A single shared index with a polite system prompt is not that control.
What is the multi-tenant leak OWASP describes?
Open Worldwide Application Security Project (OWASP) LLM08:2025 says embeddings from one group can be retrieved for a query from another group when access control on a large language model (LLM) store is inadequate. That is their scenario for vector and embedding weaknesses, not a story from our client work. If two teams, two customers, or two brands share one index, assume a query can cross the wall until you have partitioned and tested the refusal.
Can I embed a file if the model is told not to reveal it?
No. Open Worldwide Application Security Project (OWASP) LLM01:2025 Prompt Injection is when user prompts alter the large language model (LLM) behavior. A line that says "do not reveal secrets" is not a control once the secret is in the store and can be retrieved into the prompt. Keep the file out. A tone rule does not remove a vector.
Who decides a file is allowed?
A named owner, before the embed job runs, using the classes on this page. The person who can upload to the drive is not automatically the person who may approve retrieval. Tie the decision to the standard operating procedure (SOP) for that folder, then test the workflow before production.