# Agentic Competitive Research With Kill Switches

- Source: https://simeoncreatives.com/blog/agentic-competitive-research-workflow
- Hub: AI & Business Automation
- Author: Simeon Matheka, Founder & Creative Director
- Published: 2026-10-01
- Updated: 2026-10-01
- Reading time: 17 min

An agent that browses the web to research a competitor is the riskiest version of this job. If you still want one, fence it: read-only tools, a domain allowlist, hard budgets, a source requirement, a threat card, and a kill switch you can hit at lunch.

The pitch is tempting. Give an agent a competitor name, let it search, open pages, follow links, and come back with a brief. It feels like hiring an analyst who never sleeps. It also means software with a browser, a credit card’s worth of model calls, and your instructions is walking through pages written by strangers, some of whom would like to redirect it.

If you can name the pages you care about, do not build this. Use the [competitor change monitor with a review queue](https://simeoncreatives.com/blog/competitor-change-monitor-with-human-review), which has one model step and no tools. Come back here only for open questions where you do not know which pages to read.

## When an agent is the right shape

| Question | Better shape | Why |
| --- | --- | --- |
| Did this specific pricing page change? | Scheduled fetch + diff + one model step | The page is known. No tool choice is needed. |
| What does this one vendor’s docs say about limits? | Fetch the page, source-required extraction | A bounded read with a quote check. |
| Who else serves this niche, and how do they position? | Agent with a fenced toolset, then human review | The path is open. Search and page choice vary by run. |
| Send our team a Slack summary every Monday | Deterministic job, human-approved text | Delivery is a closed action. Keep it out of the agent. |

Rule of thumb from [rules, retrieval, or an agent](https://simeoncreatives.com/blog/rules-retrieval-or-agent): an agent is warranted only when the task must choose among tools. If the job ends at an answer, the smaller system wins.

## Design the fence before the prompt

Most agent advice starts with prompts. For a research agent, the prompt is the least important control, because a prompt is a request and the model can be talked out of it. The fence is made of things the model cannot negotiate: the tools it is given, the domains it can reach, and the budgets that stop it.

```mermaid
flowchart LR
    Q["Research question"] --> G["Gate: budgets + domain allowlist"]
    G --> A["Agent (read-only tools)"]
    A --> N["Notes store"]
    N --> S{"Source check: quote in fetched text?"}
    S -->|No| D["Drop and log"]
    S -->|Yes| R["Human review"]
    R -->|Approve| O["Brief"]
    K["Kill switch"] -.-> G
    K -.-> A
```

> Framework: the agent can read public pages and write notes to one store. It cannot send, post, pay, log in, or edit anything outside that store.

## The toolset: three tools, no more

| Tool | Allowed? | Limit |
| --- | --- | --- |
| Search public web | Yes | Fixed query budget per run. Results logged. |
| Fetch a page (GET only) | Yes | Domain allowlist or blocklist, size cap, no cookies, no custom auth headers. |
| Write a note to the notes store | Yes | One store. Append only. Schema-validated records. |
| Send email or post to chat | No | Delivery is a separate approved step. |
| Run code or shell | No | A research agent does not need it. |
| Access the customer relationship management (CRM) system, billing, or file storage | No | No credentials in the agent’s environment at all. |

OWASP traces excessive agency to three roots: excessive functionality, excessive permissions, and excessive autonomy. This table removes the first two by construction. The budgets below handle the third. Sourced: [OWASP LLM06:2025 Excessive Agency](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/) (checked 1 October 2026).

## Budgets that stop a run without you

A kill switch you have to be awake to press is not enough. Put hard caps in the runner, outside the model’s control, so a looping or manipulated agent runs out of road.

#### Budget config

```json
{
  "max_steps": 25,
  "max_searches": 8,
  "max_pages_fetched": 20,
  "max_page_bytes": 500000,
  "max_wall_clock_seconds": 600,
  "max_model_spend_usd": 2.00,
  "allowed_methods": ["GET"],
  "allowed_domains_mode": "blocklist",
  "blocked_domains": ["localhost", "intranet.example", "mail.example"],
  "kill_switch_flag": "agent_research_enabled"
}
```

#### Runner guard (JS)

```js
// The runner, not the model, enforces limits.
function guard(state, budget, flags) {
  if (!flags.agent_research_enabled) return stop('kill_switch');
  if (state.steps >= budget.max_steps) return stop('max_steps');
  if (state.pages >= budget.max_pages_fetched) return stop('max_pages');
  if (state.spendUsd >= budget.max_model_spend_usd) return stop('max_spend');
  if (Date.now() - state.start > budget.max_wall_clock_seconds * 1000)
    return stop('timeout');
  return 'continue';
}
// Every stop reason is logged with the run id. A cap that never fires
// is a cap nobody tested.
```

*Hypothetical run budget. Numbers are starting points for a small job. Tune them from your own logs, not from this page.*

Block internal addresses and your own services by default. A fetch tool that can reach localhost or an internal admin page can be pointed at them by a page the agent reads. Block that class of target in the runner, where a persuasive page cannot argue with it.

## The threat card: what can go wrong, and the control

This is a framework table, not a list of incidents. Each row is a door and the control that closes it. Prompt injection, in the OWASP definition, is when inputs alter a model’s behavior or output in unintended ways, including inputs a human cannot see. Sourced: [OWASP LLM01:2025 Prompt Injection](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) (checked 1 October 2026).

| Door | What it looks like | Control |
| --- | --- | --- |
| Injected instructions in a page | Hidden or visible text tells the agent to ignore rules, reveal its prompt, or visit a link. | Read-only tools. Page text labeled as data. Nothing the agent can send or write outside the notes store. |
| Scope creep | The agent follows links to login pages, forms, or private areas. | GET only, no credentials, blocklist for internal and sensitive hosts, step cap. |
| Fabricated findings | Confident claims with no support on any page. | Source-required records. Quote must appear in the fetched text. Drop what fails. |
| Runaway cost or loop | The agent searches in circles or fetches huge pages. | Step, page, byte, spend, and time caps in the runner. |
| Leaky context | Your internal notes or pack end up inside a search query or a fetched URL. | Keep the context pack minimal for this job. No secrets, no customer data in the prompt. Log every outbound query. |
| Quiet drift | A prompt or model change shifts behavior and nobody notices. | Version the prompt and budgets, rerun a golden question set after each change. |

The Leaky context row is design reasoning, not a reported incident. The safe move is the same either way: the less sensitive material in the prompt, the less there is to leak. The wider small-business version of this exercise is in the [AI automation security threat model](https://simeoncreatives.com/blog/ai-automation-security-small-business-threat-model).

## Kill switches you can actually hit

1. **Global flag: **One setting that blocks every new run. Checked at the start of the run and at every step.
2. **Run cancel: **A way to stop a run in flight, by run id.
3. **Hard caps: **Steps, pages, bytes, spend, wall clock. These stop a run when nobody is looking.
4. **Credential cut: **The agent has no real credentials to revoke. If one existed, revoking it would be a fourth switch, and the real fix is not having it.
5. **Alert on stop: **Any cap or switch stop writes a log line and pings the owner. A silent stop looks like a quiet week.

Test the switches before launch. Flip the flag in the middle of a run. Lower the page cap to three and confirm the run stops at three. The approval pattern for what happens next is in [human-in-the-loop review for n8n and LLM jobs](https://simeoncreatives.com/blog/human-in-the-loop-n8n-llm-jobs).

## From notes to a brief: the review gate

The agent never writes the final brief. It writes claim records to the notes store, each with a URL, a verbatim quote, and a retrieval date. A script keeps only the records whose quote appears in the fetched page. A person reads the survivors and writes or approves the brief. The method is in [source-required LLM research](https://simeoncreatives.com/blog/source-required-llm-research-outputs).

## Filled hypothetical: mapping a niche

Hypothetical, not a client. A studio asks the agent: “Which firms offer fixed-price brand packages to dental practices, and what do their package pages say?” The run is capped at 25 steps and 20 pages. It returns 31 claim records across 9 sites. The quote check drops 6. A reviewer rejects 4 more as marketing fluff with no concrete detail, and parks 3 that cite directory listings. What is left is 18 sourced facts about package structure. The reviewer writes a half-page summary and links every line.

The agent did the legwork a person would find tedious. It did not decide anything. If a page had tried to redirect it, the worst outcome was a wasted fetch, because it had nothing to send and nothing to break.

## What this does not prove

None of this makes an agent safe in general. It makes this agent’s worst case small. A fenced agent can still miss important competitors, over-weight whatever ranks well in search, and misread a page. It also cannot see anything behind a login or inside a sales call. Treat the output as a lead list with citations, not as a market analysis.

Print the [agentic competitive research threat card](https://simeoncreatives.com/resources/agentic-ci-threat-card) and do not start the first run until every row has a control and a tested kill switch. Before any agent touches production, run it through [evaluating an AI workflow before production](https://simeoncreatives.com/blog/evaluate-ai-workflow-before-production). If you want a second look at your fence, [start a conversation](https://simeoncreatives.com/contact).

## FAQs

### Do I need an agent for competitive research?

Usually not. If you know which pages matter, a scheduled fetch, a text diff, and a model that describes the change does the job with far less risk. An agent earns its place only when the question is open (“who else sells to this niche, and how do they position?”) and the path to the answer is not known in advance.

### What is excessive agency?

The Open Worldwide Application Security Project (OWASP) uses the term for damaging actions taken in response to unexpected, ambiguous, or manipulated outputs from a large language model (LLM). It usually traces to too many tools, too many permissions, or too much autonomy. A research agent needs read access to public pages and a place to write notes. It needs nothing else.

### What is prompt injection, and why does it matter here?

Prompt injection is when text the model reads changes what the model does. A research agent reads other people’s pages all day, so every page is a possible source of instructions. You cannot filter that perfectly, so you limit what a manipulated agent could do: no write tools, no email, no credentials.

### What belongs in a kill switch?

A single flag that blocks new runs, a way to cancel the run in flight, and hard caps on steps, pages, tokens, and wall-clock time that stop a run without anyone watching. If stopping the agent takes more than a minute, add another switch.

### Who reviews the output?

A named person, before anything is used in a deck, a brief, or a message. Pair the review with source-required output so the reviewer can click from each claim to the supporting quote.

### Can the agent write to our customer relationship management (CRM) system or send alerts?

Not in this design. Write and send are separate steps behind a human approval. The agent drops notes in a review store. A person or a deterministic step moves approved notes onward.
