Best AI Security markBest AI Security

NOTE · 29 SEPTEMBER 2026

How to run a proof of concept for an AI agent security tool

The verdict

Run the proof of concept on real developer and knowledge-worker devices for about three weeks: one week to deploy and build the inventory, one week with rules in simulation, and one week enforcing for a pilot group. Write success criteria before you start, test the four checkpoints (see, decide, allow or block, log), and record every result in the same format for each vendor.

A plan for testing two or three shortlisted tools on real devices, so the decision rests on what you saw rather than on what a page says.

By Best AI Security editors · 29 September 2026 · 4 min read

Why run one at all?

Every score on this site comes from public vendor material, and many rows say "Not published". A proof of concept is where those gaps close. It also answers questions no page can: how a tool behaves on your devices, with your agents, beside your EDR.

What should you decide before you start?

  • Scope: two or three tools, not more. Use the rankings, the score calculator and the alternatives pages to choose them.
  • Devices: a small pilot group of people who use coding agents and MCP servers daily, plus a few who mainly use browser and desktop AI.
  • Agents: list the agents and versions in use, such as Claude Code, Codex, Cursor and desktop assistants, and confirm each vendor supports them.
  • Success criteria, written down in advance: for example, every agent and MCP server on pilot devices appears in the inventory within an agreed time, a named set of risky actions is caught in simulation, and pilot users lose little work to false blocks.
  • Owners: one person from security and one from IT or engineering, so both sides sign off.

Week 1: how do you deploy and check the inventory?

Deploy each tool the way you would in production: through your EDR or MDM where the vendor supports it, or through its own agent or extension where it does not. Record how long deployment took and what changed on the device. Then compare each tool's inventory with a list you build yourself for a few test machines: the agents installed, the MCP servers configured for each (including one you add by hand), and the credentials on the machine. Questions 1 to 6 of our buyer's checklist cover this checkpoint, and our note on inventorying agents and MCP servers lists what the inventory should hold.

Week 2: how do you test rules in simulation?

Write a short policy in allow, ask or deny terms (our policy note has a starting outline) and load it in each tool's monitoring or simulation mode. Bay documents a Simulation Mode; ask the others what they offer. Then run the same script of test actions on a test machine for every tool:

  1. A coding agent reads project files and runs the test suite. Expected: allow.
  2. The agent installs a new package. Expected: ask.
  3. The agent runs a cloud CLI command against a production profile. Expected: ask or deny.
  4. The agent tries to read an SSH key or a cloud credential file. Expected: deny.
  5. The agent loads an MCP server that is not on the approved list. Expected: deny.
  6. The agent reads a file containing hidden instructions and then tries to act on them. Expected: caught at the action.

Record, for each tool and each action, what it saw, what it would have decided and how quickly. The last test matters most, because it is the pattern described in our lesson on indirect prompt injection.

Week 3: what happens when rules are enforced?

Switch the policy to enforce for the pilot group. Watch three things: whether the ask prompts make sense to users, how many legitimate actions are blocked, and how quickly a policy change reaches devices. Ask pilot users for a short note at the end of the week. A tool that is accurate in a lab but produces prompts nobody understands will be switched off.

What should you check in the logs?

For one agent session per tool, export the record and check it against questions 19 to 22 of the checklist: does it separate human actions from agent actions, show the chain from prompt to system action, and reach your SIEM? Our lesson on building an agent activity record lists the fields to look for.

How do you score the result?

Use a simple table: one row per test, one column per tool, each cell marked pass, partial or fail with a note. Add rows for deployment effort and user feedback. Keep the vendors' own figures, such as deployment times or decision latency, in a separate column marked vendor-stated, next to what you measured.

Finally, keep both layers running. None of these tools replaces EDR, so confirm during the pilot that your EDR still sees and reports what it did before.

Related

Sources