← Security Research
Complete
Security Research · Prompt-Injection Defense · The Recipe

Break the Triangle

Prompt-injection protection for the rest of us — the plain-language version. If you just want to be safe without reading forty pages, you're in the right place.

This is the short version. The full defense — the HOW behind each step — is the companion paper, Killing the Lethal Trifecta. The story of how each rule was learned, failures included, is the Lab Notes.

The danger, in one breath

AI coding agents read the web, your files, and your dependencies. Hidden text in any of those — a web page, a README, a code comment, a package — can quietly hijack the agent into leaking your private data or running an attacker's commands. This isn't a bug waiting for a patch: the AI genuinely can't tell your instructions from instructions buried in the stuff it's reading. It's the #1 security risk for AI applications, and it has already happened to real, shipping products.

You can't stop the trick. So you make the trick worthless.

The trick (how you beat something you can't out-fight)

What's the easiest way to stop a tank? You don't out-gun it — you throw a blanket over the viewport so the crew can't see out. All that armor and firepower is still there; it's just useless, because the tank depends on vision to aim. You didn't defeat the tank. You removed the one thing it needs.

Prompt injection is the tank: you can't out-gun it (undetectable, no patch coming). So don't. Blind it instead. Your AI agent depends on seeing all three danger-sides at once to hurt you — so you cover the dangerous eye on each one and leave the useful eye open. Still powerful, now harmless.

The one idea

Here's the "vision" the blanket takes away. A breach needs three things at the same time, in the same agent:

  1. Private data — something worth stealing. And it's more than you think: the biggest prize usually isn't your code, it's your browser — logged into your email, bank, and everything else. Any code running on your machine as you can read those live sessions without ever needing a password. "I'm just making a small tool, I have nothing to steal" is exactly the mistake — you're already logged into everything worth stealing.
  2. Untrusted input — poison from outside (the web, a repo, a dependency).
  3. A way out — the ability to act or send data outward (run commands, hit the network, push code).

Like a fire needs heat, fuel, and oxygen. Take away any one side and there's no fire. You don't try to detect the bad instruction — attackers always find a new one. You remove a side by design.

The recipe — your windshield

You can't stop the storm. The internet is hostile and always will be. But a windshield lets you drive through it. Here is the whole windshield:

1. Two boxes. Run two separate, throwaway boxes (a VM or sandbox):

Neither box can start a fire, because each is missing a different side of the triangle. (Two blankets, really: the settings below are a light blanket the agent could shrug off if the config got deleted; the VM is the blanket bolted on from the outside, that a hijacked agent can't reach to lift. That's why the box is the one that really counts.)

2. Two settings files — mirror images of each other. In each box, one small config tells the agent what it's allowed to do:

That's the whole switch, and the two files are opposites of each other. (Denying web in the Build box is not optional and the box's network settings can't do it for you — see the full paper for the one surprising reason why.)

3. You are the only bridge. Nothing moves from Research to Build automatically. When you carry something across, you paste it as plain text and glance at it first. You are the one part of this whole system that can actually tell your instructions from a stranger's — so you are the gate.

4. Never run unreviewed output on your real machine. Whatever the Build box produces, review it and rebuild it from source before you run it for real. This is the one crossing that can actually hurt you. In plain terms, four quick moves: (a) let the build script run its automated checks and fail loudly if anything's off; (b) glance at the flagged lines — the tool tells you what to ask ("why does this need the internet?"); (c) test-run it first in a throwaway box — an offline tool should work with no network at all; an online tool should reach only where you expect; (d) then run it as a normal (non-admin) user.

That's it. Two boxes, two mirror-image settings, one human bridge, one rule about running things.

One behavior matters as much as the boxes: don't let agents run on auto outside a box. When you're watching each step, you can be the guard — but only if you actually understand what you're approving. Clicking "yes" on a prompt you can't read isn't guarding; it's waving it through. So the real danger is any time nothing informed checks an action before it runs — whether because you automated (swarms, auto-approve), or because you're approving things you don't understand. Either way the fix is the same: the box. Inside one, there's nothing for a runaway — or a rubber-stamped — agent to reach.

Why this is enough

Notice what you didn't have to do: no filtering the web, no detecting malicious prompts, no antivirus for AI. You simply arranged things so that when the agent inevitably reads something poisonous, there's nothing to steal and nowhere to send it. The danger walks in the front door and finds an empty room.

Which tools this works with

The boxes (VMs) work with any AI coding tool — Claude Code, Cursor, Windsurf, Devin, whatever. That's the part that actually protects you, and it's universal. The settings-file fence, though, is specific to Claude Code on Windows — it can do things more locked-down tools can't. If you're on a different tool: skip the fence, keep the boxes. The box was always the real wall; the fence is a bonus that some tools let you add and others don't.

The optional insurance

There's a fancier layer you can add — automated guardrails that enforce the Research box's rules from the inside, shout loudly if they ever break, and repair themselves on restart. It's good insurance, but it is insurance, not the wall. The two boxes and two settings files are what actually protect you. If you skip it, you're still safe.

Tested, and worth knowing: that extra insurance layer works on your real machine — but not automatically inside a throwaway VM. The same guardrails, vendored a second time inside a disposable research box, did nothing at all on first live test — every rule it was supposed to enforce was silently ignored. The coding tool requires you to confirm you trust a folder before it applies that folder's own rules, precisely so a hostile downloaded folder can't silently grant itself power. A disposable VM resets completely on every launch, so it never earns that trust — the extra layer stays off by construction. This isn't a reason to skip the VM — the box's own wall held anyway, proven against real, disclosed, maximum-severity bugs in the coding tool itself. It's a reason not to count on the extra insurance once you're already inside one.

The honest catch

No skill level saves you from this, given enough volume and time — not a beginner's, not an expert's. The beginner can't spot the danger; the expert gets overconfident and tired and stops looking. The right setup and habits mitigate the risk; they don't eliminate it. This is not zero risk — nothing is. A determined, targeted attacker, or a brand-new VM-escape bug, can still get through, and no setup protects you from every mistake you personally make. What this does is drop the blast radius from "your whole machine and everything on it" to "a throwaway box you can delete." For a solo developer running AI agents, that is the difference between a bad afternoon and a real breach.

Set it up once. Then drive through the storm.


Want more? The Lab Notes are the receipts — what actually broke, what it taught, in the order it happened. Killing the Lethal Trifecta is the full defense: the threat model, the architecture, and the build walkthrough, appendices included.