← Security Research
Complete
Security Research · The Receipts

Lab Notes — Discoveries Along the Way

What actually happened while building and testing the defenses, in the order it happened, including the parts that failed.

Companion to Break the Triangle (the recipe) and Killing the Lethal Trifecta (the white paper — the HOW). These are not a paper. They are the receipts. The white paper cites these sections for the WHY behind its rules; read them when a rule sounds arbitrary and you want to know what it cost to learn. Drafted with AI assistance from build logs and live transcripts; discoveries are told as they happened, mistakes included — that's what lab notes are for.

§1 · The copy-paste that ate the fence

The first research-mode fence denied all writes — airtight on paper, and nearly useless in practice: the only way to get the agent's output out was copying text from a terminal by hand. That channel turned out to be badly lossy — line-wraps became real newlines, rendered views flattened structure, encodings mangled characters, and rich-text paste re-interpreted whatever rode along. The fence designed to leak nothing also failed to deliver anything, and the workaround it invited (raw-pasting into a word processor) was worse than the narrow exception it forced us to design: one new, inert .txt file, in the project folder, never overwriting anything.

The same lossy channel then bit a second time, harder: a copy of the enforcement hook itself got corrupted in a terminal round-trip — twelve parse errors — and the paper's own earlier draft was mangled the same way and had to be reconstructed.

Lesson: a control that doesn't survive real use gets bypassed with something worse. And never move a hook file through a chat window or a terminal — edit on disk, always.


§2 · The fence that failed open

That corrupted hook exposed something worse than lost text: a PreToolUse hook that can't launch is treated as a non-blocking error — the agent proceeds. A broken fence is no fence, silently. We had the layering backwards in our heads: the clever hook was never the defense; the boring static deny-list (which fails closed) was. The rebuild made failure loud — a guard that parse-checks the hook on every call and forces a visible prompt when it's broken, plus an out-of-band canary alarm.

Lesson: know which of your layers fail open and which fail closed, because the answer decides which one is actually the boundary. The hook is granularity; the deny-list is the backstop; the VM is the wall.


§3 · What the hand-audit caught that no test did

Days after the scoped write exception shipped, a plain re-read of the hook logic — asking "what does this check actually not check?" — found two gaps the happy-path tests never touched. Writing newfolder\notes.txt silently created newfolder: the guarantee said "one new file," the capability was "one new file plus an arbitrary directory tree." And the containment check was purely lexical — it could not see through a directory junction, so a reparse point under the project folder would have let a lexically-inside path resolve anywhere. Both fixed by narrowing: parent directory must already exist; walk the ancestor chain and refuse reparse points.

Lesson: a control has to survive contact with itself. No scanner asks "what does this check not check?" — a human does, with fresh eyes, later.


§4 · The belt-and-suspenders that was never buckled

Research mode moved into a disposable VM, and the same four-layer fence was vendored inside the box as defense-in-depth. Live probing — asking the caged agent to attempt four forbidden writes — found all four succeeded. The inner fence wasn't degraded; it had never been on. Root cause: the coding tool gates repo-controlled settings and hooks behind a workspace-trust confirmation (hardened after a real disclosed vulnerability), and a disposable VM that resets on every launch can never durably earn that trust. The property that made the box strong was the property that kept the fence off — by construction.

Two details worth preserving. The diagnosis was co-discovered with the caged agent itself — the sandboxed instance pushed back on the host session's first theory, spotted the sharper clue (its instructions loaded; only enforcement was dark), and proposed the trust-gate explanation the CVE record then confirmed. And the fix respected why the gate exists instead of defeating it: the launcher now seeds the policy into the VM user's own settings — a scope that belongs to the user, not to a possibly-hostile folder, so it loads without a trust prompt. Verified the same way the failure was found: by watching it deny things.

Lesson: a control you haven't watched fire is a control you don't have. Install is not enforce.


§5 · Loosening, on purpose

Once the inner fence was actually alive it failed in the opposite direction: it was a copy of the host fence — no reading, no sub-agents — inside a box that holds nothing private. The agent couldn't read the brief dropped for it; one reader worked serially where a swarm works in parallel. Friction with no security payoff. So the in-VM dial was loosened — reads/searches scoped to the one drop-folder, sub-agent swarms allowed (every sub-agent's tool calls re-enter the same hook) — while the host fence stayed exactly as tight as ever, because the host has something to steal.

The loosening created one new seam, worth writing down: the drop-folder became an input channel. A hijacked session could plant instructions in a findings file that a future session reads — injection carried across sessions by the rig's own plumbing. Mitigations: a hook-enforced warning header on every findings file, a loud launch-time warning listing anything already in the folder, and a human promotion checklist before anything in it is trusted.

Lesson: tightness is a dial you set by what the environment holds, not a virtue in itself. Loosening the fence inside the box is not loosening the architecture — the box didn't move.


§6 · Nobody audits a real dependency tree

"Audit dependencies before install" sounds actionable until you install PyTorch — one deliberate choice — and watch it pull roughly forty packages. No human reviews forty packages; expecting one to is the "be vigilant" fallacy wearing a checklist. What actually works at solo scale: the human audits the decision (right name, not a typosquat, maintained, mainline index — and trust is transitive: choosing torch is choosing torch's choices); the tooling audits the tree (pip-audit over the resolved set, hashes pinned so it can't drift); the structure catches what both miss (the install and first run happen inside a no-egress box).

Lesson: split the audit the way the ghostline splits review — human judgment on the few load-bearing choices, automation on the bulk, a box around the remainder.


§7 · The host's own immune system got in the way

The sandbox refused to auto-launch its agent, and the trail led somewhere unexpected: the machine's OEM-bundled endpoint security (HP Wolf) blocks PowerShell specifically when auto-started at logon — the signature of a malware persistence trick — while the identical command run by hand works fine. (It also, probably, explains the sandbox's occasionally dead keyboard until another app gets focus.) We stopped fighting it: the sandbox boots to a plain Explorer window, and a human starts the launcher.

Lesson: your own defensive stack is part of the environment your controls must survive. Sometimes the right fix is to stop resembling an attack.


§8 · The market started shipping friction

Mid-project, the coding agent's vendor shipped real guardrails: automatic prompt-injection checks, per-website fetch approvals. Genuinely welcome — and on first contact with a real workload (hundreds of fetches across a sub-agent swarm), the per-site prompt was switched off within the hour, because nobody meaningfully approves website #347. Not the vendor's failure, not the user's: safety that arrives as friction gets disabled by exactly the people running the hottest workloads. Safety that arrives as architecture — a box in which the friction can be safely deleted — doesn't depend on anyone's patience.

Lesson: the white paper's own §4, observed in production. Guardrails are welcome; they are still not the wall.


§9 · The dead guard that passed its own inspection

The build container's isolation proof runs five checks: no raw egress, no DNS, non-allowlisted domain blocked, the model API reachable, no host drives visible. First real run: four passed, one failed — the model API was unreachable. A locked box with the one road down.

The failure was the good kind, because chasing it exposed something much worse. The proxy — the single container allowed to touch the internet — was in a crash-restart loop and had never listened at all. It couldn't write its own log file (it drops privileges to a user that can't write to the container's stdout, which squid treats as fatal). And here is the part worth the whole section: every "must be blocked" check still passed. Of course they did. A dead proxy blocks everything. The suite cheerfully reported that the allowlist was working when there was no allowlist — no proxy — at all.

A deny-test cannot tell refused from nobody home. So the suite grew a new first check, before all the others: prove the guard is alive and answering. Only then does "it blocked me" mean anything.

The same run had a second, quieter version of the same disease. Once the proxy was alive, the blocked-domain check still reported failure — because it was reading the wrong signal. For an HTTPS request through a proxy, the proxy's verdict arrives as the response to the tunnel request; the test was reading the status of the tunnelled request instead, which doesn't exist when the tunnel is refused, and which therefore reads as zero — indistinguishable from dead proxy. Two different bugs, one shape: the test was not measuring the layer it claimed to be measuring. Squid's own log settled it in one line, showing the refusal and the allowed tunnel side by side.

Lesson: a control you have not watched work is a control you do not have — and a test that only ever watches it fail cannot tell the difference between a guard and a corpse. Assert the positive case, on the layer you're actually making a claim about, or you are testing nothing.


§10 · Two Claudes arguing

The review method that caught the most real defects wasn't a scanner; it was adversarial instances. First by hand: two sessions open, each red-teaming the other's conclusions (that's how §4's trust-gate diagnosis happened). Then deliberately: after the paper's revision, two agents were spawned with refutation briefs — one attacking the paper's claims against the actual code, one hunting internal contradictions. They confirmed nine real defects, including two in text written that same day, and one that was a live code bug, not a paper bug: the second sandbox tier's fence had the exact §4 problem — configured, parsed, and never seeded, so never on.

One rule keeps this safe: the critic writes objections to a notes file; a human (or a supervised session) applies them. Refute-and-edit on a real machine is an unattended write loop — the exact auto-outside-the-box pattern the architecture forbids.

Lesson: the cheapest competent reviewer you'll ever hire is another instance of the same model told to prove you wrong — as long as it can only talk, not touch.

(§9 is the same lesson without the second instance: the isolation suite was the reviewer, and it took a human asking "wait — why did that one pass?" to catch it lying.)


§11 · The boop — the trusted lane proving the whole thesis, live

Setting up the daily-driver build box, the box command "did nothing" — typed in a project folder, no box. Diagnosis: Windows ships with PowerShell scripts disabled (Restricted), so the profile that defines box had never loaded (which also meant the fence-status indicator had been silently dead this whole time). The unboxed host session — a normal Claude Code on the real machine, no VM — fixed it in one command: relaxed the execution policy to RemoteSigned and edited the user's PowerShell profile. Working as intended, approved in spirit, a two-line convenience.

The author's reply was the paper: "you've accidentally proved why I need a Docker for my work — 'let me just touch your Windows settings, boop.'"

That "boop" is the entire threat model in one word. Nothing about the edit was hostile. But the same reach — a session on the real machine, casually writing a file that executes in every future console — is exactly what prompt injection hijacks. The host lane treats capability as ambient; a poisoned page read by an unboxed agent inherits that ambience. The demo wasn't a drill we designed; it fell out of ordinary setup work, which is what made it land.

Two corrections came out of arguing it through, both worth recording because the first answer was too confident:

Execution policy is not the security boundary, and "scripts disabled" is not the win it feels like. It is documented (Microsoft's own words) as an accident-guard, not a control: powershell -ExecutionPolicy Bypass defeats it, piped and encoded commands ignore it, and it binds exactly one actor — the human at the prompt. Restricted had been disarming the defender (profile never ran) while stopping none of the threats the boxes exist for. But the tidy version of that claim ("it's pure theater") was overstated: friction against your own accidents has real value, and the mark-of-the-web gate RemoteSigned keeps is the main way "just run this script" tricks get refused. Both true at once.

Re-enabling the profile re-opened a persistence channel — and that is the cost, stated plainly. Anything host-side that can write $PROFILE now runs in every console, forever. The saving grace is structural, and it is the thesis restated: a box cannot reach the profile. $PROFILE lives outside every mount; only /work crosses in. So the re-opened channel is reachable from the trusted lane (host sessions, host software) and sealed off from the high-volume untrusted lane (everything boxed). The fix matched the doctrine already used elsewhere: don't try to prevent tampering of a file whose writer could also flip the policy — make tampering loud. The profile and the fence-indicator script were folded into the same hash-baseline guard, tamper-evidence for a channel that can't be tamper-proofed.

Lesson: the box is not for when things look dangerous — it's for the "boop," the friendly two-line convenience on the trusted machine, because that reach and an attacker's reach are the same reach. The honest confidence calibration is its own lesson: the first pass called execution policy "theater" and had to be walked back to "accident-guard, not boundary" — right conclusion, overstated middle. Red-teaming your own confident answer is the §10 rule pointed inward.


§12 · The offline box that cried wolf

The detonation box's whole promise is no network, and it proves that to itself on every launch — a self-check that tries every way out and reports what got through. First real run on the Windows lane, it flagged an escape: DNS resolved a name. An offline box that can still reach a resolver would be a hole in the one wall it sells.

It wasn't. The check was reading the wrong thing. Offline, a name can still resolve — from the local cache, from the hosts file, from a negative answer Windows hands back without a single packet leaving. The test asked "did a name resolve?" and treated yes as egress; but resolution is not a query on the wire. The honest test asks what the box actually claims: can a query reach an external resolver? Rewritten to probe reachability to a public resolver instead of mere resolution, the alarm went quiet — correctly. Every other vector had been contained all along; the escape was in the alarm, not the box.

Lesson: an alarm that can't tell resolved locally from reached the network is §9 wearing a different hat — a test measuring a proxy for the thing instead of the thing itself. Probe egress, not resolution, or you are testing nothing.


§13 · A quarter-million events, unread

With the box proven sealed, the plan was simple: run the suspect binary under Procmon, read what it did. Two minutes of capture produced roughly 250,000 events. No human reads that. The recording was complete and useless in the same way §1's fence was airtight and useless — a control that does not survive contact with a real workload.

Recording, it turned out, is only half the gate. The other half is a procedural analyst — a dumb, deterministic script that reads the dirty output and labels the handful of behaviors that matter: a reach for the network, a read of a saved-password or wallet store, a write to a persistence key, a spawned living-off-the-land binary, a dropped executable. The human reads a short verdict; the tool touches the quarter million. Never a model reasoning over the raw hostile bytes — the same split as §6's build gate: automation on the bulk, judgment on the few things that matter.

Lesson: observation without triage is just a bigger haystack. The dumb script that reads the firehose is as load-bearing as the box that produced it.


§14 · The detector that flagged Windows breathing

The analyst passed every synthetic test — planted behaviors caught, benign lines ignored. Then it met a real capture and flagged thirty-four suspicious events. Most were Windows itself: the process literally named Registry flushing its own hives, lsass writing the SAM database (its actual job), Explorer touching the copy of Procmon we had shipped into the box. The rules were technically right and practically crying wolf — matching legitimate system activity they had never seen, because the synthetic test file that vetted them contained none of it.

The fix was discipline, not a cleverer rule: trace every false alarm to its specific, legitimate cause and exclude that — the Registry process, the OS's own credential store, our own tool — never the signal. And focus on the one process you actually detonated, because the firehose across every process on the system is noise by construction. Thirty-four collapsed to the handful that were really the bogey.

Lesson: a detector tuned on synthetic data lies on real data. You tune it the only honest way there is — watch it fail against the real thing, and run each false alarm back to the legitimate behavior that caused it.


§15 · Catching the question, not the answer

The move that beats a detonation box isn't force; it's patience with a checklist. A bogey that detects it's being watched simply does nothing — behaves like a saint on its one recorded run and saves the payload for a real machine. So the analyst learned to flag the probe itself: a program reading VM driver files, hypervisor registry keys, even hunting for the analysis tool watching it. That probe is high-signal precisely because a program with nothing to hide never asks am I in a sandbox? Turned loose on a test bogey that checked a handful of VM tells, it caught every one.

But the honest limit is sharp, and worth writing down. We can see the question; we cannot hide the answer. Our box is a Windows Sandbox — it announces itself as the WDAGUtilityAccount user, it runs on Hyper-V, and it carries the very Procmon a bogey looks for. And the timing tricks — CPUID, RDTSC — happen inside the CPU, where Procmon cannot see them at all. The detector catches a bogey asking whether it's watched; it cannot make the true answer be no.

Lesson: you can detect the anti-analysis move, and it's worth doing — but detecting the question is not hiding the answer. A determined probe still learns it's in a box; what you've bought is knowing that it asked.


§16 · The one it would have missed

Asked the obvious question — could this have caught SolarWinds? — the honest answer was no, and chasing why was the most useful thing in the arc. SUNBURST sat dormant for twelve to fourteen days; a detonation run is minutes. It would have shown a clean, quiet Orion process and been waved through. Worse, the ground-truth edge the whole detonation idea leans on — your calculator has no business touching the network — collapses for Orion, whose entire job is network monitoring. The one signal that's damning for a quiet app is invisible noise for a product built to talk to the wire, and SUNBURST's beaconing was designed to blend into exactly that normal Orion traffic.

The right gate for that attack is a different discipline: reproducible builds — compile the same source in independent clean environments and compare, so a binary that doesn't match a fresh rebuild of its own source exposes an injection before it ships. Tor and Debian have run this for years. Making a build bit-for-bit deterministic is real engineering, beyond this project's scope — so the honest move was to name the gate we don't have rather than imply the box covers everything.

Lesson: behavioral detonation is strong when your product is quiet and the adversary opportunistic, and weak when the product is noisy-by-design and the adversary patient. Know which case you're in — and name the wall you're missing, out loud.


§17 · The tool that doesn't travel

The last discovery wasn't a bug; it was about the deliverable. The kit works — but it's bespoke: Windows, Docker, our paths, our skill. Handing someone the exact scripts doesn't help them, and most people using an LLM couldn't rebuild this from the recipe. For a moment that read as failure — we solved our trifecta and couldn't ship the solution.

The reframe is the point. The tool is bespoke; the doctrine is portable — contain the blast, break the exfil leg, disposable-by-trust, works ≠ safe. The kit is a reference implementation and an existence proof, not a product. And even the doctrine has a floor: a paper that says use network_mode: none still needs a reader who can. The real mass answer was never a kit anyone downloads; it's containment shipped by default in the tools people already use — which a solo dev can spec and prove, but not personally hand to everyone. Usable security has to be the default or it doesn't happen.

Lesson: a bespoke control is a proof and a spec, not a product. What generalizes is the way of thinking, not the scripts — and being honest about that ceiling on reach is its own finding, worth as much as any wall.


The recipe these lessons produced: Break the Triangle. The full defense and build guide: Killing the Lethal Trifecta. The supply-chain and detonation half — the source of §12–17 — is Contain the Blast. Everything above is why those documents say what they say.