· 5 min read · Alan Fu

What it cost me to cut the third leg of the lethal trifecta

Simon Willison’s rule says an agent can’t safely hold private data, read untrusted content, and have a way out. A coding agent needs all three, so I put a check on the third one. This is the bill.

Cutting the exfiltration leg out of a coding agent cost me a 515-line file of things the cut doesn’t cover, about six false alarms for every real secret until I tuned the detector, and a proxy that only speaks one HTTP method. That’s the price so far. None of the posts that quote the lethal trifecta mention one.

Simon Willison’s lethal trifecta is the clearest description of the problem I’ve read. An agent that has access to private data, reads untrusted content, and has a way to send data out can be made to leak that data by anyone who can get text in front of it. His advice is to never give one system all three.

A coding agent has all three before lunch. The private data is the repo, the .env file, ~/.aws/credentials, whatever is sitting in the shell history. The untrusted content is every README, issue, dependency, and web page the agent reads to do its job. The way out is git push, curl, pip install, gh api. Take away any one of those and the agent can’t do the work you hired it for.

So the diagnosis is right and the treatment it implies, remove a leg, doesn’t fit the patient. What I did instead was leave all three legs attached and put a check on the third one, because the third one is the only leg that comes with a chokepoint.

Private data and untrusted content are text. Text can talk a model into anything, and once it has, nothing upstream of the model is left to help you. Sending data out is an action. A shell command runs or it doesn’t, a socket opens or it doesn’t, and something can sit in between and say no.

That something is Doberman, the AI guard dog to stop your AI when it goes rogue. It sits on the tool-execution path of Claude Code, Codex, OpenClaw, and anything that speaks MCP, and every shell command, file write, or tool call gets one of three verdicts before it runs: PASS, AUTH, or BLOCK. AUTH means a person has to approve it.

On the exfiltration leg the verdicts look like this. Piping a secret into curl or nc is a hard BLOCK. Any other command that leaves the machine, including one aimed at a host that looks fine, steps up to AUTH. And a value Doberman saw the agent read from an untrusted source, and then sees again inside something being sent out, gets matched on its fingerprint. The same wget passes in one session and gets blocked in another, depending on what the agent read ten minutes earlier.

That’s the treatment. Now the bill.

Item one. A static check reads the text of the command. It can say “this looks like egress” and it can’t prove the host it read is the socket the process opens. I wrote the list of things that route around it into the limitations file: a redirect file, curl’s --resolve and --connect-to flags, an HTTP_PROXY environment override, DNS rebinding, a URL built at runtime, a git push to a remote that was configured before the agent showed up, a package’s install-time script, a trusted service used as a relay, and anything launched from a child process. Every one of those is in the doc because a reader who finds one on their own stops trusting the rest of the README.

Item two. The secret detector that feeds the exfil check has a weak heuristic in it, Shannon entropy, and Shannon entropy can’t tell a base64 secret from a variable name, a relative path, or a build tag like py311. They all land in the same 3.6 to 4.5 bits per character range. Before I tuned it, that check fired about six times for every real hit. Six false alarms per real one is how you train a person to click approve without reading, which turns the whole AUTH tier into theatre.

So I added exemptions, and every exemption is a hole you chose. I measured the hole. For a bare base64url secret with no key name next to it, the chance that every piece of it happens to look exempt rises from about 0.8% to 3.6% at 24 characters. The stronger credential-shape check isn’t affected. That number is in the doc too, because an exemption I can’t put a number on doesn’t ship.

Item three. Real containment of the third leg means controlling the connection itself, and the command text is only a guess at it. The building blocks for that exist in Doberman now and they’re opt-in, unregistered by default. A broker interface the engine consults on every egress-classified action. A reference broker with a default-deny allowlist and a forward proxy, built on Python’s standard library, that enforces the allowlist at the socket so a blocked destination’s connection is never opened. The proxy only supports HTTP CONNECT. No SOCKS, and no way to transparently intercept traffic that wasn’t routed to it.

The static check only hands out a PASS on the broker’s say-so when the broker proves it enforces egress and reports that it will enforce this exact destination. A broker’s connection history can raise a decision to AUTH when it disagrees with the text. It can never lower one.

What survives all of that is two promises, and they’re the only two I’ll put in writing. Fail closed: any error, any unhandled case, any approval prompt nobody answers within two minutes on the desktop or ten minutes as the backstop, resolves to a denial, logged as timeout rather than denied so you can tell nobody-was-there from somebody-said-no. Raise-only: the checks can tighten themselves and never silently loosen, and a permanent weakening needs a person proving they hold something, a TOTP code or the local password.

The trifecta is a diagnosis. The bill above is what treating one leg of it costs.

Which leg have you actually cut in your setup, and what did it cost you?

The limitations file: https://github.com/dobermancore/doberman-core/blob/main/docs/LIMITATIONS.md

Also published on Substack ↗

Doberman gates every tool call your AI agent makes — PASS / AUTH / BLOCK before anything executes. Open source, fail closed.

$ pip install doberman-core
View on GitHub ↗

More field notes