· 4 min read · Alan Fu

When my guardrail crashes, the agent gets a no

The easiest way past most security checks is to make them fall over. I built Doberman so that its own failure is a denial, and that choice has a cost worth being honest about.

The easiest way past a security check is to make it fall over. Throw an exception inside it, let the caller catch the error and carry on, and the check was never there. Most guardrails around AI agents work like that, because the guardrail is a paragraph in the system prompt and a paragraph can’t throw.

Doberman is the check I built for coding agents. It runs before every shell command, file write, and tool call and returns PASS, AUTH, or BLOCK. The rule I built the whole thing around is that any error, any uncertainty, any case the code didn’t anticipate, resolves to a denial. If Doberman crashes, the agent gets a no.

That includes silence. When an action needs approval and nobody is at the keyboard, the desktop dialog gives up after two minutes and the whole approval flow gives up after ten. Both are denials. They’re logged as timeout rather than denied, because a person saying no and no person being there are different facts, and I want to see which one it was when I read the log later. Agents run unattended most of the time, so that entry shows up a lot.

Compare that with the guardrail in the prompt. Once an injection has talked the model past it, nothing upstream of the model is left to help. A check on the action doesn’t care what the model was told or what it believes. The command runs or it doesn’t.

There are two ways an agent can go after the check itself, and I had to close both.

The first is turning it off. Doberman’s state lives in a .doberman directory and its hooks live in the host’s config, .claude/settings.json for Claude Code and .codex/hooks.json for Codex. A write, delete, or read of any of those through the engine is a hard BLOCK, and so is a shell command that merely names one of the paths. rm -rf .doberman gets caught. echo > .claude/settings.json gets caught.

The limit there is honest and I’d rather say it than have you find it. This stops the agent from disabling its own oversight, and it doesn’t stop a human. Someone at the keyboard can always turn a hook off, Codex even ships a flag for it. And a control-plane path built at runtime, out of a shell variable or a python -c payload, isn’t something a static read of the command text can see.

The second is going after the neighbours. An agent that could only delete a file under a prompt injection, but could kill the operator’s database, IDE, or CI runner with no check at all, was a real gap. So kill, pkill, killall, taskkill, Stop-Process, and the interpreter one-liners that call os.kill or psutil’s terminate all step up to AUTH now. A probe like kill -0 stays PASS, but only when that flag is the only option present. kill -0 -9 is treated as a real signal, because I don’t know which kill implementation is running and some of them read the second flag.

There’s a third property holding this together, and it’s the one I’d keep if I could only keep one. The checks may tighten themselves and may never silently loosen. The rule that merges verdicts can only go up. A permanent weakening of policy needs a person, and a person proving they hold something, a TOTP code or the local password, so the approval can’t come from the thing being approved.

Now the cost, because fail closed has one and the posts that praise it never mention it.

Fail closed means real interruptions. I installed the hooks globally on my own machine to test them and my own coding session got gated the same afternoon. Commands that named the .doberman directory were blocked, and git push needed an approval. Safe to say the dog bites its owner too. [VERIFY: this is from your own session this week; cut the paragraph if you’d rather not tell it.] That’s the behaviour I want and it’s still friction, and if the friction fires six times for every real problem, people click approve without reading and the AUTH tier stops meaning anything. I measured that ratio on one detector, fixed it, and wrote the fix’s own cost into the limitations file. That’s a separate post.

The other cost is that a denial you didn’t expect looks like a bug. doberman log --why shows the reason codes and a one-line explanation for every BLOCK and AUTH, because a check that says no without saying why gets disabled by the person it’s protecting.

What does your guardrail do when it throws?

The fail-closed section of the README: https://github.com/dobermancore/doberman-core#why-doberman

Also published on Substack ↗

Doberman gates every tool call your AI agent makes — PASS / AUTH / BLOCK before anything executes. Open source, fail closed.

$ pip install doberman-core
View on GitHub ↗

More field notes