Pacing the Frontier Is Not Enough
Dario Amodei’s We Must Pace the Frontier makes a strong case that AI capability is beginning to move faster than our ability to understand and control it.
Dario Amodei’s We Must Pace the Frontier is basically an argument for buying time. Independent evaluations, stronger monitoring, better operational security and eventually coordination between labs and governments are all ways of slowing capability growth when our ability to understand or control these systems starts falling behind.
I agree with the premise. Where I think the conversation needs to go further is what we actually do with that time.
Even if frontier development slows, we are still going to deploy models that are imperfect, difficult to fully understand and increasingly capable of doing things in the real world. The problem is no longer limited to whether a model gives a bad answer. Agents are beginning to operate software, infrastructure and entire workflows on our behalf.
So the question cannot just be how we make the model safer. We also need to ask what happens when the model is wrong anyway.
That sounds like a subtle distinction, but I think it changes the security model completely.
The failure mode has changed
A few years ago, most model failures were informational. A model hallucinated a fact, wrote broken code or gave you bad advice. That could still be harmful, but there was usually some distance between the model output and whatever happened next. A person copied the code, acted on the advice or decided whether to trust the answer.
Agents are removing that distance.
A coding agent can now execute shell commands. A browser agent can navigate unfamiliar websites and submit forms. An infrastructure agent can touch cloud resources. A business agent can update databases, call internal APIs and make changes across systems without a human manually translating every output into an action.
The model is not just telling you what it thinks should happen anymore. Increasingly, it is the thing making it happen.
That means a wrong answer and a wrong action are no longer remotely the same kind of failure, especially when the action is being taken with production credentials.
The model does not have to be malicious
I think this is one of the more important lessons from the recent agent incidents because the public conversation still tends to jump immediately to deception, rogue models or some system actively deciding to act against us.
Those scenarios are worth studying, but there is a much more ordinary failure mode that I suspect will matter far more often: the model is simply wrong.
It misunderstands where it is. It misreads an instruction. It gets hit by prompt injection. It makes a reasonable inference from incomplete context and that inference happens to be false.
Then it acts on it.
That last part is what matters.
This is not really a new engineering problem. Humans misunderstand systems. Operators make mistakes. Software has bugs. We have never solved those problems by assuming the actor will eventually become perfect. We build controls around them because failure is expected.
AI agents should not be an exception.
Alignment and authorization are different problems
A lot of current AI safety work is still focused on shaping model behavior: better training, stronger system prompts, classifiers, evaluations, interpretability and all the rest. That work is necessary, but it solves a different problem from authorization.
Security systems have always separated intent from permission.
If an employee wants to delete a production table, their intention does not determine whether the database lets them do it. If a service asks for access to a resource, the request is checked against its permissions. We use roles, authentication, network controls and audit logs precisely because we do not treat good judgment as the final security boundary.
There is no reason AI agents should work differently.
I think it helps to separate the problem into three layers. Alignment is about what the model should want to do. Containment is about where the model is allowed to operate. Authorization is about whether a specific action should actually be allowed to happen.
Those layers overlap, but they do not replace each other.
A well-aligned model can misunderstand its environment. A sandbox can be misconfigured. An evaluation can miss a behavior that only emerges under a strange combination of tools, context and permissions.
The point of layering security is not that every layer is perfect. It is that one failure should not immediately turn into a system failure.
We may be overestimating how much we can predict
There is a broader issue underneath this.
A lot of AI safety still depends on predicting what the system might do before deployment. Can we identify dangerous capabilities? Can we enumerate the failure modes? Can we benchmark the model well enough? Can we detect risky behavior in advance? Can interpretability tell us what is happening internally?
We should keep pushing on all of those questions, but there is an obvious tension here. The more useful agents become, the more we specifically want them to handle situations we did not script.
That is the whole point of an agent.
Traditional software follows a path that somebody wrote. An agent gets a goal, looks at the environment and figures out the path itself. Every increase in autonomy expands the space of possible behavior, including behavior nobody thought to test.
At some point, trying to anticipate every bad path stops being a realistic security strategy.
This is normal everywhere else in security. We do not assume the firewall catches every attack or that the application contains no vulnerabilities. We layer controls because we expect individual defenses to fail sometimes.
AI security is eventually going to have to adopt the same assumption.
Capability and authority are growing together
One part of the pacing discussion that I think deserves more attention is that capability and authority are not independent.
As models become better, we naturally give them more access because doing so becomes useful.
A weak coding model gets autocomplete. A better coding agent gets access to the repository and maybe a terminal. Once it becomes reliable enough, you start connecting CI/CD, cloud infrastructure, databases and production systems because that is where the economic value is.
This creates an interesting dynamic. The model’s error rate can decrease at the same time that the consequences of each error increase.
A system that fails 10 percent of the time while suggesting text is annoying. A system that fails 0.1 percent of the time while executing millions of privileged actions can still be a serious security problem.
That is why I do not find “the models will get better” particularly satisfying as a long-term answer.
Of course they will get better. But better models are exactly what will convince us to trust them with more.
Capability creates trust. Trust creates access and access increases blast radius.
The security architecture has to improve along the same curve.
Sandboxing is necessary, but it is not the whole answer
Amodei is right to put more emphasis on sandboxing and operational security. Agents need much better containment than they have today.
The limitation is that containment only works if the boundary you defined matches the boundary that actually exists.
If there is an exposed credential, an unexpected network route or some external service the developers did not realize was reachable, an agent may discover and use it. In fact, as agents get better at exploring unfamiliar environments, we should expect them to find more of these things, not fewer.
That does not mean sandboxing is the wrong approach. It means sandboxing should be one layer.
Train the model to behave safely. Restrict the environment it can operate in. Limit what credentials it receives. Add independent checks around sensitive operations. Log what happens. Escalate high-risk actions when necessary.
None of this is especially novel if you come from cybersecurity.
What is novel is the actor sitting in the middle of the system.
We are giving probabilistic models control over deterministic infrastructure. The model can be uncertain, confused or operating with incomplete context. The shell command it emits is not uncertain. The database does not interpret DROP TABLE probabilistically.
There is a mismatch there that we have not fully designed around yet.
Pacing only matters if we use the time well
This is where I think Amodei’s argument is strongest.
Slowing down is not particularly useful on its own. The value of pacing is that it gives us time to build the things that capability development is currently outrunning: better evaluations, better interpretability, better monitoring, stronger containment, clearer governance and a much more mature security model for agents operating real systems.
The mistake would be to spend that time chasing perfect reliability.
I do not think perfect reliability is a realistic assumption for these systems, especially once we deliberately give them autonomy and place them in environments their developers cannot completely predict.
We should assume models will make mistakes. We should assume prompt injection will continue to exist. We should assume an agent will eventually encounter something we did not test and that some of our safety mechanisms will fail.
Then build around those assumptions.
That changes the question from:
How do we stop the model from ever making a bad decision?
to:
How do we make sure a bad decision does not automatically become a bad action?
That is the problem I have been working on with Doberman, but I think the broader idea matters far more than any one project.
There is going to be a security layer around autonomous agents. It will probably look familiar in hindsight because most of the concepts already exist elsewhere in computing: permissions, least privilege, runtime controls, auditability, escalation and independent enforcement.
What we have not figured out yet is how all of those pieces should work when the thing asking for permission is an AI agent making thousands or millions of decisions on its own.
Right now, we are moving very quickly toward giving these systems real authority while still relying heavily on the model itself to decide what is safe.
I do not think that scales.
Pacing the frontier may give us some time. The useful question is whether we spend that time trying to make models incapable of failure or building systems that remain secure when failure inevitably happens.
Also published on Substack ↗