The security layer for agents is filling with detectors. A prompt-injection classifier. A tool-poisoning scanner. A monitor for a corrupted memory. Each one has to recognize the attack before it can stop it, and the attacker only has to find the technique the catalog has not named yet.
That is an arms race, and in an arms race the defender is a step behind by construction.
The detector has to know the attack.
A detector is a control that must recognize the thing it stops. It carries a model of what an attack looks like, whether that model is a rule, a pattern, or a trained classifier. It fires when the input matches. It stays quiet when the input does not.
Every attack technique that has been catalogued can be matched. Every technique that has not been catalogued passes. The public MITRE ATLAS catalog of adversarial techniques against AI systems keeps growing for a reason. Each new entry is an attack that worked before anyone had a name for it.
The agent case sharpens the problem. A subverted agent does not break out of anything. It uses the tools it was already cleared to use, in an order it was already able to walk. There is no breach for a perimeter to catch. The only thing that changed is what the agent is now trying to do.
What the adversary actually changes.
Goal-hijacking is the technique the last few months have made concrete. A poisoned document. A tampered tool result. A planted instruction sitting inside retrieved content. The model reads it as part of its task and adopts the attacker's objective as its own. From that point the agent pursues a goal it was never given, using authority it was legitimately granted.
The permission set does not move. It was drawn before the run, and it says the same thing to the hijacked action that it said to every honest one. The identity is intact. The credentials are valid. What has changed is not visible at the boundary at all. It is visible only in the behavior. The agent has left the trajectory it was assured to follow.
A different variable.
A runtime authority control reads that variable directly. It does not ask what the input was or whether the input matches a known attack. It asks whether the agent's realized behavior has diverged from the trajectory it was assured to hold, and it lowers the authority of the next action as the divergence grows.
The cause of the divergence does not enter the calculation. A novel injection nobody has named produces a divergent trajectory in exactly the way a catalogued one does, because the attacker's goal is not the agent's assured goal. The control does not have to name the attack to bound its effect. It reads the effect the attack produces and clamps what the agent is allowed to do next.
This is why the substrate is not in the arms race. The detector's problem scales with the number of attack techniques, and that number only grows. The governor's problem is fixed. There is one thing to watch, which is whether behavior stayed inside the envelope, and one lever to pull, which is how much authority the next action carries.
What this does not claim.
The governor is not a detector and does not replace one. Naming an attack is still worth doing, because an attack you can name is one you can stop before it ever moves the agent. The two controls sit in different places and answer different questions.
There is an honest boundary on what divergence catches. An attack that produces no divergence, one that keeps the agent perfectly on its assured trajectory, is not caught by watching the trajectory. What bounds that case is the envelope itself. Authority was already attenuated to what the task required, so the effect an on-trajectory action can land is capped no matter whose goal it serves. The security layer catches what it can name. The substrate bounds the effect of what it cannot.
The record either way.
Whether the attack was named or not, the run seals a record of it. The divergence that was observed. The authority that was applied to each action as a result. The point in the sequence where behavior left the envelope, and what the governor did about it. That record is signed onto a tamper-evident chain at the moment each action runs, so it cannot be shaped after the incident by anyone, including the party under review.
SR 26-2 took effect on April 17, 2026 and placed agentic AI outside its scope, which leaves each institution to define the record its own systems keep. After an agent incident, the board asks what happened, the regulator asks what governed it, and the carrier asks whether the account can be replayed. The evidence they read cannot depend on whether anyone had named the attack in advance. It has to hold for the attack that had no signature yet.
What we are building.
Wayfinder Systems Group sits between the agent and the record. It does not scan inputs for known attacks, and it does not replace the tools that do. It reads each action the agent proposes against the trajectory the run has actually walked, sets the authority of that action against the divergence it measures, and signs the value onto a tamper-evident chain before the action lands. The attack does not need a name for the effect to be bounded and the record to be sealed. Patents held in The Wayfinder Trust. We call her Velma.
Thirty minutes. Architecture, not sales.
A conversation about which of your agent controls have to recognize an attack to stop it, and what your incident record holds for the technique that was not named yet.
JonathanLuethke@WayfinderSystemsGroup.com
