A governor sets how much authority the next action carries, and the amount it grants changes what the system does next. That is the point of putting it there. It is also the reason the usual validation move does not survive contact with it.
Take the log of a run that already happened. Stand the candidate control over it. Count the actions it would have stopped. That procedure is sound for a component that watches. It breaks on a component that acts.
The trace was produced by a system that was never bounded.
Every historical agent log is a record of a run where nothing intervened. The agent proposed, the action executed, the next state followed from it, and the agent proposed again from there. The sequence is causally closed on itself.
Put a governor on that path and the first reduced grant breaks the chain. The action that executes is narrower than the one proposed. The state that follows is not the state in the log. The agent's next proposal is made from a position it never occupied in the recorded run.
So a replay is faithful up to the first intervention and fiction after it. Everything past the first grant the governor lowers is a record of a system that no longer exists. The longer the run, the smaller the fraction of the log that still means anything, and agent runs are long by construction.
What a shadow-mode number is worth.
The standard answer is to run the governor beside the path and discard its grants. It decides, nothing enforces, and the decisions get scored against what actually happened.
That answers one narrow question. At the moment of a known bad action, did this control lower authority. It is worth knowing and it is not the deployment question. The deployment question is what the run looks like once the grants land, and a shadow run holds no evidence about that, because in a shadow run the grants never land.
The anticipatory case is where scoring against history goes furthest wrong. A governor can grant less authority on an action whose present trust reads higher, because its forward look caught a divergence before the outcome arrived. At that step nothing has gone wrong yet. The telemetry is clean, the labels say benign, and a scorer comparing the pullback against the recorded outcome marks the correct call an error. The governor is penalized precisely for the judgment that is worth paying for.
A control whose output is its own next input.
A detector is a function. Telemetry goes in, a score comes out, and the world is unchanged by the evaluation. That is what makes a corpus a fair test of it.
A governor is a loop. It reads the run, sets the authority of the next action, and the action that executes under that grant produces the telemetry it reads on the following pass. The object under validation is a system in feedback with the thing it governs, and no fixed corpus stands outside that loop.
What can be established about a loop is the set of properties that hold on every pass. Whether the enforcement point sits on the only path to a durable effect. What the action carries when no grant arrives. Which direction an uncertain forward look rounds. Whether the envelope the grants are measured against can be rewritten from inside the run. Those are read off the construction rather than estimated from a sample, and they hold for actions no corpus contained.
The observable the run does produce.
The counterfactual is unavailable. The governed run is not.
On every action there are two quantities. What the agent proposed, and what the governor granted. The difference between them is the intervention, stated at the moment it happened, in units the action was measured in. Scope requested against scope allowed. Amount proposed against amount permitted. Reach sought against reach granted.
That series is the thing a validation function can actually work with. It shows where authority was withheld, how often, against which class of action, and how the grants moved as the run consumed its envelope. A validator reading it is examining the governor's judgment in the run that was governed, rather than inferring it from a run that was not. The chain the governor signs is what carries that series out of the run intact.
The practice this collides with.
Model validation as it stands is built around outcome analysis and benchmarking. Compare predictions to realized results. Compare the challenger to the champion on the same held-out data. Both procedures rest on an assumption that is never written down, because for a scoring model it is always true. The assessment does not change the outcome being assessed.
SR 26-2 took effect on April 17, 2026 and placed generative and agentic AI outside its scope while pointing institutions at existing model risk practice pending an interagency request for information. The NAIC Model Bulletin, now adopted in more than half the states, asks carriers for evidence that AI systems were tested before deployment and monitored after. Both frameworks are written for components that estimate. Neither has a template for a control that sits in the decision path and changes the sequence it is being judged on.
Independent validation of a governor is a different exercise. Read the construction for the properties that must hold on every action. Read the governed run for the grants it issued and what they cost the agent. A backtest of a bound measures a run the bound would never have allowed to happen.
What we are building.
Wayfinder Systems Group builds a runtime governor. It sits above control and below intelligence, sets how much authority a system is granted on every action against a safety envelope declared before the run, and bounds autonomy so the system acts only inside what it grants, through a path it cannot step around. Its forward look is what lets it pull authority back before an outcome lands rather than after. Every grant it issues is signed onto a tamper-evident chain as it happens. We call her Velma.
Thirty minutes. Architecture, not sales.
A conversation about how a runtime governor gets validated when the historical trace stops describing the system, and what your validation function should be reading instead.
JonathanLuethke@WayfinderSystemsGroup.com
