malav shah
a thesis

The Accountability Layer

Capability stopped being the bottleneck. Being able to prove what an autonomous system did, and that it was allowed to, is the bottleneck now.

The most expensive number in enterprise AI right now is not the cost of inference. It is the number of pilots that will never reach production, and the reason is almost never that the model was not good enough.

The structural shift

For three years the constraint on deployment was capability. Everyone optimised against it, because it was legible and it moved. That constraint has quietly stopped binding for most enterprise workloads. What is binding instead is that a regulated institution cannot deploy a system whose actions it cannot reconstruct, attribute, and defend.

This is not a compliance checkbox bolted on at the end. It is a different question than the one the industry has been answering. “Did the model behave well?” is an evaluation problem. “Can you prove, eighteen months later, to a regulator or a court, what this system did, on whose authority, under which policy version, and what a human would have done instead?” is an infrastructure problem, and almost nothing in the current stack answers it.

The moment agents move from drafting text to moving money, the second question becomes the only question.

The wedge

Everyone is building guardrails, which prevent the bad action. Very few are building the record, which survives the bad action.

Guardrails are a feature. Every model provider will ship them, and they will be free within two years, because they are a product-safety cost the platforms must absorb. The record is different: it is adversarial, it must be tamper-evident, it must outlive the vendor that produced it, and it is worthless unless a third party trusts it. That combination has historically produced durable infrastructure businesses, and it does not get absorbed into a model API, because the buyer specifically needs it to not be controlled by the party being audited.

That is the whole wedge, in one sentence: an audit trail owned by the model vendor is not an audit trail.

The framework

Four things an institution must be able to produce, in order of how hard they are:

  1. What happened. The action log. Easy, and mostly solved.
  2. Under whose authority. The permission chain from a human principal to the agent that acted. Harder, and where most systems fall apart, because the delegation was implicit.
  3. Under which policy. The exact version of the rules in force at the moment of the action, not the current version. Almost nobody stores this.
  4. What the counterfactual was. What the system considered and rejected. This is the one regulators will eventually ask for, and the one that is impossible to reconstruct retroactively, which means it has to be captured at write time or never.

A vendor that only does the first is selling logging. Point four is where the category is going and it has to be designed in from the start.

My position

Accountability infrastructure for autonomous systems is a picks-and-shovels business with regulatory tailwind and an unusually defensible position: it is one of the few AI categories where being independent of the model vendors is a feature the buyer will pay for rather than a disadvantage.

The risk is timing, not direction. Regulation arriving slower than expected is the thing that kills companies here, and I do not think that risk is small.

I am building in this space. Disclosure, not hedge.

The prediction

The first significant enforcement action against an institution for an autonomous agent’s decision, rather than against the model provider, resets the entire procurement conversation. Everything before that is a pilot budget. Everything after is a line item.

The signal I actually track is duller and earlier: pilot-to-production time in regulated buyers. When that number starts falling, the bottleneck has moved, and the market is real.


← all theses