Before reliance: what must be true before consequential authority
A practical test for any system you are about to grant room to act: objective, evidence, boundary, owner, conditions.
12 September 2026 ยท 8 min read
Most teams ask whether a system is accurate enough. That is the wrong first question. The first question is: what would have to be true for us to let it act, and what would make that permission expire?
Accuracy is a property of a model or a vendor test. Authority is a property of a decision. Reliance infrastructure starts from the decision: the objective, the evidence a reasonable operator would require, the boundary that must not be crossed, the named owner of residual risk, and the conditions under which the conclusion remains valid.
Five conditions before authority
First, the objective is explicit. 'Resolve customer inquiries' is not an objective. 'Resolve billing inquiries under $500 without creating a regulatory exception' is. Vague goals produce unbounded systems.
Second, evidence is current and specific. A pilot from March does not license a September deployment. Name the tests, dates, populations and contrary results.
Third, authority is bounded in a way that can be checked. Refunds up to $500 is a boundary. 'Use judgement' is not.
Fourth, a named human owns the residual risk. A committee cannot suspend a system at 2 a.m. A role with a name can.
Fifth, revisit triggers are armed. A model swap, a vendor change, a spike in overrides or a policy edit should move the record from Ready to Conditional until the evidence is re-tested.
What the public evidence is already showing
Source finding
Stanford HAI's 2026 AI Index reports that agent task success on OSWorld rose from about 12% to about 66% in a year, while agents still fail roughly one in three structured attempts. McKinsey's work on trust in the age of agents describes agency as a transfer of decision rights, not a feature.
Solarascope analysis
Solarascope analysis
Capability moving quickly is not an argument for skipping the five conditions. It is an argument for writing them down. A system that fails one in three structured attempts can still be used, if the boundary, owner and fallback are real. It cannot be used on the basis of a demo that went well.
Sources
- McKinsey. Trust in the age of agents (2026)
- Stanford HAI. AI Index Report 2026: Technical performance (April 2026)