The AI Evaluation-to-Reliance Gap
A test can tell you what happened under test conditions. It does not, by itself, justify granting this system this authority here.
18 September 2026 ยท 7 min read
Organisations are getting better at testing AI systems. That is necessary. It is not the same as deciding whether the available evidence is sufficient to grant consequential authority in a live setting.
The gap is simple: evaluation produces evidence about a test. Reliance is the organisational decision that follows. Permission, a passing score and a completed action can all exist while the reliance basis remains unresolved.
What changed
Source finding
Stanford HAI's 2026 AI Index reports that agent task success on OSWorld rose from about 12% to about 66% in a year, while agents still fail roughly one in three structured attempts. McKinsey describes agency as a transfer of decision rights, not a feature. Reuters reported on 16 September 2026 that OpenAI will publish regular reports on unexpected or unauthorised behaviour after increased scrutiny of agent testing.
What the evidence supports, and what it does not
Solarascope analysis
The public record supports this much: capability is moving quickly, agents fail often enough that authority still matters, and vendor disclosure after an event is not the same as a prior organisational decision. It does not establish that any named organisation is ready to grant a specific agent a specific authority. It does not certify a system, approve a deployment or replace identity, IAM, security controls or independent evaluation.
The framework
Test. Evidence. Reliance context. Authority decision. Monitoring. Re-examination.
Evaluators test. Standards set expectations. Security systems, identity, IAM and runtime controls produce constraints and telemetry. Solarascope examines whether those inputs justify this authority, for this use, under these conditions, with a named owner and a way to revoke or reopen the conclusion.
What an organisation should re-examine
If an agent can already act, ask whether the original evidence still maps to the authority now in force. If a model, tool, permission, owner or incident has moved, the earlier conclusion is a candidate for Decision Coverage, not a permanent licence.
Sources
- Stanford HAI. AI Index Report 2026: Technical performance (April 2026)
- McKinsey. Trust in the age of agents (2026)
- Reuters. OpenAI plans regular reports on unexpected AI behavior (16 September 2026)