Research · August 28, 2026
Measuring accountability: the benchmark gap in agentic evaluation
Current benchmarks score capability. Enterprises are liable for authority, evidence, and cost.
Summary
A technical note on the gap between how agentic systems are evaluated and what enterprises are answerable for: capability benchmarks say little about who authorised an action, what evidence supports it, or what it cost to produce.
Full text
The record
Benchmarks measure whether a model can. Enterprises are answerable for whether an action should have been taken, by whom, and at what cost.
What current evaluation covers
Accuracy, reasoning, tool use, task completion. All useful, all about capability, all silent on accountability.
What it omits
Four things an enterprise is asked about after an incident: which named human authority the action was executed under; what evidence existed at the moment of commit; whether the action was consistent with the systems of record it claimed to follow; and what it cost to produce relative to the budget that governed it.
Toward an accountability evaluation
An honest evaluation would score authority binding, evidence completeness, verification against sources of record, boundary compliance, and cost discipline — alongside capability, not instead of it. These are properties of the control plane, not of the model, which is precisely why capability leaderboards cannot report them.
Position
Dipp AI publishes the measures it applies to Orcher and welcomes correction in writing. Where a figure we publish is wrong, we want the citation.
Corrections welcome, in writing.
If a figure we published is wrong, we want the citation. Research correspondence goes straight to the founding team.
