Research · August 28, 2026

Measuring accountability: the benchmark gap in agentic evaluation

Current benchmarks score capability. Enterprises are liable for authority, evidence, and cost.

Summary

A technical note on the gap between how agentic systems are evaluated and what enterprises are answerable for: capability benchmarks say little about who authorised an action, what evidence supports it, or what it cost to produce.

Full text

The record

Benchmarks measure whether a model can. Enterprises are answerable for whether an action should have been taken, by whom, and at what cost.

What current evaluation covers

Accuracy, reasoning, tool use, task completion. All useful, all about capability, all silent on accountability.

What it omits

Four things an enterprise is asked about after an incident: which named human authority the action was executed under; what evidence existed at the moment of commit; whether the action was consistent with the systems of record it claimed to follow; and what it cost to produce relative to the budget that governed it.

Toward an accountability evaluation

An honest evaluation would score authority binding, evidence completeness, verification against sources of record, boundary compliance, and cost discipline — alongside capability, not instead of it. These are properties of the control plane, not of the model, which is precisely why capability leaderboards cannot report them.

Position

Dipp AI publishes the measures it applies to Orcher and welcomes correction in writing. Where a figure we publish is wrong, we want the citation.

Corrections welcome, in writing.

If a figure we published is wrong, we want the citation. Research correspondence goes straight to the founding team.