Measuring accountability: the benchmark gap in agentic evaluation
Agent benchmarks measure whether a task was completed. Enterprises need to know whether it should have been attempted, by whom, and at what cost.
By Dipp AI Research — The editorial desk behind the Dipp AI record
Current agentic evaluation optimises for task success on synthetic environments. It says nothing about attribution, refusal correctness, boundary integrity or cost per outcome. This piece proposes four accountability measures an enterprise can actually run against its own traffic.
Agentic benchmarks have improved quickly and measure the wrong thing for enterprise purposes. They ask whether an agent completed a task in a constructed environment. An enterprise, facing an auditor, asks a different set of questions: was this action within someone's authority, did the system correctly refuse the ones that were not, did any payload cross a boundary it should not have, and what did the completed work actually cost.
Those four questions are measurable, and unlike task-success benchmarks they can be measured on live production traffic rather than on a synthetic suite. What they require is a substrate that records each action with its authority, routing decision, boundary decision and realised cost — which is precisely what a governed execution path produces as a by-product.
FIG. A01-01
Adoption has outrun accountability
Select a measure
78%
Run AI agents in production
Agents are acting on live systems of record, not sandboxes. Adoption is effectively universal across the enterprise estate.
Dipp AI · Orcher
Fig. 1 — Deployment against traceability. The measurement gap in one view.
Four accountability measures
Measure
Definition
Failure it exposes
Attribution completeness
Share of actions resolvable to a named authorising role
Autonomy running outside any grant
Refusal correctness
Share of refusals that were warranted, and of permitted actions that should have been refused
Policy too loose, or too blunt
Boundary integrity
Crossings permitted by policy versus crossings observed
Data leaving via an ungoverned path
Cost per completed outcome
Realised spend attributed to finished work
Defaulting to the top tier by habit
Table 1 — Measures an enterprise can compute on its own traffic.
Why task success is a weak proxy
A high task-success score is compatible with an agent that succeeded at something it had no authority to attempt, using data it should not have seen, on the most expensive tier available. In regulated settings, that is not a good outcome with a caveat; it is an incident that happened to produce the right answer. Evaluation that cannot distinguish the two is not measuring enterprise risk.
FIG. A03-01
The Verified Execution Cycle
Tap a stage
Directive
A named professional states intent. Nothing runs anonymously.
Dipp AI · Orcher
Fig. 2 — The record required to compute all four measures: one sealed entry per action.
Refusal correctness deserves particular attention
It is the only measure that penalises both directions of error. A system that refuses nothing is not enforcing; a system that refuses constantly is mis-scoped and will be routed around by frustrated teams, which is worse than either failure alone. Tracking refusal reasons over time turns policy design into an empirical practice: clusters show where a grant was drawn too tightly, and gaps show where it was never drawn at all.
Compute them on production traffic, not on a curated suite, and state the population.
Treat refusal as a measured outcome rather than an error to be minimised.
Report cost per completed outcome, not cost per token or per call.
“An agent that completes every task and can attribute none of them has not passed an evaluation. It has failed a different one.”
Sources
Sources for every figure in this article.
Where a number comes from Dipp AI's own analysis or an observed deployment, it is labelled as such and is not presented as an independently audited third-party finding.
Dipp AI Technologies, The Enterprise Superintelligence Report, Vol. I (August 2026)
Dipp AI Research, agentic evaluation survey, Q3 2026
Case studies
Organisations that ran this argument in production.
Modelled reference scenarios with the measurement window, the components enforced and the numbers attached. Each one downloads as a PDF.
The desk that edits, sources and dates every piece Dipp AI publishes, and holds the line on what may be claimed. Every figure in this piece carries a source, and corrections are published on the record rather than made quietly.