News · September 14, 2026
Introducing the Verified Execution Cycle: A New Way to Meter, Bill, and Audit AI
The industry has spent 2026 trying to bill AI agents for outcomes instead of tokens. Almost none of it has solved the harder problem underneath: proving an outcome actually happened before anyone gets billed for it.
By Odero Otieno — Founder, CEO & CTO, Dipp AI Technologies, Inc.
The Verified Execution Cycle is the unit Orcher meters, bills, and audits: one directive, one bound role, gates in sequence, one committed action, and a permanent record. Nothing partial counts, and a cycle that halts produces no charge and no ambiguity about what happened.
Today we are introducing the Verified Execution Cycle, or VEC — the unit Orcher meters, bills, and audits. A VEC is one directive carried through Orcher's full control plane, start to finish: a named professional's intent, a bound role, gates that verify authority, action, and data boundary in sequence, a committed action, and a permanent record. Nothing partial counts. A cycle that halts at any gate produces no charge and no ambiguity about what happened, because the halt itself is recorded with the same rigor as a completed action.
Orcher closes the accountability gap enterprises face when they cannot prove what an agent did, and the structural failure of reviewing actions after the fact, through seven components across two layers. This piece is about what all of that produces when it runs once, on one directive — and what makes that unit different from how the rest of the industry is currently trying to price agentic work.
FIG. A03-01
The Verified Execution Cycle
Tap a stage
Directive
A named professional states intent. Nothing runs anonymously.
The industry's actual problem is not pricing. It's proof.
Outcome-based pricing has become the dominant framing for AI agent monetization in 2026. Salesforce Agentforce charges per conversation. Intercom's Fin lists from $0.99 per outcome. Zendesk charges only when an AI agent resolves a ticket without human intervention and the resolution passes verification. The logic is sound: when an agent acts on a customer's behalf, the customer wants to pay for the result, not for the compute it took to get there.
The trade-off nearly every guide to this shift names honestly is attribution. Outcome pricing requires defining what success means, measuring it reliably, and handling disputes when an agent partially succeeds. A widely cited 2026 pricing guide states the problem plainly: "completed" is not always binary, and tool attribution gets genuinely difficult once a single customer-visible outcome was actually produced by search, storage, OCR, email, CRM, and model calls working together underneath. Deloitte's June 2026 technical accounting guidance for outcome-based agentic contracts shows how granular this gets: $1.25 per invoice processed straight through, $1 per duplicate payment prevented, $3 per successful autonomous fraud intervention — each requiring an auditable definition of success written into the contract itself.
“The unit to bill on is usually not tokens. It is a protected unit of work the buyer already understands.”
That is the correct instinct, and it is also exactly where most outcome-based pricing quietly runs out of architecture. A "protected unit of work" is only protected if there is a reliable way to prove it happened as claimed, independent of the vendor's self-reported resolution count. The specialist customer-service tier shows what happens without that proof: one widely used platform counts an "assumed resolution" whenever a customer simply stops responding — a definition that also captures customers who gave up and went elsewhere. That is not a pricing model failure. It is an evidence failure wearing a pricing model's clothes.
A 2026 metered-billing guide states the consequence in one concrete image: a customer accepts variable pricing when the value is obvious, and panics when one vague "agent run" turns into a $3,400 invoice because the system quietly retried a failing tool call for six hours before anyone noticed. Nothing about that invoice was fraudulent. It was simply unverifiable at the moment it mattered.
single invoice produced by a six-hour retry loop nobody could see
2026 metered billing guide
of AI companies that started with usage-based pricing have already changed the model
at least once
per invoice processed straight through, in a real outcome-based agentic contract
Deloitte, June 2026
charge produced by a cycle that halts at any Orcher gate
nothing partial is billed
What a VEC actually verifies, gate by gate
Select a component to reveal its Vol. I excerpt.
Fig. CP-01 · The seven-component control plane
A Verified Execution Cycle does not ask an enterprise to trust a resolution count after the fact. It produces the proof as a structural byproduct of the gates every directive passes through before it can be metered at all.
- Role Identity Fabric — is this person, in this role, entitled to ask for this? Authority and context are captured before anything runs, and anything outside the role halts here, before a token is spent.
- Logic Scrubber — does the proposed action hold up against the systems of record? Eligibility, policy, licensure, and limits are checked before anything is written, not reconciled afterward.
- Data Control Gateway — may this payload cross the boundary, to this class of model, in this region, under these training-exclusion terms? Verified on this call, not at onboarding.
- Immutable Audit Ledger — the committed action and the full decision trace are hashed permanently, replayable months later as evidence rather than as logs.
The consequence for billing is unusual and deliberate: the invoice line and the audit record are the same object. An enterprise does not reconcile what it was charged against what it can prove, because the charge exists only where the proof does.
Why halting must be free
If a halted cycle carried a charge, every gate would acquire a quiet commercial incentive to pass. That is the failure mode we designed against first. A halt is recorded with the same rigor as a completion — who asked, under what role, which gate refused, and why — and it costs nothing. The enterprise gets the evidence without paying for the non-outcome, and Orcher gets no economic reason to loosen a check.
Other companies building AI billing infrastructure in 2026 have converged on part of the same answer independently. One metering platform describes its approach as zero-trust billing: every usage record pushed to an append-only, tamper-proof log at the moment it is created, with the pricing rule stamped onto each credit so an auditor can verify the total against the line items. That is a genuinely different design goal from most usage-based billing, which tracks consumption accurately without making it independently verifiable. A VEC treats the pricing question and the proof question as the same problem from the start.
The market is already moving toward a unit-of-work model. Most of it still can't verify the unit.
Cost Governance
Routing by stakes, enforced — not recommended
Pick a stakes class
Small / open class
1× relative cost · permitted
Classification, extraction, summarisation, routine drafting.
Mid class
halted by ceiling
Multi-step reasoning where an error is recoverable and reviewable.
Frontier class
halted by ceiling
Irreversible, high-stakes, or externally binding work only.
A directive with no external effect and a recoverable error path is capped at the small class. The frontier is not an option the router can pick.
Kyle Poyar's study of more than 240 software companies found hybrid pricing — a base fee plus consumption — rose from 27% to 41% of vendors in a single year. Salesforce has gone granular with Flex Credits, where individual Agentforce actions draw from a metered credit pool. Microsoft shipped an Agent Governance Toolkit on 2 April 2026: open-source runtime security covering the OWASP Agentic Top 10 and EU AI Act requirements. Every one of these moves treats the unit of work as the thing worth metering precisely, which is the same instinct behind a VEC. None of them, on current public documentation, binds that metered unit to the same cryptographic chain of role, verification, and data boundary that produced it in the first place.
of 240+ software companies moved to hybrid usage-plus-consumption pricing in one year
Kyle Poyar
Intercom Fin's per-outcome starting price, defined by the vendor's own resolution criteria
record a VEC produces for both the billing unit and the audit trail
not two systems reconciled later
A VEC is also an answer to a question academic benchmarking is asking right now
Everything so far treats the Verified Execution Cycle as a billing unit. It is also, by the same construction, an evaluation unit — and the research field measuring agent trustworthiness spent 2026 arriving independently at the distinction a VEC already enforces: whether an agent completed a task is a different, weaker question than whether the trajectory that produced the completion is trustworthy and provable.
“In 2026, public benchmarks are saturating and gameable… task success vs. trajectory accuracy.”
Academic work has moved the same direction with more formal machinery. DEMM-Bench, a cross-regime benchmark for agent-runtime governance-evidence sufficiency, tests directly whether an agent's own runtime evidence is sufficient to satisfy a given regulatory regime. A companion determinism-faithfulness assurance harness for tool-using agents evaluates replayability across 74 configurations and 12 models. A third framework, AEMA, proposes process-aware, auditable evaluation for multi-agent systems, built on the same premise: a result without an auditable process behind it is not evidence of anything.
None of these efforts were built by Dipp AI, and none were built with Orcher in mind. That independence is the point. A well-resourced enterprise automation vendor and at least three separate 2026 academic papers each concluded, on their own, that the field needs a way to measure whether an agent's trajectory is provably faithful — not just whether its output looked correct. A VEC does not evaluate that property after the fact, the way a benchmark applied to a completed transcript does. It enforces the property structurally, at the moment of execution, because the gates that verify role, action, and data boundary are the same gates that decide whether a cycle completes at all. There is no ungoverned trajectory for a benchmark to go looking for evidence in, because an ungoverned trajectory never became a completed cycle.
A related risk this research makes concrete is worth naming directly: evaluation itself can be gamed. A 2026 study of retrieval-augmented systems found near-perfect benchmark scores when evaluation elements were leaked or predictable — closely related to the self-reported "assumed resolution" problem already documented in the specialist customer-service tier above. A benchmark that scores a transcript after the fact can be optimised against, deliberately or not. A cycle that has to clear its gates before it exists cannot be gamed the same way, because there is no later transcript being scored; there is only a record that either shows the gates cleared or shows the cycle halted.
What the unit gives an enterprise beyond an invoice
Dipp Intelligence
What compounding actually looks like
Step through the cycles
A claims adjuster's directive is verified against policy language and precedent. The outcome, approved with specific reasoning, becomes a Dipp Intelligence entry.
Because a VEC is a whole unit rather than a stream of tokens, it becomes the natural denominator for everything else an enterprise wants to know. Cost per verified outcome, not cost per million tokens. Utilisation measured against cycles that actually cleared, not against raw request volume — the basis of Elastic Compute Governance. And every completed cycle writes an entry into Dipp Intelligence, the permanently owned record of what an organisation's own people authorised and why.
Sources
Sources for every figure in this article.
Where a number comes from Dipp AI's own analysis or an observed deployment, it is labelled as such and is not presented as an independently audited third-party finding.
- 2026 industry guides to outcome-based AI agent pricing, covering Salesforce Agentforce, Intercom Fin, and Zendesk resolution-based billing.
- Deloitte, technical accounting guidance for outcome-based agentic contracts, June 2026.
- "Metered Billing for AI Agents: 2026 Guide," on protected units of work and retry-loop invoice exposure.
- 2026 analysis of zero-trust billing architecture and append-only usage records.
- Kyle Poyar, study of 240+ software companies on hybrid pricing adoption, cited in Nevermined, "AI Agent Billing in 2026: Patterns & Playbooks," May 8, 2026.
- Salesforce Flex Credits documentation, 2026; Microsoft Agent Governance Toolkit, released April 2, 2026.
- Automation Anywhere, "AI Agent Benchmarks: The 2026 Enterprise Evaluation Guide," on the shift from task success to trajectory accuracy.
- "DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency" and the companion "Decision Evidence Maturity Model for Agentic AI," 2026.
- "Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents," 2026.
- "AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems," 2026.
- "Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?," 2026, on benchmark gaming through metric overfitting.
- The Enterprise Superintelligence Report, Vol. I, Dipp AI Technologies, August 2026.
Written by Odero Otieno.
Writes the Dipp AI record on enforced governance for agentic systems — authority, enterprise data boundary, cost and compute. Every figure in this piece carries a source, and corrections are published on the record rather than made quietly.
