Route by stakes, not by habit. A technical account of the decision that determines cost, compliance and latency in a single step.
By Dipp AI Research — The editorial desk behind the Dipp AI record
Model selection is the highest-leverage decision in an agentic stack, and in most enterprises it is not a decision at all — it is a default. This piece specifies routing as a constrained optimisation across stakes, data class, jurisdiction, latency and an enforced budget envelope.
In an agentic stack, the single decision that most determines cost, compliance posture and user-perceived quality is which model executes a given task. In most enterprises that decision is made once, at integration time, by a developer choosing a default — and then inherited by every task the workflow will ever run, regardless of stakes.
Treating it as a live decision changes the economics and the risk profile at the same time, because the constraints that should govern model choice are exactly the constraints that govern exposure: what the task is worth getting wrong, what class of data it carries, which jurisdictions may process it, how fast the answer is needed, and how much authority remains in the budget envelope.
Multi-step reasoning where an error is recoverable and reviewable.
Frontier class
halted by ceiling
Irreversible, high-stakes, or externally binding work only.
A directive with no external effect and a recoverable error path is capped at the small class. The frontier is not an option the router can pick.
Dipp AI · Orcher
Fig. 1 — Cost spread across tiers for an identical task class. Defaulting is a spending decision made by omission.
The routing decision, stated
Stakes — the cost of an incorrect action, expressed by directive class rather than inferred from prompt length.
Data class — determined before routing, restricting eligible models, regions and execution modes.
Jurisdiction — which legal regime may process the payload, and therefore which providers are eligible at all.
Latency budget — an interactive directive and an overnight batch do not warrant the same tier.
Spend envelope — the remaining authority attached to the role and the period, enforced rather than reported.
Given those inputs, routing selects the cheapest tier that satisfies every constraint. Where no tier satisfies them, the correct behaviour is refusal with the failing constraint named — not silent selection of the closest available option, which converts a policy failure into an invisible exposure.
Directive class
Typical stakes
Routing outcome
Summarise internal thread
Low
Small tier, batch latency, no external boundary crossing
Draft customer correspondence
Medium
Mid tier, redacted payload, reviewed by role on exception
Adjudicate a claim
High
Top tier permitted, strict jurisdiction, ledger-backed
Move funds or amend a record
Critical
Top tier plus explicit standing grant, refusal by default
Table 1 — Stakes drive tier. Tier is not a global default.
Why the ceiling belongs in the same step
Cost control implemented downstream of routing can only report or interrupt. Implemented inside routing, the envelope becomes another constraint in the same evaluation, which means the system can degrade along a path the accountable role has already approved. That is the difference between a cost control that saves money and one that quietly damages outcomes to hit a number nobody agreed to.
Select a component to reveal its Vol. I excerpt.
Fig. CP-01 · The seven-component control plane
Fig. 2 — Routing as a pre-dispatch gate, evaluated with identity, cost and boundary in one pass.
Measuring the router
Cost per outcome
the metric that matters
Not cost per token or per call
Escalation rate
how often a cheaper tier proved insufficient
Unroutable rate
how often policy admitted no compliant model
Escalation rate is the honest test. A router that never escalates is over-provisioning; one that escalates constantly is under-tiering and paying twice. The target is a stable, low escalation rate at a falling cost per completed outcome — and that target only becomes measurable once every action carries its directive class, its tier and its realised cost in the same record.
Dipp Intelligence
What compounding actually looks like
Step through the cycles
A claims adjuster's directive is verified against policy language and precedent. The outcome, approved with specific reasoning, becomes a Dipp Intelligence entry.
Dipp AI · Orcher
Fig. 3 — Routing accuracy improves as the verified record grows: the compounding effect of governed execution.
“Choosing the strongest model for every task is not a quality strategy. It is the absence of one.”
Sources
Sources for every figure in this article.
Where a number comes from Dipp AI's own analysis or an observed deployment, it is labelled as such and is not presented as an independently audited third-party finding.
Dipp AI Technologies, The Enterprise Superintelligence Report, Vol. I (August 2026)
Dipp AI Research, model tier price spread analysis, Q3 2026
Case studies
Organisations that ran this argument in production.
Modelled reference scenarios with the measurement window, the components enforced and the numbers attached. Each one downloads as a PDF.
Every claims directive was going to the most expensive model available. Stakes-based routing moved 82% of them to a cheaper sufficient tier and left the hard ones where they belonged.
The desk that edits, sources and dates every piece Dipp AI publishes, and holds the line on what may be claimed. Every figure in this piece carries a source, and corrections are published on the record rather than made quietly.