Orcher · Component 06 of 07

Cost Governance

Routes by stakes and halts runaway loops.

Layer 2 — continuous, non-blocking

A routing manifold sorting requests by stakes, with a circuit breaker halting a runaway loop
Routes by stakes and halts runaway loops.

Not every directive deserves the frontier. Cost Governance routes by the stakes of the action rather than the habits of the developer, enforces budgets at the role and department level, and terminates loops that would otherwise spend without producing an outcome. Cost is a governance property, not a finance report written after the fact.

  • Selects the model tier from the consequence of the action, not the default.
  • Enforces budget ceilings per role, department, and directive class.
  • Detects and halts non-terminating agent loops.
  • Escalates to a higher tier only when verification demands it.

Without it, spend outruns value

Ungoverned agent fleets exhaust annual budgets in a quarter and produce no auditable outcome to show for it.

What it emits

A routing decision and a cost record per call, reconciled to the directive.

Evidence

Cost Governance in the 2026 enterprise data#

Figures drawn from the Enterprise Superintelligence Report, Vol. I, August 2026.

40–85%

spend reduction observed under Orcher cost governance

range across deployments — Vol. I

4,500×

price spread across 400+ models and 70+ providers

Vol. I, p.25

63%

cost reduction at 91% accuracy match in internal routing tests

Dipp AI Research: internal analysis, not independently verified

68%

of enterprise AI programs are over budget

Vol. I, p.9

Mechanism

How it works#

Four stages, in order. Layer 1 components gate execution; Layer 2 components run continuously and never block.

  1. 01

    Score the stakes

    Each directive is scored by consequence: what changes in a system of record, whose authority is exercised, and what a wrong outcome would cost. Stakes, not developer habit, set the tier.

  2. 02

    Route to a tier

    Across 400+ models and 70+ providers the price spread is roughly 4,500×. Orcher selects from that range deliberately, and escalates only where verification demands the stronger model.

  3. 03

    Enforce ceilings

    Budgets are enforced per role, per department, and per directive class at call time — not reported after the quarter closes.

  4. 04

    Detect and halt loops

    Non-terminating agent loops are detected and stopped, and the halt is written to the ledger like any other outcome rather than disappearing into a bill.

Operational contract

Input
Directive stakes score, role budget, and live provider pricing
Output
Routing decision and per-call cost record reconciled to the directive
Mode
Layer 2 — continuous, non-blocking
Controls
Tier selection, escalation policy, budget ceilings, loop termination
Attribution
Cost is attributed to a Verified Execution Cycle, not to an API key
Observed effect
40–85% spend reduction across deployments

What it is not

It is not a FinOps dashboard

Dashboards describe spend that already happened. Cost Governance decides, at call time, whether the spend is permitted at all.

It is not a race to the cheapest model

High-stakes directives are routed to the strongest available tier on purpose. The discipline is matching the tier to the consequence, in both directions.

Failure behaviour

When a ceiling is reached or a loop is detected, execution halts and the directive is recorded as halted with its cost trace. Spend never continues silently past a limit.

Questions

What enterprises ask first

How do you decide what a directive is worth?
From the action it proposes: the system of record it writes to, the value at risk, the role exercising authority, and the regulatory exposure. Enterprises tune the scoring; Orcher enforces it.
Where does the 40–85% range come from?
Observed spend reduction across deployments. Internal routing tests matched 91% of frontier-tier accuracy at 63% lower cost; that figure is Dipp AI research and is not independently verified.
Do reasoning models break the model?
They make it more necessary. Reasoning tokens multiply effective cost per call, which is why the tier decision has to be governed rather than defaulted.
Can a team be given its own budget?
Yes. Ceilings are set per role, department, and directive class, and they are enforced at call time rather than reconciled later.

Where it shows up

Solutions and industries that depend on this

Surfaced automatically from the components each solution engages and each industry relies on.

The strategy behind the component

Dynamic model routing.#

Binding a platform to one heavy frontier model wastes tokens and leaks proprietary data. Dipp AI's strategy — and the reason Orcher's seven components exist — is intelligent routing: decouple orchestration from any single provider, match each task to the cheapest model that can carry it, and keep prompts and data inside the enterprise's own tenant boundary.

DIPP AI / ORCHER

Seven components. One enterprise-owned control plane.

01 + 02 / AUTHORITY

A role-bound directive

Intent, named professional, permitted actions and systems of record.

06 / COST GOVERNANCE

Dynamic model routing

Select the least costly sufficient model within the task’s stakes, latency target and spend ceiling.

Explore Cost Governance

05 / DATA CONTROL GATEWAY

Boundary before egress

Region, redaction, retention and training-exclusion requirements determine which destinations qualify.

Frontier reasoning

  • GPT-6 Astra
  • Claude Fable 5.1
  • Claude Mythos 5.1
  • Muse Spark 1.3

Models discussed in Dipp AI Research · September 6, 2026

Open-weight & private

  • Llama
  • Qwen
  • DeepSeek
  • Mistral
  • Gemma
  • Enterprise-tuned models

Model families · exact versions and licences assessed per deployment

Compute destinations

  • AWS
  • Microsoft Azure
  • Google Cloud
  • CoreWeave
  • Lambda
  • Crusoe
  • Private infrastructure

Hyperscalers, neoclouds and enterprise infrastructure

03 / Logic Scrubber

Verify proposed actions before commit.

04 / Immutable Audit Ledger

Preserve the decision and refusal evidence.

07 / Observability

Trace cost, latency and outcomes across routes.

Illustrative routing architecture, not a live connection status or an exhaustive integration catalogue. Provider availability, model access and deployment readiness are confirmed during a technical briefing.
01

Decoupled orchestration

Switch providers in a week, not a quarter.

An orchestration layer sits in front of every AI task rather than inside one vendor's SDK. Directives, roles, policy and evidence live in Orcher, so a provider change is a routing-table change — a week of work rather than a quarter of re-platforming — and no model vendor becomes the system of record for how the enterprise operates.

The protocol stack
02

Task-to-model matching

Roughly 90% of workloads never need the frontier.

Cost Governance routes by the stakes of the action rather than the habits of the developer. Standard, routine queries — roughly nine in ten workloads — go to cheaper open-weight or small in-house models; expensive frontier models are reserved for genuinely complex, high-reasoning tasks. Across 400-plus models and 70-plus providers the price spread is 4,500× (Vol. I, p.25), so the routing decision is the economics.

Cost Governance
03

Private tenant boundaries

Providers never learn from your interaction data.

The Data Control Gateway forces data and prompt histories to stay inside the company's own cloud environment — a private tenant, in-region — and verifies training exclusion and the data boundary on every call before a token leaves. Using frontier capability never means donating enterprise interaction data to the firm that sells it back.

Data Control Gateway

Runtime workflow

Authority per directive.#

Routing is not a developer preference resolved in a config file. For every directive, Orcher resolves who is accountable, selects the cheapest model class that authority permits, carries the data boundary along the chosen route, and hashes the whole decision into one evidence chain that reads the same across every provider.

  1. 01

    Authority is resolved before a model is chosen

    Who is accountable, and what may they authorize?

    The directive arrives with a named issuer. The Role Identity Fabric resolves that professional's current entitlements and binds them to this cycle, so the routing decision is made against a known authority rather than a service account. Anything outside the role halts here — before a single token is spent.

    Enforced by

    Emits
    A role-bound authorization token scoped to this directive.

  2. 02

    The routing decision is a policy decision

    What is the cheapest model class this directive's stakes permit?

    Cost Governance classifies the workload by stakes, sensitivity and reasoning depth, then selects the cheapest sufficient model class within the ceiling in force for that role. Roughly nine in ten workloads clear on routine classes; frontier capacity is reserved for complex work. The chosen class, the alternatives considered and the ceiling applied are all recorded as part of the decision.

    Enforced by

    Emits
    A signed routing decision: class chosen, ceiling applied, rationale.

  3. 03

    The data boundary travels with the route

    May this payload reach that provider, in that region?

    The Data Control Gateway applies data-boundary, redaction and training-exclusion rules to the route that was chosen, not to a default path. If a class is otherwise optimal but its provider cannot satisfy the boundary, the route is refused and the next sufficient class is selected. Consistency across providers is enforced by the gateway, not negotiated per vendor SDK.

    Enforced by

    Emits
    A per-call boundary attestation: region, redactions, no-training terms.

  4. 04

    Every decision lands in one evidence chain

    Can this be replayed and defended months later?

    The Immutable Audit Ledger hashes the directive, the authority, the routing decision, the boundary attestation and the committed action into a single tamper-evident chain. Observability streams the same run across whichever providers were involved, so one query answers what was asked, who authorized it, where it ran and what it cost — regardless of vendor.

    Enforced by

    Emits
    A replayable, hashed execution record with cross-provider cost and latency.

Interactive · routing simulator

Enter a workload. See where it routes.#

Each workload type carries a different level of stakes, sensitivity and reasoning depth. Select one to see which Orcher component owns the decision, which model class it routes to, which gates run, and what that does to tokens and cost against a frontier-default baseline.

Workload type

Routing decision

Document classification and routing

High-volume intake: label the document, extract identifiers, route it to a queue. No irreversible write, no judgement call.

Decided by

Cost Governance

Routed to

In-house small model

Runs inside the tenant. No egress, lowest unit cost, deterministic latency.

Token reduction

42%

1,800 → 1,044 tokens

Cost per directive

$0.00013

baseline $0.0432 on frontier model

Cost saving

100%

2360ms faster median

Gates that run before commit

  1. Role scope check
  2. Cheapest-sufficient class
  3. In-tenant execution
  4. Hashed record

Stays inside the tenant entirely — the payload never reaches an external provider.

Indicative class prices per 1M tokens: In-house small model $0.12 · Open-weight mid model $0.60 · Commercial mid-tier $4.50 · Frontier model $24.00. Illustrative only; the 4,500× spread across 400+ models and 70+ providers is from The Enterprise Superintelligence Report, Vol. I, p.25.

Interactive · economics

What routing does to the bill.#

A frontier-default estate pays the top price band for every workload, including the nine in ten that never needed it. Move the inputs to your own volumes and see the token waste removed and the cost impact of the routine-versus-frontier split.

Your estate

250,000

One directive is one verified execution cycle.

5,000

Prompt plus completion, before scoping and redaction.

90%

Dipp AI's routing model puts this at roughly 90%; the remainder is frontier-only.

32%

Directive scoping, field redaction and cache reuse before a model is called.

$24

Your blended frontier rate.

$0.60

In-house or open-weight class hosted in your tenant.

Estimated impact

Monthly cost, frontier default

$30.0K

1.25B tokens at the top band

Monthly cost with routing

$2.5K

$459 routine · $2.0K frontier

Token waste removed

400.0M

32% of 1.25B never reaches a model

Cost reduction

92%

$27.5K per month

Annualised

$330.0K

Retained by routing 90% of directives to a routine class and reserving the frontier for the 10% that genuinely needs it — with every routing decision hashed into the same audit chain as the action it paid for.

Estimates only, for planning conversations rather than contract terms. The 4,500× price spread across 400+ models and 70+ providers, and the 68% of AI programmes running over budget, are from The Enterprise Superintelligence Report, Vol. I (pp. 25, 9).

Proof points

The strategy, in measurable terms.#

Dynamic model routing is only a strategy if it moves numbers a finance team and an auditor both recognise. These are the outcomes it is accountable to.

Cost

4,500×

price spread across 400+ models and 70+ providers

The spread between the cheapest sufficient model and the frontier default is the entire economic case for routing. A directive sent to the wrong class does not fail — it just costs orders of magnitude more than it had to.

Vol. I, p.25

Token waste

~90%

of workloads never need a frontier model

Routine classification, extraction, summarisation and drafting clear on in-house or open-weight classes. Reserving frontier capacity for genuinely complex work is what turns AI spend from a run-rate into a budget.

Dipp AI routing model · Vol. I, p.25

Budget

68%

of enterprise AI programmes run over budget

Overrun is a control failure, not a forecasting failure. Cost Governance applies a ceiling per role and per directive and halts loops at the step boundary rather than at the invoice.

Vol. I, p.9

Latency

10×

faster on routine classes than a frontier default

Cheaper classes are also materially faster. Routing the routine nine-tenths off the frontier shortens the median directive as a side effect of the economics.

Indicative class latencies, Orcher routing table

Auditability

28%

of enterprises can trace an agent action end to end today

Every routing decision Orcher makes — class chosen, alternatives considered, ceiling applied, boundary attested — is hashed into the same evidence chain as the committed action, so the economics are as auditable as the outcome.

Vol. I, p.6

Portability

1 week

to change providers, not a quarter of re-platforming

Because orchestration is decoupled from any vendor SDK, a provider change is a routing-table change. Directives, roles, policy and evidence stay in Orcher.

Dipp AI protocol stack contract

The other six

Orcher is one control plane

Further reading · September 2026

The argument behind this page, published in full.

Two long-form pieces per week through September 2026 — sourced, dated, and attributed. Each one links back to the components it argues about.

Verified execution, or none at all.

Orcher is deployed with named enterprises under the Human-in-the-Role model. Request a technical briefing with the founding team.