Dipp AI Research · September 21, 2026

Cost Governance: Why Recommending a Cheaper Model Isn't Governing Cost

Model routing is not new. Almost every gateway vendor offers it. What is still missing is the difference between a router suggesting the cheaper option and a system that will not let the expensive one run unless the stakes justify it.

By Odero OtienoFounder, CEO & CTO, Dipp AI Technologies, Inc.

IBM finds 87% of organizations claim an AI governance framework and fewer than 25% have implemented the controls it describes. Cost is where that gap is easiest to measure and hardest to argue with. This piece sets out what enforced routing actually means, gate by gate, and traces the same fallacy through Glean's 95% frontier-usage estimate, a $1.3 million OpenClaw bill, Microsoft's Claude Code license cuts, a live Gemini-led price war, and the FinOps industry's own walk-stage/run-stage distinction.

We should be precise about what Cost Governance actually is, because the honest answer is more useful than a bigger claim would be. Matching a task to a cheaper model that can still handle it is not a new idea. RouteLLM is peer-reviewed. Salesforce meters Agentforce actions against a Flex Credit pool. A wave of AI gateway vendors, TrueFoundry among them, now ship routing as a default feature. None of that is Dipp AI's invention, and this piece will not pretend otherwise.

What Cost Governance is built around is a narrower, more defensible claim: nearly every one of those systems recommends a route. Very few enforce one, inside the same evidence chain that already verifies who authorized the directive and whether its data stayed in bounds. That gap — between a router that suggests and a system that will not let an unjustified frontier call happen at all — is where the governance problem actually lives, and 2026's data on AI governance broadly shows the same gap recurring everywhere organizations claim a control exists.

Cost Governance

Routing by stakes, enforced — not recommended

Pick a stakes class

Small / open class

1× relative cost · permitted

Classification, extraction, summarisation, routine drafting.

Mid class

halted by ceiling

Multi-step reasoning where an error is recoverable and reviewable.

Frontier class

halted by ceiling

Irreversible, high-stakes, or externally binding work only.

A directive with no external effect and a recoverable error path is capped at the small class. The frontier is not an option the router can pick.

Fig. 1 — The same token costs 4,500 times more depending only on which model answers it. Roughly nine in ten enterprise workloads never need the frontier. Pick a stakes class to see which tiers the ceiling permits and which are halted outright.

Claimed governance versus enforced governance, in the industry's own numbers

IBM's own research states the distinction more starkly than a vendor pitch would dare: 87% of organizations report having a clear AI governance framework in place, while fewer than 25% have actually implemented the specific controls needed to manage the risks that framework claims to cover. IBM's 2026 Cost of a Data Breach Report, released 29 July 2026, found 68% of breached organizations had no AI governance policy at all, and that five of the six governance controls the study measured in both years lost adoption year over year, moving backward while AI exposure kept growing. Only 19% reported governance and security teams working together at all.

Claiming governance and running governance are two entirely different things.

Evolvance Market Research, AI Governance Statistics 2026, on the gap between IBM's 87% claimed-framework figure and its under-25% implementation figure

The pattern is not unique to security governance broadly. It recurs specifically wherever a policy exists on paper without an enforcement mechanism behind it. A separate 2026 governance survey found organizations with the largest security budgets, including global enterprises, had the lowest representation in the top governance-maturity tier of any size band measured, while mid-sized organizations with far smaller budgets outperformed them. The reason given is specific and directly relevant to cost: the controls that actually determine whether a governance program holds up under a real incident — tested termination capability, attribute-based access enforcement, centralized audit logging — are not something budget alone buys. A large enterprise can afford the best routing dashboard on the market and still have no mechanism that stops an engineer from overriding it.

87% → <25%

claim a clear AI governance framework; have actually implemented the controls it requires

IBM

68%

of breached organizations in 2026 had no AI governance policy in place at all

IBM, July 29, 2026

5 of 6

governance controls IBM measured actually lost adoption year over year

19%

report governance and security teams working together at all

The fallacy, restated precisely

The Big Model Fallacy is not the belief that frontier models are good. It is the fact that no single decision to use one ever looks wrong in isolation, so the decision never gets challenged, one directive at a time, until an aggregate invoice makes the pattern visible months later. Glean CEO Arvind Jain estimated in June 2026 that roughly 95% of enterprise AI usage still runs on the most expensive frontier models available, even for tasks a cheaper alternative could handle just as well. CNBC's framing of that estimate is worth sitting with: model routing has become enough of a threat to frontier-model revenue that two executives close to the buildout, at companies whose business depends on frontier usage staying high, are the ones now telling enterprises to route away from it.

The scale behind that estimate is not abstract. Enterprise LLM API spending doubled in six months, from $3.5 billion in late 2024 to $8.4 billion by mid-2025. Gartner now forecasts $2.52 trillion in worldwide AI spending for 2026 alone, a 44% year-over-year increase. Token prices have never been lower, and the bill keeps exploding anyway, because agentic workflows consume 10 to 100 times more tokens per session than a chatbot interaction ever did, and usage volume is growing faster than prices are falling. Sixty to eighty percent of enterprise LLM cost concentrates in the 20 to 30% of use cases that are high-volume and low-complexity — exactly the category where a frontier model is most dramatically over-specified for what the task actually requires.

~95%

of enterprise AI usage still runs on the most expensive frontier models

Glean CEO Arvind Jain, via CNBC

$3.5B → $8.4B

enterprise LLM API spending, late 2024 to mid-2025 — a doubling in six months

$2.52T

Gartner's 2026 worldwide AI spending forecast, +44% year over year

60–80%

of enterprise LLM cost concentrated in just 20–30% of use cases — high-volume, low-complexity work

The scale problem is not exclusive to large enterprises, which makes it worth naming a smaller, named example directly. In May 2026, software engineer Peter Steinberger published the costs of running roughly 100 coding-agent instances on his own OpenClaw project. Over 30 days, the agents generated more than 603 billion tokens and 7.6 million requests, producing over $1.3 million in API spending. Steinberger later traced much of that spend to a single configuration choice, a high-rate "Fast Mode" setting, and estimated that disabling it alone would have cut costs by roughly 70%. Nobody had to make a bad model-selection decision for that number to happen. One default, left unexamined, was enough. Microsoft encountered a version of the same dynamic at enterprise scale and responded by canceling most internal Claude Code licenses, citing runaway token bills that had become unsustainable at the volume its own engineers were generating.

$1.3M

in 30 days of API spend on Peter Steinberger's OpenClaw project

603B

tokens generated across 7.6 million requests in that same 30-day window

~70%

of that cost traced to a single unexamined default — a high-rate "Fast Mode" setting

The price spread is not a fixed fact. It is a live, widening price war.

Elastic Compute Governance

Utilisation is a governance outcome, not a procurement problem

Toggle governance

5%

Average enterprise GPU utilisation across roughly 23,000 clusters, Cast AI 2026. The other 95% is paid for and idle.

Fig. 2 — An agent that re-plans has no natural stopping point. Without a ceiling enforced in the execution path, the only signal is the invoice.

Google has spent the second half of 2026 pushing budget-priced Gemini models directly at Anthropic and Microsoft, with CNBC's own coverage describing the strategy as a deliberate attempt to win enterprise customers on cost rather than raw capability. Gemini 3.6 Flash ships using up to 17% fewer tokens than its predecessor at a lower price per token, and Gemini 3.5 Flash Cyber, a specialized security model, is priced explicitly below Google's own larger general-purpose models. The competitive pressure is not limited to the three largest US labs: open-weight models out of China, Moonshot AI's Kimi K3 among them, are pushing the same price floor down further, and CNBC's own market analysts describe current AI pricing as still "unsophisticated," a market that has not yet settled into anything resembling equilibrium.

Enterprises are already reacting, and at least one reaction shows exactly what happens without a governed routing layer: an all-or-nothing switch rather than a calibrated one. Flo Crivello, CEO of AI startup Lindy, moved 100% of the company's traffic off Anthropic's Claude models onto DeepSeek, a cheaper Chinese open-weight alternative, specifically to bring runaway expenses under control. That is a legitimate response to a real cost problem, and it is also the same undisciplined, all-or-nothing pattern this piece has already argued against, just running in the opposite direction: instead of every task defaulting to the most expensive model out of habit, every task now defaults to the cheapest one, regardless of whether a given directive's actual stakes warranted the switch. Both OpenAI and Anthropic have had to respond with their own spend-control tooling under the same pressure — OpenAI with new analytics letting administrators break down credit spend and set usage limits, Anthropic with organization- and individual-level spending controls rolled out in August. Ramp co-CEO Eric Glyman's framing of why enterprises keep getting caught out is direct: most CFOs did not plan for AI's spend growth in their annual budgets and do not have the tools to manage it once it arrives.

A story every finance team will recognize

A widely cited 2026 account of what happens without enforcement reads less like a cautionary tale and more like a description most enterprises will recognize from their own budget review. A mid-size enterprise rolls out its first customer-facing AI agent in March. Three separate teams connect it to a frontier model using their own API keys, with no token tagging, no per-team budget, and no routing policy in place. By May, the CFO asks why the AI line on the cloud invoice grew 11 times over two months. Finance runs a week-long forensic review across four separate dashboards and still cannot determine which team owns 60% of the spend. Nothing about that scenario required a missing dashboard. Dashboards existed. What was missing was a mechanism that stopped the routing decision from being made ad hoc in the first place, by whichever engineer was closest to the API key.

After giving roughly 5,000 engineers access to an AI coding agent in December 2025, Uber saw usage nearly double by February 2026; by March, 84% of developers were classified as agentic coding users, and by April the company had burned through its entire 2026 AI budget. A separate enterprise reportedly spent $500 million in a single month after deploying AI access with no usage caps. Neither failure required a security breach. Both were a governance layer that could report spend after the fact but could not stop it while it was happening.

A July 2026 survey of 107 enterprise respondents found 21% track agent spend only through post-hoc logs, with no real-time mechanism to halt a runaway loop, and another 30% depend entirely on whatever cap their model provider happens to ship. A single runaway agent can consume $50 to $500 before anyone notices, multiplied by every concurrent user hitting the same failure pattern.

MarchMay
Routing policyNone — three teams, three sets of API keysStill none, discovered only in the invoice review
Cloud/AI invoice lineBaseline11× baseline
Spend attributionNot trackedUnresolved for 60% of spend after a week-long, four-dashboard forensic review
Fig. 3 — A mid-size enterprise's first customer-facing agent rollout, reported as representative of 2026 enterprise deployments generally.

Akshat Agrawal, a GenAI architect writing days before this piece was published, frames the diagnostic test plainly: is every request going to the same model regardless of difficulty; can spend be attributed by tier or only as a single monthly total; and is anyone measuring the cost of escalations that did not actually improve the answer — the purest form of leakage, because it is usually invisible until someone goes looking for it. His conclusion matches the distinction this piece is built around: the cost of an AI system is set by its architecture, not by which model happens to appear on the invoice.

The industry's own maturity model already names this distinction

The FinOps Foundation's 2026 State of FinOps report, a survey of 1,192 practitioners stewarding more than $83 billion in annual cloud spend, found 98% of FinOps teams now manage some aspect of AI spend, up from 63% the year before and just 31% two years earlier. That adoption curve looks like progress until a separate 2026 review of 127 enterprise agentic AI implementations found 73% went over budget anyway, some by more than 2.4 times their original estimate, burning roughly $2.3 million on costs nobody had anticipated. Managing AI spend and controlling AI spend turned out to be different claims, at the exact ratio this piece has already documented for AI governance broadly.

The unit economics of every AI feature your company ships will be set by people whose KPI is uptime, not margin.

THE D*AI*LY BRIEF, on why only 8% of the discipline controlling AI spend reports into finance rather than technology

That finding is the most useful one in this section, because it explains why the walk-stage-versus-run-stage gap persists even as FinOps adoption approaches saturation. A monitoring dashboard answers to whoever is accountable for uptime. A spend cap that can halt a directive mid-execution answers to whoever is accountable for margin. Ninety-two percent of the discipline currently sits with the former, which is a structural reason, not a tooling gap, for why alerts keep firing after the money is already spent. IDC's forecast expects the resulting gap between expected and actual AI spend at the largest global enterprises to reach 30% by 2027 if the incentive structure does not change alongside the tooling.

Industry FinOps guidance now formalizes the distinction this piece is built around, using its own maturity language rather than Dipp AI's. A widely cited 2026 AI FinOps guide describes a "walk stage," where organizations have workflow-level cost attribution and apply basic optimization like model-tier right-sizing, typically reducing blended inference costs 30 to 50%, and a separate "run stage," where the organization operates real-time cost controls: hard spend caps that automatically pause execution when a limit is reached, not an alert that fires after it. Most enterprises entering 2026 with active AI deployments, on the guide's assessment, sit at walk stage or below. The AWS FinOps Agent, launched at FinOps X 2026, is a useful marker of where even the most recent tooling still lands: it monitors costs, detects anomalies, and routes alerts to the responsible team without waiting for end-of-month reporting — a genuine advance over a monthly invoice, and still, by its own design, a detection system rather than a blocking one.

1,192

FinOps practitioners surveyed, stewarding $83B+ in annual cloud spend

FinOps Foundation, State of FinOps 2026

98%

of FinOps teams now manage some aspect of AI spend, up from 63% a year earlier and 31% two years earlier

73%

of 127 reviewed enterprise agentic AI implementations went over budget anyway

some by more than 2.4×

92%

of the discipline controlling AI spend reports into technology, not finance

What "enforced" actually means, gate by gate

FIG. A03-01

The Verified Execution Cycle

Tap a stage

1234567VECone billable unit

Directive

A named professional states intent. Nothing runs anonymously.

Fig. 4 — Where the ceiling actually binds: inside the cycle, before the action commits, not in a monthly reconciliation.
  1. Stakes classified against systems of record — the directive is scored using the same records the Logic Scrubber already checks, not a developer's subjective read on how hard a task feels.
  2. Model class selected, not suggested — the cheapest class satisfying both the stakes classification and the Data Control Gateway's boundary requirement is chosen. There is no button to click past it.
  3. Ceiling enforced in the inference path — a per-directive spend ceiling applies before execution, in the same path the directive runs through, not reconciled against a budget after the invoice arrives.
  4. Convergence monitored continuously — a directive looping without converging halts mid-execution, the moment its behavior stops making progress: the tested kill-switch capability the 2026 governance research found budget alone does not buy.
  5. Routing decision hashed, not logged separately — class chosen, alternatives considered, ceiling applied: written into the same evidence chain as the action it paid for, not a cost dashboard reconciled against an audit log after the fact.

The market's own data on what disciplined routing achieves, once it is actually enforced, is consistent across independent sources. Teams that implement a tuned routing layer report bill reductions in the 40 to 85% range without a visible drop in answer quality, because most production traffic never needed a frontier model in the first place. Peer-reviewed research on RouteLLM found 85% cost savings while retaining 95% of frontier-model quality in controlled evaluation. A separate 2026 guide found frontier models can cost 20 to 50 times more per token than a lightweight model handling the same category of request, and that roughly 70% of a typical enterprise's query distribution is simple enough for a budget model to handle correctly. None of those savings require a system to be smarter than the routers already on the market. They require the routing decision to actually hold.

Recommendation-only routingEnforced Cost Governance
When the cheaper model is selectedSuggested; a developer or agent can override itSelected automatically; nothing to override
Runaway loop detectionVisible in a monthly report, after the spendHalted mid-execution, the moment convergence stops
Spend attributionReconstructed later, often across several dashboardsHashed into the same record as the directive, in real time
What a governance audit findsA policy document and a partial control setA routing decision provable at the moment it was made

A representative case, reported by SSNTPL in July 2026, gives a concrete before-and-after: a US-based mid-market SaaS company, roughly 120 engineers, deployed internal coding and support agents in Q1 2026, initially routing essentially everything to a single frontier model. Initial spend ran approximately $40,000 a month on that one model class, for every task regardless of actual complexity. After implementing tiered routing and prompt caching, task completion and human review acceptance stayed flat or improved, because the right model was matched to the right step, and cache hit rate stabilized between 65% and 78% on repeated prefixes. SSNTPL reports the case as representative of the outcome when cost governance is treated as an architectural layer rather than an afterthought — not an isolated success story.

Regulation is starting to price the same gap

FIG. A01-03

What the regimes now require

Select a milestone

EU AI Act enforcement powers

The Commission can demand model evaluations and source-code access, restrict market access, and fine up to €15M or 3% of worldwide annual turnover.

Fig. 5 — Regulatory milestones that price the same gap — disclosure, attribution, and an auditable trail of who authorized the spend.

Cost control is not usually framed as a compliance matter, but the regimes arriving through 2026 increasingly require an organization to demonstrate that a stated control is operative, not merely documented. Colorado's AI Act took effect 30 June 2026, and California's Automated Decision-Making Technology rules begin full enforcement 1 January 2027, both requiring documented risk assessments and technical controls for high-risk AI systems, not policy statements about intent. Global spending on AI governance and compliance is projected to reach $2.54 billion in 2026, growing to $8.23 billion by 2034, and 27% of organizations have not even technically verified whether the AI vendors they already rely on use their data for model training, extending sensitive information into third-party systems on trust alone. Every one of those figures describes the same underlying condition Cost Governance is built to close for the cost dimension specifically: a policy that exists in a document is not the same claim as a control that holds when a real directive tests it. A spend ceiling that halts is demonstrable; a dashboard that reports is not. The same evidence chain that satisfies an auditor asking who authorized an action answers a CFO asking why a class of work was routed where it was.

What enforced routing actually protects

The saving is real, but it is the second-order effect. The first is that an enterprise regains the ability to say yes to more agentic work, because the downside of a bad directive is bounded by construction rather than by attention. Every routing decision Cost Governance makes runs through the same infrastructure layer, and lands in the same record that proves who authorized the directive and whether its data stayed in bounds. Cost is not a separate ledger reconciled after the fact. It is one more property of the same evidence chain. Observability then shows the whole run across providers in a single trace, so the routing decision is legible after the fact as well as enforced during it.

Sources

Sources for every figure in this article.

Where a number comes from Dipp AI's own analysis or an observed deployment, it is labelled as such and is not presented as an independently audited third-party finding.

  1. Evolvance Market Research, "AI Governance Statistics 2026: Key Data & Insights," May 31, 2026, citing IBM data on claimed versus implemented AI governance frameworks.
  2. IBM and the Ponemon Institute, "2026 Cost of a Data Breach Report," released July 29, 2026, as summarized by ComplexDiscovery, "Policy without control."
  3. Kiteworks, "The 2026 Annual Survey Report Is In: The AI Governance Gap Didn't Close. It Widened."
  4. CNBC, "Model routing is a fix for AI overspending. That's a problem for OpenAI and Anthropic," June 5, 2026, including Glean CEO Arvind Jain's estimate.
  5. TrueFoundry, "LLM Cost Optimization: Why an AI Gateway Is the Missing Layer," June 2026, including Gartner's 2026 worldwide AI spending forecast and cost concentration in high-volume, low-complexity use cases.
  6. Growth Acceleration Partners, "AI Cost Optimization: Controlling Runaway Token Spend," June 15, 2026, including Peter Steinberger's OpenClaw case study.
  7. Portal26, "AI Agent Cost Control: Stop Agents Burning Budget," including Microsoft's cancellation of internal Claude Code licenses.
  8. CNBC, "OpenAI and Anthropic face new AI reality as users shift from 'tokenmaxxing' to efficiency," June 26, 2026, including Lindy CEO Flo Crivello's switch to DeepSeek, Ramp co-CEO Eric Glyman's remarks, and Google's Gemini 3.5 Flash Cyber and Gemini 3.6 Flash budget pricing.
  9. TrueFoundry, "AI Cost Optimization: A Practical Guide for 2026," April 2026, including the illustrative mid-size enterprise routing failure.
  10. Akshat Agrawal, "Why Enterprise AI Costs Are an Inference Problem, Not a Training One," BigDATAwire / HPCwire, August 24, 2026.
  11. FinOps Foundation, "State of FinOps 2026," survey of 1,192 practitioners stewarding more than $83 billion in annual cloud spend.
  12. THE D*AI*LY BRIEF, "AI FinOps in 2026: 73% Blow Budget, 98% Now Track," July 6, 2026, including the review of 127 enterprise agentic AI implementations and IDC's 2027 spend-gap forecast.
  13. Vitaloralife, "AI FinOps: Enterprise Cost Governance Guide 2026," June 16, 2026, including the walk-stage/run-stage maturity model and the AWS FinOps Agent, launched at FinOps X 2026.
  14. Digital Applied, "LLM Model Routing in 2026: Cost-Quality Optimization," June 14, 2026, including RouteLLM peer-reviewed results.
  15. Kosmoy, "Smart LLM Routing: Cut Your AI Bill up to 40% (2026 Guide)."
  16. SSNTPL, "AI Agent Cost Reduction 2026: Model Routing & Prompt Caching Guide," July 2026, including the representative Q1 2026 mid-market case study.
  17. Kiteworks, "AI Data Governance Enforcement 2026: Compliance Guide," May 20, 2026, on Colorado's and California's 2026–2027 enforcement deadlines.
  18. SQ Magazine, "AI Compliance Cost Statistics 2026," April 1, 2026, on global AI governance spending projections.
  19. The Enterprise Superintelligence Report, Vol. I, Dipp AI Technologies, August 2026.

Case studies

Organisations that ran this argument in production.

Modelled reference scenarios with the measurement window, the components enforced and the numbers attached. Each one downloads as a PDF.

Written by Odero Otieno.

Writes the Dipp AI record on enforced governance for agentic systems — authority, enterprise data boundary, cost and compute. Every figure in this piece carries a source, and corrections are published on the record rather than made quietly.