News · September 16, 2026

Introducing the Data Control Gateway: Why “We Won't Train on Your Data” Isn't a Control

A promise in a contract is not enforced anywhere a machine can check it. The Data Control Gateway redacts, routes, and verifies an enterprise's own data boundary at the point of every call.

By Odero OtienoFounder, CEO & CTO, Dipp AI Technologies, Inc.

Every payload that would otherwise leave the tenant is redacted, region-pinned, and checked against training-exclusion terms before it crosses the boundary — on every call, not once at onboarding. If a model class cannot satisfy the boundary a directive requires, the route is refused and the next sufficient class is selected automatically.

Today we are introducing the Data Control Gateway, the Orcher component that makes an enterprise's data boundary a property of the system a directive runs through, rather than a sentence in a contract nobody can verify at runtime. Every payload that would otherwise leave the tenant is redacted, region-pinned, and checked against training-exclusion terms before it crosses the boundary — on every call, not once at onboarding. If a model class cannot satisfy the boundary a directive requires, the route is refused and the gateway selects the next sufficient class automatically.

We built this component because the industry's current answer to "will you train on our data" is a promise, and a promise is not something a machine can check.

Data Control Gateway

The boundary is a decision the system makes on every call

Select a stage

The payload is classified before anything leaves the tenant. Regulated fields, identifiers, and privileged material are tagged at the field level, not at the document level.

Fig. 1 — The gateway is the boundary itself: classification, redaction before egress, region pinned in place, training exclusion verified per call, and a refuse-by-default route decision. Select a stage to read what it enforces.

The promise the rest of the industry is still selling

A recent industry analysis of enterprise AI data security states the core problem more plainly than most vendors would want stated: even when a provider says corporate data won't be used for training, there is no reliable way to verify it, because the data crosses a web of servers and APIs an enterprise will never have direct access to. Most AI platforms offer little to no visibility into where that processing physically happens, and unless an enterprise specifically negotiates regional hosting, its data may sit in a jurisdiction that conflicts with its own legal obligations.

The clearest live demonstration of that gap is recent. Starting 17 August 2026, Atlassian began using data from its cloud products — Jira, Confluence, Jira Service Management and others — to train its AI offerings by default, affecting approximately 300,000 customers worldwide. Legal analysis of the change flags two questions: a shift in whether Atlassian is acting as a data controller or a processor for that training use, and a residency gap for customers who assumed regional hosting also meant regional training exclusion. The opt-out settings exist. They are off by default for most tiers, and an enterprise has to know to go find them before the change takes effect.

300,000

customers affected by one vendor's default training-data policy change

Atlassian, August 17, 2026

40%

of organizations report using access controls on their AI models and data at all

38%

of AI leaders report high confidence in their own cloud security posture

Per call

frequency at which the gateway re-verifies exclusion and residency

not per contract

Even the labs that build these systems have leaked data by accident

The risk is not limited to third-party vendors reading the fine print differently than customers expected. Frontier labs have themselves disclosed incidents in which conversation data, titles, or payload fragments were exposed to the wrong accounts through infrastructure faults rather than attacks. Nothing in those incidents required an adversary; they required only that the boundary existed in a policy document rather than in the execution path. A gateway that redacts before egress limits what an infrastructure fault can expose in the first place, because the material was never sent.

A breach in the gateway layer itself reached thousands of companies

The integration layer between enterprises and AI providers has already proven to be the most consequential single point of failure in this stack. The Salesloft–Drift compromise saw OAuth tokens stolen from one third-party integration and used to reach Salesforce instances across more than 700 potentially affected organizations. The lesson is not that integrations are dangerous. It is that a boundary enforced inside a single integration is only as strong as that integration's worst week — which is why Orcher enforces it in the control plane, above any one connector.

Healthcare shows what the boundary actually protects against

In regulated care settings the boundary question is not abstract. A payload containing clinical detail that leaves an enterprise's estate for inference does not merely create commercial exposure; it creates a disclosure with statutory consequences and a patient on the other end of it. The gateway's field-level redaction exists precisely for this case: the directive still succeeds, because the model rarely needs the identifying fields to do the work it was asked to do. What crosses the boundary is what the task requires, and nothing else.

No sector makes the stakes more concrete. Healthcare has been the most expensive industry for data breaches for fourteen consecutive years, at an average of $7.42 million per incident, and healthcare breaches take the longest of any sector to find and contain — an average of 279 days, roughly five weeks longer than the global average. Those two figures are related: the longer protected health information sits exposed before anyone notices, the more of it leaks and the more the eventual cleanup costs. The 2024 attack on Change Healthcare exposed the records of roughly 192.7 million individuals — the largest healthcare breach ever recorded, and on its own responsible for two-thirds of that year's total industry exposure.

Enforcement is not slowing down. In March 2026, HHS's Office for Civil Rights announced a financial penalty after finding a covered entity had failed to conduct a required risk analysis, failed to meet breach-notification obligations, and impermissibly disclosed the electronic protected health information of 15 million individuals. OCR closed 21 enforcement actions in 2025, its second-highest annual total, and every case in its current initiative cites the identical root cause: the entity failed to conduct an accurate, thorough assessment of the risks to the confidentiality of its own electronic health information. The single-tier civil monetary penalty cap now stands at $2,190,294, adjusted annually for inflation, and cumulative multi-million-dollar fines are routine once a violation is found to have persisted for years.

Every one of those actions traces back to the question the Data Control Gateway answers automatically: did this specific system, at this specific moment, have a documented, verifiable basis for the disclosure it made? A risk analysis performed once at deployment and never revisited is exactly the kind of static, contractual answer that has already failed — in Atlassian's training-policy shift, in Anthropic's exclusion-list oversight, and now in fourteen straight years of healthcare paying the most to learn the same lesson twice.

Sovereignty is now a named barrier, not a side conversation

FIG. A01-03

What the regimes now require

Select a milestone

EU AI Act enforcement powers

The Commission can demand model evaluations and source-code access, restrict market access, and fine up to €15M or 3% of worldwide annual turnover.

Fig. 2 — Sovereignty requirements as they land. Each regime asks for a provable boundary, not an assurance in a contract.

Data sovereignty has moved from a procurement footnote to a named blocker in enterprise AI programmes across the EU, the Gulf, and Asia-Pacific. NTT DATA's 2026 Global AI Report found more than 95% of enterprise AI leaders call private and sovereign AI important, while only 29% are prioritising it in any concrete, near-term way. Nearly 60% cite cross-border data restrictions as a major adoption challenge, and about 35% of Chief AI Officers name building and managing models inside private or sovereign environments as their single largest barrier. The gap NTT DATA describes is architectural, not aspirational: enterprise systems were built to move data across clouds, applications, and borders with increasing speed, and AI is the first workload that requires the opposite — keeping data inside a defined jurisdiction while still being useful.

IBM's 2026 Cost of a Data Breach report, built on 21 years of Ponemon Institute research and 602 breached organisations this year, states the shift in blunt economic terms: AI is compressing the time between exposure and impact, and only 40% of organisations report using access controls on their AI models and data at all. Legal practice is reaching the same conclusion from the contracts angle — but a contract clause is still a promise, even in writing. It answers what a vendor said it would do. It does not answer whether the system enforced it on the specific call that mattered.

The real danger is irreversible IP leakage, not a compliance checkbox, and the only fix is getting vendors to put 'no training on our data' in writing, not buried in the tool's settings.

Jake Vollebregt, Senior Counsel, Dykema, August 3, 2026

How the gateway decides, on every call

Select a component to reveal its Vol. I excerpt.

Layer 1 — sequential gatesLayer 2 — continuous controls01Directive InterfaceIntent captured02Role Identity FabricAuthority bound03Logic ScrubberAction verified04Immutable Audit LedgerEvidence hashed05Data Control GatewayData boundary + no-training06Cost GovernanceCeilings on autonomy07ObservabilityCross-provider traceOrcher™ · Dipp AI Technologies

Fig. CP-01 · The seven-component control plane

Fig. 3 — The gateway inside the wider control plane: a continuous Layer 2 control that governs every call for the life of the directive.
  1. Classify the payload at field level, before anything leaves the tenant.
  2. Redact everything the directive does not require to succeed — member_id and dob tokenised, a diagnosis code replaced with a policy-safe token, the task itself preserved so the directive can still be fulfilled.
  3. Pin processing to the region the enterprise's obligations require; a route that cannot hold the region is not offered.
  4. Verify training-exclusion terms for the specific model class about to be called, at call time, not from a contract signed months earlier.
  5. Route to the cheapest sufficient class that satisfies all of the above — or refuse, select a sufficient alternative automatically, and record the refusal.

The same principle hardware security is converging on

Gartner's 2026 forecast names the hardware version of the identical principle — confidential computing — as one of only three core "Architect" technologies expected to shape enterprise infrastructure over the next five years, projecting that by 2029 more than 75% of processing run on untrusted infrastructure will be secured in use, not just at rest or in transit. NIST's 2026 draft guidance on hardware-enabled confidential computing describes the shift in language that could describe this gateway directly: moving from "trust the infrastructure because access is restricted" toward "verify the execution environment before allowing the workload to access sensitive assets." NVIDIA's reference architecture for Confidential AI Factories combines CPU trusted execution environments, confidential GPUs, and remote attestation across an ecosystem spanning Red Hat, Intel, Anjuna, Fortanix, Edgeless, OPAQUE, Dell, HPE, Lenovo, Cisco, and Supermicro.

75%+

of processing on untrusted infrastructure projected to be secured in use by 2029

Gartner

3

core "Architect" technologies Gartner names through 2030 — confidential computing among them

12+

named infrastructure vendors building toward the same zero-trust reference architecture

NVIDIA ecosystem

The Data Control Gateway operates at a different layer than a hardware enclave, and the two are complementary rather than competing: a TEE protects data while a model computes on it; the gateway decides whether that data should have reached that model, in that jurisdiction, under that training-exclusion status, in the first place. An enterprise running Orcher inside a confidential-computing environment gets both guarantees at once. Encryption at rest and in transit was never the gap. The gap — as Gartner, NIST, and NVIDIA each concluded independently this year — was always the moment of use, and that is the exact moment the gateway governs.

Regulation is starting to make the same distinction

The regimes converging through 2026 — the EU AI Act's high-risk provisions, the revised Product Liability Directive, sectoral rules in health and financial services, and state statutes in Texas and California — increasingly distinguish between a stated safeguard and an enforced one. California's newest automated decision-making technology rules apply once a system materially influences a consequential decision, with no exclusion for a human being nominally in the loop: developers must document intended uses, the categories of personal data used in training, and known limitations, while deployers must give consumers a pre-use notice and, if an adverse outcome results, a plain-language explanation within 30 days. Each of those obligations depends on stating accurately what data actually crossed which boundary for a given decision. A gateway that verifies the crossing at call time produces that answer as a byproduct; a gateway that only enforces policy in a document produces a compliance project every time a regulator asks. Trust & standards sets out how Orcher maps to each regime.

What a verified boundary makes possible

Dipp Intelligence

What compounding actually looks like

Step through the cycles

A claims adjuster's directive is verified against policy language and precedent. The outcome, approved with specific reasoning, becomes a Dipp Intelligence entry.

Fig. 4 — A verified boundary is what makes accumulation safe: corrections stay inside the enterprise instead of leaving with the model.

The point of enforcing the boundary is not caution for its own sake. It is that enterprises which can prove the boundary held can safely put far more of their real work through agents — and everything that work produces stays theirs, compounding into Dipp Intelligence rather than into someone else's training corpus.

Sources

Sources for every figure in this article.

Where a number comes from Dipp AI's own analysis or an observed deployment, it is labelled as such and is not presented as an independently audited third-party finding.

  1. 2026 industry analysis of enterprise AI data security and provider training-data verification.
  2. Atlassian cloud products AI training policy change effective August 17, 2026, and accompanying legal analysis.
  3. Salesloft–Drift OAuth token compromise and downstream Salesforce exposure reporting, 2025–2026.
  4. 2026 survey data on AI access-control adoption and cloud security confidence among AI leaders.
  5. Confidential computing and attested-enclave literature, 2016–2026.
  6. NTT DATA, "2026 Global AI Report: A Playbook for Private and Sovereign AI," May 14, 2026.
  7. IBM, "2026 Cost of a Data Breach Report," with the Ponemon Institute, 602 organizations studied.
  8. Pahi Mehra, interview with Jake Vollebregt, Senior Counsel, Dykema, "Keeping Proprietary Data Out of AI Training Models," August 3, 2026.
  9. HIPAA Journal, "Healthcare Data Breach Statistics — Updated for 2026"; March 2026 Healthcare Data Breach Report on the OCR action involving 15 million individuals.
  10. Cryptonomist, "AI Supply Chain Breach Exposes 2,500+ Companies in 2026," August 11, 2026, citing CloudSEK's LiteLLM investigation.
  11. Gartner, "Top Strategic Technology Trends for 2026," October 18, 2025, ID G00829643, on confidential computing.
  12. Prolifics, on NIST's 2026 draft guidance on hardware-enabled confidential computing; NVIDIA, "Building a Zero-Trust Architecture for Confidential AI Factories," 2026.
  13. California automated decision-making technology (ADMT) regulations, as summarized by Hinshaw & Culbertson LLP, 2026.
  14. Regulation (EU) 2024/1689; EU Product Liability Directive; Texas TRAIGA; California SB 942.
  15. The Enterprise Superintelligence Report, Vol. I, Dipp AI Technologies, August 2026.

Case studies

Organisations that ran this argument in production.

Modelled reference scenarios with the measurement window, the components enforced and the numbers attached. Each one downloads as a PDF.

Written by Odero Otieno.

Writes the Dipp AI record on enforced governance for agentic systems — authority, enterprise data boundary, cost and compute. Every figure in this piece carries a source, and corrections are published on the record rather than made quietly.