Back to Insights

Every SaaS Tool You Use Just Added AI. Nobody Asked Where Your Data Goes.

## AI-Powered Software: The Hidden Data Flows in Your Stack ### The Problem Microsoft Copilot is embedded in your Outlook. Slack's AI summarizer is analyzing every channel conversation. Your CRM's...

Every SaaS Tool You Use Just Added AI. Nobody Asked Where Your Data Goes.

AI-Powered Software: The Hidden Data Flows in Your Stack

The Problem

Microsoft Copilot is embedded in your Outlook. Slack's AI summarizer is analyzing every channel conversation. Your CRM's "AI insights" features are sending customer data to external APIs — by default, enabled automatically on the day the feature launched.

Nobody asked you. Nobody asked your security team. Nobody asked your data protection officer. The features were added to software you had already approved, under terms that were updated in changelog footnotes, through an opt-out mechanism that requires knowing the feature exists in order to disable it.

This is Vector 3 of the shadow AI exposure problem: AI embedded in your approved software. Not shadow usage by employees. Not unauthorized API access. Approved, licensed, IT-sanctioned software that began processing your data through external AI infrastructure on an update date you cannot identify, with retention terms you did not negotiate, under jurisdiction you did not evaluate.

The most dangerous data flows are the ones your users do not know are happening.

What Your Approved Software Is Doing

Microsoft Copilot for Microsoft 365 processes every email you summarize, every meeting transcript it generates, every document it reviews. The processing runs on Microsoft's infrastructure. Microsoft's data residency commitments and retention policies govern what happens to the data. These commitments are in the enterprise agreement — which was signed before Copilot features existed at their current scope, under terms that have been updated by changelog.

Salesforce Einstein processes your CRM data — every lead, every deal, every customer interaction — to generate AI insights and predictions. The processing happens on Salesforce infrastructure. The model that generates those insights trains on aggregated data from Salesforce's entire customer base, under terms that permit Salesforce to use aggregated, anonymized data for model improvement. Your customer relationships are part of the training dataset.

Slack's AI features, launched in 2024, made workspace conversations available for external AI processing. Every message in every channel — including channels where deal strategy, client data, and personnel decisions are discussed — became processable by AI infrastructure outside the organization's perimeter. The feature was enabled by default for paid workspaces.

SAP Joule, Google Workspace Duet AI, HubSpot's AI features, Zendesk's AI-powered support tools, Notion AI, Grammarly's enterprise offering — each adds an AI processing layer to software that was evaluated and approved before that layer existed. The approval covered the software's original function. The AI features were not part of the evaluation.

Netskope's January 2026 research identified 223 sensitive data incidents per company per month related to AI usage, with the top quartile experiencing 2,100 incidents per month and the trend accelerating at plus 6 percent per month. A meaningful portion of these incidents trace to embedded AI features, not to deliberate employee use of unauthorized tools.

The Architecture of Invisible Exposure

The storybook analysis of shadow AI distinguishes three vectors. Vector 1 is employee behavior — deliberate use of unauthorized AI tools. Vector 2 is vendor behavior — law firms, accounting firms, and consultants using AI to process their clients' confidential data. Vector 3 is software behavior — AI embedded in approved tools, enabled by default, operating without user knowledge.

Vector 3 is architecturally distinct from the others. It is not addressed by AI usage policies, because the users are not making choices about AI usage. It is not addressed by employee training, because employees are not doing anything wrong. It is not addressed by unauthorized access controls, because the software is authorized. Vector 3 is addressed only by modifying the software configuration to disable AI features, by replacing the software with sovereign alternatives, or by intercepting the data flows at the infrastructure layer.

The interception approach is the SIA Router's function. A Router deployed at the perimeter of the organization's infrastructure can classify data flows by destination, inspect whether outbound requests are headed to known AI processing endpoints, and block or redirect them according to the organization's governance policy. Microsoft Copilot traffic destined for Azure AI infrastructure can be intercepted. Salesforce Einstein traffic can be identified and logged. The sovereign infrastructure creates visibility into data flows that the approved software creates without announcement.

This is not the same as blocking the software. A Router-based approach can permit the software to function for non-sensitive data categories while blocking AI processing of data classified as confidential. The SIA Hybrid Intelligence principle — routing between local and cloud by sensitivity level — applies to Vector 3 as much as to deliberate API usage.

The Consent Problem

When Microsoft adds Copilot to Outlook and a user clicks "Summarize this email," the user is consenting to Copilot processing that email. The user is also, without necessarily knowing it, creating a data processing event subject to GDPR. The user is the data subject. The email likely contains third-party personal data — the sender's information, the names of people discussed in the conversation, client contact details. Those third parties did not consent to their data being processed by Microsoft Copilot.

GDPR Article 6 requires a lawful basis for processing. The processor — Microsoft — has a lawful basis in the form of the contract with the organization. The organization's lawful basis for processing third-party personal data is the organization's legitimate interests or the performance of a contract with those third parties. When a lawyer pastes a client's personal data into Copilot, the lawful basis for that processing requires the client's personal data to be processed under terms the client was not told about, for purposes the client did not authorize.

The professional secrecy issue is parallel. Once case strategy flows through US AI cloud infrastructure, professional secrecy is structurally compromised. The architecture creates the breach. Microsoft's data processing terms do not restore professional privilege. Salesforce's enterprise agreement does not restore confidentiality obligations. The data left the perimeter. The obligation attached to it does not follow.

The LayerX 2025 analysis found that 92 percent of enterprise AI usage converges on OpenAI infrastructure — either directly or through software that uses OpenAI's API as its backend. Salesforce Einstein uses OpenAI models. Microsoft Copilot uses OpenAI models. GitHub Copilot uses OpenAI models. Organizations that have deliberately avoided direct OpenAI API usage may find that 80 percent of their software stack is routing through OpenAI infrastructure through embedded features.

The SIA Router as the Remediation Architecture

The SIA methodology's response to Vector 3 is the Router, deployed as an infrastructure-level data flow controller.

The Router's sensitivity classification function is what distinguishes it from a traditional firewall. A traditional firewall blocks traffic based on source and destination — it cannot evaluate the sensitivity of the content being transmitted. The SIA Router classifies content at the data level: this document contains personal data, this query includes financial projections, this message discusses client strategy. The classification drives the routing decision: does this data flow go to an external AI processing endpoint, or does it route to sovereign inference infrastructure instead?

For Vector 3 exposure specifically, the Router creates a new category of organizational visibility. Before Router deployment, the organization's answer to "how much of our data is being processed by Microsoft, Salesforce, and Slack AI features right now?" is "we don't know." After Router deployment, the answer is a measurement. Every data flow is logged. The Recorder captures every inference. The governance function can distinguish authorized AI processing from Vector 3 exposure.

The visibility enables governance. An organization cannot govern data flows it cannot see. The SIA Recorder's audit trail — immutable, stored inside the sovereignty perimeter, accessible to compliance — is the instrument that makes Vector 3 governance possible.

The Audit Question

When regulators examine an organization's AI data handling, they will ask about software AI features. The EU AI Act's transparency requirements apply to AI system operators — and an organization that enables Copilot, Einstein, and Slack AI has, in the regulatory sense, become an operator of AI systems that process personal data. The organization is responsible for being able to demonstrate that appropriate technical and organizational measures are in place.

"The software vendor handles this" is not sufficient for GDPR Article 28 purposes if the organization has not conducted a Data Protection Impact Assessment for the AI features it has enabled, has not documented the lawful basis for processing by those features, and has not implemented controls that can be demonstrated to an auditor.

Eighty-nine percent of enterprise AI usage is invisible — no logs, no SSO records, no oversight traces, per LayerX's 2025 research. For Vector 3 exposure, the invisibility is structural: the approved software is creating data flows that no existing IT monitoring is configured to capture.

The SIA methodology's six-phase deployment approach addresses this sequentially. Phase 1 is Discovery and Data Audit: establishing inventory of what AI features are active in the software stack, what data they process, and where that data goes. This is a prerequisite for phases that follow. Organizations that have not completed this inventory cannot answer regulatory questions about their AI data handling.

The path from Vector 3 exposure to governed sovereign AI is the same path as for any other AI governance gap: inventory, assess, classify, route, log, govern. The SIA methodology provides the framework for each step. Certified practitioners provide the implementation capability. The organizational outcome is the same: every data flow visible, every inference logged, every compliance question answerable.

The question "where does your software's AI send our data?" has a specific, auditable answer. Or it does not. The difference between those two positions is sovereign infrastructure.

---

SIA-certified practitioners assess software-embedded AI exposure and architect data flow controls for the complete software stack. Information on the certification program is available at thesovereigninstitute.org.

← Previous If Your AI Model Is Open-Source, Ask Who Funded It Next → Your AI Logs Will Be Subpoenaed. Here's What They'll Find.

Full SIA methodology documentation and certification programs at thesovereigninstitute.org