Showing Posts From
Data protection
- 01 Jul, 2026
Who Owns AI Data Protection in a European Enterprise: CIO, CFO, and CTO
Most AI programmes in European enterprises have a data protection gap. When I ask executives to show me the data flow map for their AI programme, I get the vendor's marketing materials. When I ask who owns the audit trail the DSK guidance requires, I get a reference to the compliance team. When I ask who signed off on the accepted residual risk from cross-border transfers, I get silence. The gap isn't usually a failure to read the regulation. Legal reads it. The DPO reads it. Compliance consultants read it. The gap is that reading the regulation and owning the outcomes are different things. Compliance reviews correctly identify obligations and distribute them across shared responsibility matrices where no single person owns any of them. Shared responsibility in compliance is, in practice, nobody's responsibility. The regulation doesn't get filed under "legal's problem" or "IT's problem." Regulators issue fines to organisations, not to functions. The accountability has to be assigned clearly, at the executive level, before the AI programme starts — not reconstructed after an enforcement action asks who was responsible for what. This is the accountability map. Not a compliance checklist — a decision and ownership structure that tells three people specifically what they are responsible for and how those responsibilities connect. Why diffuse ownership produces consistent failures The pattern is predictable. GDPR review is assigned to legal. Technical implementation goes to IT. Vendor contracts go through procurement. The DPO is informed. Nobody owns the complete picture — the data flows, the contractual protections, the technical controls, and the audit trail — as a single coherent programme with a single accountable executive. Regulators, when they investigate, don't assess functions. They assess programmes. They want to know who was responsible for the data classification decision, who approved the vendor without completing a DPIA, who was accountable for the audit trail the DSK requires, and who accepted the residual cross-border transfer risk. When the answer to any of those questions is "that was a shared responsibility across teams," the investigation continues until it finds something it can attach liability to. The organisations that have closed the AI data protection gap aren't the ones with the best compliance teams or the most thorough GDPR reviews. They're the ones where a named executive owns the outcome — not the review process, not the documentation, but whether the programme is actually compliant. There is also a practical motivation. The June 2025 Federal Labour Court ruling confirming that GDPR compensation applies in employment data protection cases, combined with the EU AI Act penalty structure that exceeds GDPR, means the financial consequence of a gap is no longer a distant possibility. The risk register entry that says "GDPR violation — maximum penalty €20M, probability: low" needs to be updated to account for AI Act maximums at 7% of global turnover, German-specific BDSG exposure, and individual employee compensation claims. When the numbers change, the ownership question becomes more urgent. What the CIO must own The CIO owns the operational data protection programme — the running of it, not just the existence of documentation. Data classification before deployment. The organisation must know what data is sensitive, personal, confidential, and commercially restricted before any AI tool goes into production. This is not a one-time classification exercise — it is an ongoing operational discipline that determines what can enter an AI prompt, what can be used for fine-tuning, and what cannot leave the organisation. The CIO owns this because it requires knowledge of how data flows operationally, not just what the policy documents say should happen. Vendor due diligence on data processing terms. Where data is stored, what retention periods apply, whether data is used for training, who the subprocessors are, and what happens in a breach — these need to be verified against the current version of the vendor's data processing agreement, not the version that was signed at contract execution. Vendors update their terms. The DPA that was adequate twelve months ago may not reflect the current architecture. The CIO owns the ongoing assurance that vendor terms match the risk assessment the programme was built on. The audit trail. The June 2025 DSK guidance requires documentation across the AI system lifecycle. That documentation needs to exist as a live operational record — not a static document, not something reconstructed in response to a regulatory inquiry. The CIO who owns the AI programme owns this audit trail as an operational artefact that is maintained while the system is running. In Germany specifically: works council engagement. Any AI system that touches employee data or workflow in German operations requires formal works council approval before deployment. This is not a legal function initiative — it is an operational deployment decision, and the CIO who is deploying the system owns the process of obtaining works council agreement. Starting that process three weeks before planned go-live is too late. It needs to be a parallel workstream beginning when the tool is evaluated, not when it has been selected and purchased. What the CFO must own The CFO owns the financial exposure — understanding it, quantifying it, and ensuring it is accurately represented in the risk register and in the business cases for AI programmes. Financial exposure quantification. GDPR maximum penalties are €20 million or 4% of global annual turnover. EU AI Act maximum penalties for prohibited practices are €35 million or 7% of global annual turnover. These are not theoretical risk register entries — they are the ceiling of a realistic exposure range that needs to be calibrated to the actual state of the AI programme. If the data flow mapping hasn't been done, the DPIA hasn't been completed, and the works council approval is pending, the probability weighting on that exposure is not low. The CFO owns putting numbers on this — not the CISO, not legal, not the DPO. AI vendor contract review. Standard commercial AI agreements were not written with GDPR in mind. The CFO needs to ensure that data processing agreements are in place, standard contractual clauses are signed, subprocessor lists are contractually controlled, and exit clauses cover data deletion obligations. A procurement team that signs a vendor agreement without these terms in place creates financial exposure that sits on the CFO's balance sheet. The commercial relationship is a CFO ownership item, and the data protection terms are part of the commercial relationship. Business case accuracy. Any AI business case that assumes data flows freely across borders — without accounting for anonymisation costs, pseudonymisation infrastructure, DPA requirements, DPIA costs, or potential remediation — is financially incomplete. The compliance cost is part of the programme cost. A business case that ignores it doesn't become compliant because the numbers look better without the cost line; it becomes a business case that will need to be revised when the cost materialises. The CFO who approves the business case owns the accuracy of what is in it. What the CTO must own The CTO owns the architecture decisions that determine the organisation's data protection posture — and those decisions are made once, at design, with consequences that persist for the life of the system. Architecture decisions with data residency and jurisdiction implications. Choosing between a US-hosted foundation model API and an EU-hosted alternative is not a technical preference. It determines the transfer mechanism burden, the US CLOUD Act exposure profile, and the complexity of the compliance programme that must sit around it. The CTO needs to make this choice with the data protection implications explicitly on the table — not as an afterthought following a performance or cost decision. Once the architecture is built on a US-hosted model, the cross-border transfer problem is baked in. Pseudonymisation infrastructure. For AI workloads where the organisation is processing personal data on external models, pseudonymisation before data leaves the organisation is the most practical risk reduction available. Building and maintaining that infrastructure — the pseudonymisation pipeline, the on-premise storage of re-identification keys, the audit of what is and isn't pseudonymised before transmission — is a CTO ownership item. It is also an architecture decision made at design. Retrofitting a pseudonymisation layer onto a production AI system that was built without one is significantly more expensive than building it in. Membership inference and re-identification risk assessment. For AI systems trained or fine-tuned on internal data, the CTO needs to understand what patterns the model could learn that allow indirect re-identification of individuals in the training set — and what technical controls reduce that risk. This is not a post-deployment question. It is an architecture question that needs to be answered before the model is built, because the answer changes what data goes into the training set and how the model is structured. I've seen organisations discover a re-identification risk in production that would have been addressed by a different data pipeline choice made six months earlier. Connecting the three ownership areas The three ownership areas are not independent. They create dependencies that need to be coordinated. The CTO's architecture decision determines what transfer mechanism the CFO needs to account for in the vendor contracts. The CIO's data classification determines what data the CTO can include in AI training sets. The CFO's financial exposure quantification depends on knowing what the CIO's vendor due diligence has found and what residual risks the CTO's architecture choices carry. The connection point is a shared risk assessment that all three own as inputs. In practice, this means one document — not three separate function-level risk assessments — that maps the data flows, the architecture, the contractual protections, the accepted residual risks, and the financial exposure. The DPO reviews it. Legal reviews it. But the three executives who own the inputs own the conclusions. The audit trail the DSK guidance requires is effectively this document — maintained, updated, and available for regulatory review at any time. What to take from thisAssign named executive ownership before the AI programme starts. Data classification and audit trail (CIO), financial exposure and vendor contracts (CFO), architecture and technical controls (CTO) — with a named person, not a function, accountable for each. The DPO advises and audits — the DPO does not own the programme. Data protection programme ownership sits with operational executives because the programme is operational. The DPO's role is to assess compliance, not to run the systems being assessed. Build the shared risk assessment document before deployment. One document mapping data flows, architecture, contractual protections, accepted residual risks, and financial exposure — owned jointly by CIO, CFO, and CTO, reviewed by DPO and legal. This is the audit trail regulators will look for. Recalibrate the risk register to reflect AI Act exposure. The penalty ceiling at 7% of global annual turnover for prohibited practices means the financial exposure calculation has changed. Update the probability weightings to reflect the actual state of the AI programme — not a theoretical best-case assessment. For German operations: works council engagement is CIO-owned from the moment a tool is evaluated, not from the moment a rollout date is set. The approval process cannot be compressed into the final weeks of a deployment project. Review the business cases for AI programmes already approved against the actual compliance cost. Programmes approved without accounting for pseudonymisation infrastructure, DPA requirements, or DPIA costs were approved on incomplete financial models. The cost will materialise — the question is whether it appears as a planned expense or an unplanned remediation.The executives who have closed the AI data protection gap in their organisations didn't do it because they understood the regulation better than their peers. They did it because someone in the C-suite owned the outcome. The regulation tells you what is required. The accountability structure determines whether it happens. A compliance review that ends with a report to the board but no named owner for each obligation is not a programme — it is documentation of intent. And intent, when examined by a regulator, is not a defence.
Read full article
- 30 Jun, 2026
Building a Privacy-Safe AI Pipeline: What Goes Between Your Data and the Model
Most European companies running AI programmes have a data protection policy that says something like: "personal data should not be sent to external AI models without appropriate controls." What most of those same companies don't have is a technical architecture that enforces this — at the point where data actually leaves the organisation, before it reaches the model, in a way that doesn't depend on every employee making the right decision every time. Policy without technical enforcement is an aspiration. I've seen teams build careful data governance frameworks, train employees on what can and cannot enter a prompt, and roll out usage guidelines. Then I've seen the same teams, six months later, discover that customer records, HR files, and contract documents had been flowing into an external model because the work was faster that way and nobody had built anything to stop it. The technical answer is a processing layer between your data and the model — a pipeline that intercepts data before it reaches the LLM, applies privacy controls, and only passes through what meets the organisation's data protection requirements. Building it is an architecture decision, not a compliance decision. This article covers what that architecture looks like, the components that make it work, and where its limits are. Why policy-only approaches fail at the prompt The problem is structural. Employees use AI tools to get work done faster. The faster path almost always involves sending more context — a full customer record rather than selected fields, a complete document rather than the relevant section, a full email thread rather than the key paragraph. The model produces better outputs with richer input, so the incentive runs directly against the minimisation principle. Enforcing minimisation through policy means asking people to voluntarily produce worse outputs to protect data they don't directly manage. That is a losing proposition. The violation isn't usually malicious — it is the product of a rational trade-off that the person making the decision doesn't see as a trade-off at all. Architectural enforcement removes the trade-off. The pipeline intercepts data before it reaches the model and applies one of two controls: stripping personal data entirely where the use case doesn't require specific identity references, or replacing identifiers with on-premise-keyed tokens where the model genuinely needs to track the same entity across a document. In both cases, what reaches the model contains no original personal data. This is the design goal: make the compliant path identical to the convenient path. But it is important to be precise about what that compliance achieves. Neither stripping nor tokenisation makes data anonymous under GDPR. Stripping is the stronger control — no key exists, no reversal is possible — and it is what the pipeline should default to. Tokenisation is a fallback for use cases that require entity consistency, and tokenised data remains personal data under GDPR, as covered in detail in the anonymisation article in this series. The pipeline reduces risk significantly. It does not eliminate the organisation's GDPR obligations for what reaches the model. The core architecture: a processing layer the LLM never bypasses The pipeline has five logical components. All personal data must pass through all five before reaching the model. The LLM has no direct access to raw data — only to what the pipeline has processed and approved. Data ingestion: Documents, database records, API responses, or user prompts enter the pipeline. This is the boundary where classification begins — what type of data is this, what sensitivity level does it carry, and what processing rules apply. PII detection: A detection layer identifies personal data in the ingested content. This covers direct identifiers (names, email addresses, phone numbers, national identification numbers, passport numbers, IBAN and financial account numbers, IP addresses, biometric identifiers) and indirect identifiers that could contribute to re-identification (age combined with location, job title combined with department size, medical codes combined with demographic data). Detection uses a combination of Named Entity Recognition (NER) models and pattern-matching rules. NER models identify contextual entities — a name in a sentence, a location in a paragraph. Pattern-matching rules catch structured identifiers — a correctly formatted IBAN, an email address, a German Steuernummer. PII removal and tokenisation layer: This is the core control, and it operates in two modes depending on what the task requires. Stripping (default): Detected identifiers are removed entirely. A name disappears. An email address is deleted. No key is created, no mapping is stored, no reversal is possible. This is the stronger control and should be the default for use cases where specific identity is not required — summarisation, document classification, general drafting, sentiment analysis. Stripped data is the closest the pipeline can bring data to the GDPR anonymisation standard, though it still may not meet the full legal threshold if indirect re-identification from remaining attributes remains possible. Tokenisation (exception): Where the AI task genuinely requires consistent entity references across a document — contract analysis referencing the same counterparty throughout, an HR review tracking a specific employee's record — detected identifiers are replaced with consistent tokens (PERSON_001, EMAIL_001, ORG_002). The token mapping is stored in a key vault that lives entirely on-premise and is never transmitted. The model sees consistent references without knowing the real identity. This is pseudonymisation. Pseudonymised data is still personal data under GDPR. The pipeline's GDPR obligations — transfer mechanisms, DPA requirements, breach notification — apply to any tokenised content that reaches the external model. LLM processing: The stripped or tokenised content is sent to the external model. No original personal data is present. The model has no access to the key vault. Output de-tokenisation (tokenisation mode only): The model's response is returned to the pipeline. Where tokenisation was applied, the pipeline replaces tokens with the original identifiers using the on-premise key vault before delivery to the user. The key vault never left the organisation. What the vendor receives, what any external disclosure would expose, and what any CLOUD Act order could compel them to produce is the stripped or tokenised version. The reduction in exposure is real. It is not, however, a GDPR exemption. Building the PII detection layer The detection layer is where most implementations have their weakest point. NER models trained on general text perform well on common identifiers — names, locations, organisations in standard formats. They perform poorly on domain-specific identifiers, non-standard formats, and contextual indirect identifiers that require understanding of the data environment. For enterprise deployments, the detection layer needs to be tuned for the organisation's specific data. A healthcare provider will have medical record identifiers that a general NER model won't recognise. A financial services firm will have account reference formats specific to its systems. A logistics company will have shipment identifiers tied to individual customers. Microsoft Presidio is an open-source PII detection and anonymisation framework that provides a base NER capability with extensible recogniser architecture — custom recognisers can be added for organisation-specific identifier formats. spaCy provides the underlying NER model that can be fine-tuned on domain data. GLiNER is a more recent generalist NER model that performs well on zero-shot entity recognition, useful for catching entity types that weren't anticipated in the original configuration. The detection layer should be treated as imperfect by design. False negatives — personal data that is not detected and therefore passes through to the model as-is — will occur. The residual risk from false negatives needs to be part of the organisation's data protection impact assessment. For high-sensitivity data categories (health data, financial data, data on minors), a higher detection threshold or additional validation step is warranted. For lower-sensitivity data, the residual risk from occasional false negatives may be acceptable. Logging is essential. Every item that passes through the pipeline should generate a log entry: what data type was detected, what action was taken (stripped, tokenised, blocked), what was transmitted to the model, and when. This log is the audit trail that GDPR and the June 2025 DSK guidance require. It also lets the organisation identify patterns where specific data types are consistently slipping through the detection layer. Synthetic data as an alternative for AI development For organisations building, fine-tuning, or testing AI models on internal data, the pipeline approach described above is a runtime control — it operates on data flowing to the model during inference. A different problem arises at development time: the organisation needs representative data to develop, test, and fine-tune AI systems, and real personal data carries risk even in development environments. Synthetic data generation creates datasets that have the same statistical properties as real data without containing records about real individuals. For a customer dataset, the synthetic version preserves the distribution of ages, purchase patterns, geographic spread, and behavioural signals — but every record is generated, not derived from an actual customer. The GDPR position on genuinely synthetic data is more favourable than on pseudonymised data, because data generated from statistical distributions rather than from individual records may qualify as anonymous if the generation process cannot be reversed to recover the underlying individuals. This requires that the generative process genuinely produces new data rather than perturbing or recombining existing records — a meaningful technical distinction that affects both the data's legal status and its usefulness as a development dataset. For model testing and development environments, synthetic data removes the need to apply runtime privacy controls and reduces the risk that development data leaks through version control, shared notebooks, or development tooling. For production fine-tuning, synthetic data has quality limitations — complex behavioural patterns and rare edge cases are difficult to preserve through synthesis — but it is a viable alternative for many use cases where the fine-tuning goal is improving general performance rather than capturing rare behaviours. What this architecture doesn't solve The pipeline addresses the most direct GDPR exposure in AI: personal data flowing to an external model without appropriate controls. It does not address all exposure. Usage metadata is not affected. The vendor still sees query frequency, document structure patterns, workflow characteristics, and feature usage — the competitive intelligence risk described in the first article in this series. The pipeline controls what is in the data; it doesn't control that data is being sent. Transfer mechanisms still apply. Pseudonymised data is still personal data under GDPR. The pipeline reduces the sensitivity and impact of a disclosure event, but it does not remove the organisation's obligation to have a valid cross-border transfer mechanism in place for the pseudonymised content that does reach the external model. Fine-tuning risk remains for the training phase. If the organisation fine-tunes a model on pseudonymised data, membership inference risk applies to the trained model — the model may encode patterns that allow indirect re-identification, as covered in the anonymisation article in this series. The runtime pipeline prevents personal data from reaching the model during inference. It does not address patterns the model has already learned from training data. Homomorphic encryption is the emerging technical approach that would allow data to be processed by the model while remaining encrypted throughout — the model would never see plaintext data at any point. This would represent a fundamentally stronger privacy guarantee than pseudonymisation. In practice, homomorphic encryption adds significant computational overhead and is not yet deployable at enterprise scale for LLM inference. It is being actively researched and is recognised in EU policy frameworks as a potential compliance tool for the medium term. It is not a practical option for most enterprise AI programmes today. What to take from thisBuild the privacy pipeline before deploying AI to end users. The decision to deploy an AI interface to employees without technical privacy controls in place is a decision to rely on policy enforcement alone — which consistently fails. The pipeline is the enforcement mechanism. Start PII detection with a general-purpose NER model and extend it with custom recognisers for your organisation's specific identifier formats. General models will miss domain-specific and proprietary identifiers. Tuning for your data is required, not optional. Default to stripping, not tokenisation. Stripping removes PII entirely — no key, no reversal, no personal data transmitted. Use tokenisation only when the AI task genuinely requires entity consistency across a document, and treat tokenised data as personal data throughout, because under GDPR it is. Where tokenisation is applied, the key vault must live on-premise, isolated from all network-facing components, and never transmitted. If the key vault is in the same cloud environment as the model API calls, the architectural separation is broken. Treat the detection layer as imperfect and log everything. Build the DPIA assumption that some personal data will pass through undetected. Log all pipeline activity so you can identify detection gaps and demonstrate the audit trail regulators will ask for. Use synthetic data for AI development and testing environments. Remove real personal data from development pipelines entirely where possible — synthetic data preserves enough statistical structure for most development use cases while eliminating the exposure. Understand what the pipeline does and doesn't cover. Transfer mechanisms, metadata exposure, and fine-tuning data risk sit outside what a runtime pseudonymisation pipeline addresses. The pipeline closes the most direct exposure. The other risks need separate controls.The gap between "we have a data protection policy for AI" and "our AI systems are technically enforcing that policy" is where most European enterprises currently sit. Policy documents don't intercept a prompt. Architecture does. The pipeline is not complex to build in its basic form — a PII detection layer, a tokeniser, and an on-premise key vault are available as open-source components or commercial services. What requires investment is tuning detection for your specific data environment, integrating the pipeline into every AI touchpoint in the organisation, and maintaining it as the AI programme grows. That investment is also what makes the compliance position defensible when examined — because a working technical control is a fundamentally different answer to a regulatory question than a policy that assumed people would follow it.
Read full article
- 24 Jun, 2026
European Company Data in US AI Models: What Your Executives Don't Know Is Already Out
Six months after a European CTO signed off on an AI programme, the DPO sent a single-line email: which data has been sent to the model, where is it stored, and how long is it kept there? Nobody had an answer. The programme had been running on a standard enterprise API agreement. The vendor was reputable. The legal team had reviewed the terms. But the question of what data was leaving the organisation — in what form, to which infrastructure, under what retention policy — had never been mapped. I've seen versions of this conversation more times than I'd like. The AI programme gets approved at the business case level. The vendor agreement goes through procurement. The DPO gets looped in late, or not at all. By the time anyone asks the data flow question seriously, the model has been running for months and the answer is complicated. This isn't a scare piece. It's a map of what actually leaves a European organisation when it runs AI workloads on external models — because the gap between what executives believe is happening and what the contract actually governs is wider than most teams realise. GDPR has been in force since 2018. The enforcement precedent in the AI space is now live. And the organisations that haven't done the data flow mapping are running programmes that haven't been assessed for compliance. What enters a prompt has left the organisation Every query sent to an AI model — every instruction, every document, every piece of text — is transmitted to and processed on the vendor's infrastructure. This is obvious in principle and consistently underestimated in practice. The problem isn't that employees use AI tools. The problem is that the data flowing through those tools hasn't been classified before it flows. When a salesperson pastes a customer contract into a prompt to generate a summary, that contract has left the organisation. When an HR manager uploads a performance review to ask for a draft response, that review has left the organisation. When a finance analyst feeds a quarterly forecast into a model to restructure the narrative, that forecast has left the organisation. None of this requires a policy violation. It happens in ordinary, productive use — which is exactly why it's difficult to control retrospectively. Most enterprise AI APIs retain prompt data for abuse monitoring purposes unless the customer has explicitly opted out — and opted out in writing, in a way the contract records. Retention periods vary by vendor. They are subject to change. I have reviewed enterprise AI agreements where the data retention section was three lines long and referred to a separate policy document that had been updated four times in eighteen months. The business owner who approved the tool had never read either version. Data classification before any AI tool goes into production is not optional — it is the prerequisite. Without knowing what data is sensitive, personal, confidential, and commercially restricted, there is no basis for deciding what can enter a prompt. Most European enterprises have data classification policies. Most of those policies were written before AI prompt data was a category that needed to be governed. Personal data in prompts — minimisation is not enough When the data entering a prompt includes names, email addresses, identification numbers, health records, or financial information, the transfer isn't just a data governance concern — it is a GDPR event. Personal data has moved to an external processor, potentially outside the EU, under whatever transfer mechanism (or lack of one) applies to that vendor relationship. The standard response inside most European enterprises is to invoke data minimisation: GDPR's Article 5(1)(c) principle that personal data should be adequate, relevant, and limited to what is necessary. In theory, this means sending only the fields required for the AI task. In practice, it almost never works that way. Teams send full customer records because the model produces better outputs with more context. HR workflows send complete personnel files because partial information leads to incomplete responses. Customer service AI is fed entire interaction histories because isolated records produce generic answers. The business logic for sending more is always compelling, and the minimisation principle gets applied at the policy level while the actual data flows ignore it. Even when minimisation is applied properly, it is not a compliance answer for personal data. A dataset that contains only a customer's name, email address, and account status is minimal — and it is still fully personal data under GDPR. Minimising what you send does not change the legal classification of what arrives at the model. The transfer still requires a lawful basis. The vendor still needs to be a compliant data processor. The cross-border transfer rules still apply. The only way personal data falls outside GDPR scope when it reaches an external AI model is if it has been genuinely anonymised — not minimised, not pseudonymised, but anonymised to the standard where re-identification is not reasonably possible for anyone, including the model provider. That standard is far harder to meet than most teams assume, and AI models introduce a specific re-identification risk through pattern learning that conventional processing does not. This is covered in full in a separate post in this series. The practical position for European enterprises: names, email addresses, and any other direct identifiers should not enter an external AI model unless a valid GDPR lawful basis, a compliant data processing agreement, and a valid cross-border transfer mechanism are all in place — and verified, not assumed. The metadata picture your vendor is accumulating Prompt content is the obvious exposure. Usage metadata is the one that organisations consistently miss. AI vendors collect usage patterns: query frequency, document types, feature usage, workflow structures, which teams use the model for what. Individually, these data points look low-sensitivity. Collectively, over time, they build a detailed picture of how the organisation operates — what it's working on, where it's directing AI resources, which functions are scaling their usage and in what direction. This is a competitive intelligence risk that exists independently of whether any personal data is involved. A vendor who can see that a company's legal team started using the model heavily in March, its M&A function spiked usage in April, and its HR team has been running high volumes of termination-related drafts since May has learned something meaningful — without any named individual appearing anywhere in the data. Enterprise agreements typically don't restrict vendors from using aggregated, anonymised usage data for product improvement and benchmarking. The data processing agreement and the main commercial contract often address prompt data and metadata under different provisions, negotiated at different points in the contracting process, with different protections attached to each. The metadata protections are almost always weaker. Read both documents carefully — not just the DPA summary the vendor sends as a compliance checkbox. The training data assumption most enterprise teams get wrong There is a persistent assumption in European enterprises that the enterprise version of an AI tool doesn't use customer data for training. Sometimes this is correct. Often it is an assumption that hasn't been verified against the current version of the data processing agreement. Consumer-grade AI tools — the non-enterprise versions of products your employees are also using — may by default use user inputs to improve the underlying model. This is processing for a new purpose. Under GDPR, processing for a new purpose requires a separate lawful basis. When an employee uses the consumer version of a tool on work data because it's faster or the enterprise licence hasn't been set up on their machine, the organisation has a problem it didn't create deliberately but owns entirely. Enterprise agreements typically include data processing agreements that prohibit training on customer data. Three things determine whether that prohibition is actually protecting the organisation: whether it applies to all data types or only certain categories; whether it covers all infrastructure the vendor uses, including subprocessors; and whether the DPA reflects the current model architecture. I have seen enterprise agreements where the DPA predated a major vendor platform change by over a year. The prohibition on training was still in the original terms — which no longer accurately described what was happening with the data. The check is straightforward. Request the current, signed DPA. Confirm it covers training, all subprocessors, and the current model version. Don't assume continuity because the vendor's name hasn't changed. When model outputs carry a breach risk This risk sits with the deployer — and it is the one most European organisations haven't included in their AI risk registers. If a model has been trained on data that included personal information, it can under certain conditions reproduce elements of that information in outputs. Not reliably, not on demand — but under specific prompting patterns, models have been shown to surface fragments of training data. If European employee or customer data was part of the training set, and the model surfaces it in a response, the deploying organisation has a data breach on its hands. Under GDPR, that breach sits with the deployer. The model provider may carry liability, but the deployer cannot transfer its own. For foundation models from major providers, the deployer typically has no visibility into training data composition. The risk is difficult to quantify and impossible to eliminate entirely. For systems where the deployer has fine-tuned or augmented a foundation model with internal data — retrieval-augmented generation, fine-tuning on internal documents, embedding proprietary content — the deployer has direct control over what went in. That control increases responsibility. Before augmenting or fine-tuning a model with internal data, the architecture review needs to answer: what personal data is in these documents, what happens if the model later surfaces fragments of it, and what technical controls reduce that risk? This is a question most teams are not asking at the design stage. It becomes the question they wish they had asked after an incident. What the enterprise agreement covers — and what it doesn't The enterprise agreement creates contractual obligations. It doesn't create infrastructure controls. The data processing agreement requires the vendor to handle data in certain ways. It doesn't prevent US law enforcement from issuing a lawful demand to a US-incorporated vendor under US law — a structural problem I address in the next article in this series, focused specifically on cross-border transfer risk. What the enterprise agreement also doesn't replace is the audit trail that European regulators now expect to exist. The June 2025 guidance from Germany's data protection conference (DSK) requires documentation across the AI system lifecycle. That documentation needs to exist before regulators ask for it — not in the forty-eight hours after they do. For European companies with German operations, this requirement is not optional and it is not the vendor's responsibility to create. The enforcement precedent is live. A major AI provider was fined €15 million by Italian regulators in December 2024 for GDPR violations — failure to establish a lawful processing basis and inadequate transparency around how user data was handled. European regulators now have precedent, process, and political pressure to continue in this space. The organisations currently running AI programmes without a data flow map are running programmes that have not been assessed for compliance. The gap between those two states is not a legal technicality — it is the gap between a defensible position and one that cannot be defended when examined. What to take from thisMap data flows before deployment, not after. Every AI tool processing external data should have a documented flow — what enters, what is retained, where it's stored, for how long — completed before the tool goes live, not reconstructed months later. Read the current DPA, not the version in the procurement file. Data processing agreements are updated. Request the current signed version and confirm it covers training, subprocessors, and the current model architecture. Classify data before it enters a prompt. Sensitive, personal, commercially restricted, and confidential data should be classified before it enters any AI tool — not assessed after six months of use. Do not treat data minimisation as a substitute for anonymisation. Sending fewer fields does not change the legal classification of what arrives at the model. Personal data that has been minimised is still personal data — names, email addresses, and identifiers stripped down to the "minimum necessary" still carry full GDPR obligations. If personal data cannot be genuinely anonymised before it reaches an external model, the transfer needs a lawful basis, a compliant DPA, and a valid cross-border mechanism — not a smaller dataset. Treat usage metadata as a competitive risk, not just a data protection issue. Understand what your vendor collects on usage patterns, what restrictions apply, and what they're permitted to do with aggregated data. Enforce a clear policy on consumer versus enterprise tool versions. If employees have access to both, the policy needs to specify which applies to which data — because the data handling terms are materially different. Include model output re-identification in your AI architecture review. For any system fine-tuned or augmented with internal data, answer the re-identification risk question before build — not after the system is in production.GDPR has been in force since 2018. The data protection gap in AI programmes isn't a new legal problem — it is an existing problem that the pace of AI adoption is making visible at scale. Most of what European companies need to do to close this gap doesn't require waiting for new regulation. It requires applying existing obligations to a class of tools that moved faster than the governance frameworks around them.
Read full article