Showing Posts From
Gdpr
- 01 Jul, 2026
Who Owns AI Data Protection in a European Enterprise: CIO, CFO, and CTO
Most AI programmes in European enterprises have a data protection gap. When I ask executives to show me the data flow map for their AI programme, I get the vendor's marketing materials. When I ask who owns the audit trail the DSK guidance requires, I get a reference to the compliance team. When I ask who signed off on the accepted residual risk from cross-border transfers, I get silence. The gap isn't usually a failure to read the regulation. Legal reads it. The DPO reads it. Compliance consultants read it. The gap is that reading the regulation and owning the outcomes are different things. Compliance reviews correctly identify obligations and distribute them across shared responsibility matrices where no single person owns any of them. Shared responsibility in compliance is, in practice, nobody's responsibility. The regulation doesn't get filed under "legal's problem" or "IT's problem." Regulators issue fines to organisations, not to functions. The accountability has to be assigned clearly, at the executive level, before the AI programme starts — not reconstructed after an enforcement action asks who was responsible for what. This is the accountability map. Not a compliance checklist — a decision and ownership structure that tells three people specifically what they are responsible for and how those responsibilities connect. Why diffuse ownership produces consistent failures The pattern is predictable. GDPR review is assigned to legal. Technical implementation goes to IT. Vendor contracts go through procurement. The DPO is informed. Nobody owns the complete picture — the data flows, the contractual protections, the technical controls, and the audit trail — as a single coherent programme with a single accountable executive. Regulators, when they investigate, don't assess functions. They assess programmes. They want to know who was responsible for the data classification decision, who approved the vendor without completing a DPIA, who was accountable for the audit trail the DSK requires, and who accepted the residual cross-border transfer risk. When the answer to any of those questions is "that was a shared responsibility across teams," the investigation continues until it finds something it can attach liability to. The organisations that have closed the AI data protection gap aren't the ones with the best compliance teams or the most thorough GDPR reviews. They're the ones where a named executive owns the outcome — not the review process, not the documentation, but whether the programme is actually compliant. There is also a practical motivation. The June 2025 Federal Labour Court ruling confirming that GDPR compensation applies in employment data protection cases, combined with the EU AI Act penalty structure that exceeds GDPR, means the financial consequence of a gap is no longer a distant possibility. The risk register entry that says "GDPR violation — maximum penalty €20M, probability: low" needs to be updated to account for AI Act maximums at 7% of global turnover, German-specific BDSG exposure, and individual employee compensation claims. When the numbers change, the ownership question becomes more urgent. What the CIO must own The CIO owns the operational data protection programme — the running of it, not just the existence of documentation. Data classification before deployment. The organisation must know what data is sensitive, personal, confidential, and commercially restricted before any AI tool goes into production. This is not a one-time classification exercise — it is an ongoing operational discipline that determines what can enter an AI prompt, what can be used for fine-tuning, and what cannot leave the organisation. The CIO owns this because it requires knowledge of how data flows operationally, not just what the policy documents say should happen. Vendor due diligence on data processing terms. Where data is stored, what retention periods apply, whether data is used for training, who the subprocessors are, and what happens in a breach — these need to be verified against the current version of the vendor's data processing agreement, not the version that was signed at contract execution. Vendors update their terms. The DPA that was adequate twelve months ago may not reflect the current architecture. The CIO owns the ongoing assurance that vendor terms match the risk assessment the programme was built on. The audit trail. The June 2025 DSK guidance requires documentation across the AI system lifecycle. That documentation needs to exist as a live operational record — not a static document, not something reconstructed in response to a regulatory inquiry. The CIO who owns the AI programme owns this audit trail as an operational artefact that is maintained while the system is running. In Germany specifically: works council engagement. Any AI system that touches employee data or workflow in German operations requires formal works council approval before deployment. This is not a legal function initiative — it is an operational deployment decision, and the CIO who is deploying the system owns the process of obtaining works council agreement. Starting that process three weeks before planned go-live is too late. It needs to be a parallel workstream beginning when the tool is evaluated, not when it has been selected and purchased. What the CFO must own The CFO owns the financial exposure — understanding it, quantifying it, and ensuring it is accurately represented in the risk register and in the business cases for AI programmes. Financial exposure quantification. GDPR maximum penalties are €20 million or 4% of global annual turnover. EU AI Act maximum penalties for prohibited practices are €35 million or 7% of global annual turnover. These are not theoretical risk register entries — they are the ceiling of a realistic exposure range that needs to be calibrated to the actual state of the AI programme. If the data flow mapping hasn't been done, the DPIA hasn't been completed, and the works council approval is pending, the probability weighting on that exposure is not low. The CFO owns putting numbers on this — not the CISO, not legal, not the DPO. AI vendor contract review. Standard commercial AI agreements were not written with GDPR in mind. The CFO needs to ensure that data processing agreements are in place, standard contractual clauses are signed, subprocessor lists are contractually controlled, and exit clauses cover data deletion obligations. A procurement team that signs a vendor agreement without these terms in place creates financial exposure that sits on the CFO's balance sheet. The commercial relationship is a CFO ownership item, and the data protection terms are part of the commercial relationship. Business case accuracy. Any AI business case that assumes data flows freely across borders — without accounting for anonymisation costs, pseudonymisation infrastructure, DPA requirements, DPIA costs, or potential remediation — is financially incomplete. The compliance cost is part of the programme cost. A business case that ignores it doesn't become compliant because the numbers look better without the cost line; it becomes a business case that will need to be revised when the cost materialises. The CFO who approves the business case owns the accuracy of what is in it. What the CTO must own The CTO owns the architecture decisions that determine the organisation's data protection posture — and those decisions are made once, at design, with consequences that persist for the life of the system. Architecture decisions with data residency and jurisdiction implications. Choosing between a US-hosted foundation model API and an EU-hosted alternative is not a technical preference. It determines the transfer mechanism burden, the US CLOUD Act exposure profile, and the complexity of the compliance programme that must sit around it. The CTO needs to make this choice with the data protection implications explicitly on the table — not as an afterthought following a performance or cost decision. Once the architecture is built on a US-hosted model, the cross-border transfer problem is baked in. Pseudonymisation infrastructure. For AI workloads where the organisation is processing personal data on external models, pseudonymisation before data leaves the organisation is the most practical risk reduction available. Building and maintaining that infrastructure — the pseudonymisation pipeline, the on-premise storage of re-identification keys, the audit of what is and isn't pseudonymised before transmission — is a CTO ownership item. It is also an architecture decision made at design. Retrofitting a pseudonymisation layer onto a production AI system that was built without one is significantly more expensive than building it in. Membership inference and re-identification risk assessment. For AI systems trained or fine-tuned on internal data, the CTO needs to understand what patterns the model could learn that allow indirect re-identification of individuals in the training set — and what technical controls reduce that risk. This is not a post-deployment question. It is an architecture question that needs to be answered before the model is built, because the answer changes what data goes into the training set and how the model is structured. I've seen organisations discover a re-identification risk in production that would have been addressed by a different data pipeline choice made six months earlier. Connecting the three ownership areas The three ownership areas are not independent. They create dependencies that need to be coordinated. The CTO's architecture decision determines what transfer mechanism the CFO needs to account for in the vendor contracts. The CIO's data classification determines what data the CTO can include in AI training sets. The CFO's financial exposure quantification depends on knowing what the CIO's vendor due diligence has found and what residual risks the CTO's architecture choices carry. The connection point is a shared risk assessment that all three own as inputs. In practice, this means one document — not three separate function-level risk assessments — that maps the data flows, the architecture, the contractual protections, the accepted residual risks, and the financial exposure. The DPO reviews it. Legal reviews it. But the three executives who own the inputs own the conclusions. The audit trail the DSK guidance requires is effectively this document — maintained, updated, and available for regulatory review at any time. What to take from thisAssign named executive ownership before the AI programme starts. Data classification and audit trail (CIO), financial exposure and vendor contracts (CFO), architecture and technical controls (CTO) — with a named person, not a function, accountable for each. The DPO advises and audits — the DPO does not own the programme. Data protection programme ownership sits with operational executives because the programme is operational. The DPO's role is to assess compliance, not to run the systems being assessed. Build the shared risk assessment document before deployment. One document mapping data flows, architecture, contractual protections, accepted residual risks, and financial exposure — owned jointly by CIO, CFO, and CTO, reviewed by DPO and legal. This is the audit trail regulators will look for. Recalibrate the risk register to reflect AI Act exposure. The penalty ceiling at 7% of global annual turnover for prohibited practices means the financial exposure calculation has changed. Update the probability weightings to reflect the actual state of the AI programme — not a theoretical best-case assessment. For German operations: works council engagement is CIO-owned from the moment a tool is evaluated, not from the moment a rollout date is set. The approval process cannot be compressed into the final weeks of a deployment project. Review the business cases for AI programmes already approved against the actual compliance cost. Programmes approved without accounting for pseudonymisation infrastructure, DPA requirements, or DPIA costs were approved on incomplete financial models. The cost will materialise — the question is whether it appears as a planned expense or an unplanned remediation.The executives who have closed the AI data protection gap in their organisations didn't do it because they understood the regulation better than their peers. They did it because someone in the C-suite owned the outcome. The regulation tells you what is required. The accountability structure determines whether it happens. A compliance review that ends with a report to the board but no named owner for each obligation is not a programme — it is documentation of intent. And intent, when examined by a regulator, is not a defence.
Read full article
- 30 Jun, 2026
Building a Privacy-Safe AI Pipeline: What Goes Between Your Data and the Model
Most European companies running AI programmes have a data protection policy that says something like: "personal data should not be sent to external AI models without appropriate controls." What most of those same companies don't have is a technical architecture that enforces this — at the point where data actually leaves the organisation, before it reaches the model, in a way that doesn't depend on every employee making the right decision every time. Policy without technical enforcement is an aspiration. I've seen teams build careful data governance frameworks, train employees on what can and cannot enter a prompt, and roll out usage guidelines. Then I've seen the same teams, six months later, discover that customer records, HR files, and contract documents had been flowing into an external model because the work was faster that way and nobody had built anything to stop it. The technical answer is a processing layer between your data and the model — a pipeline that intercepts data before it reaches the LLM, applies privacy controls, and only passes through what meets the organisation's data protection requirements. Building it is an architecture decision, not a compliance decision. This article covers what that architecture looks like, the components that make it work, and where its limits are. Why policy-only approaches fail at the prompt The problem is structural. Employees use AI tools to get work done faster. The faster path almost always involves sending more context — a full customer record rather than selected fields, a complete document rather than the relevant section, a full email thread rather than the key paragraph. The model produces better outputs with richer input, so the incentive runs directly against the minimisation principle. Enforcing minimisation through policy means asking people to voluntarily produce worse outputs to protect data they don't directly manage. That is a losing proposition. The violation isn't usually malicious — it is the product of a rational trade-off that the person making the decision doesn't see as a trade-off at all. Architectural enforcement removes the trade-off. The pipeline intercepts data before it reaches the model and applies one of two controls: stripping personal data entirely where the use case doesn't require specific identity references, or replacing identifiers with on-premise-keyed tokens where the model genuinely needs to track the same entity across a document. In both cases, what reaches the model contains no original personal data. This is the design goal: make the compliant path identical to the convenient path. But it is important to be precise about what that compliance achieves. Neither stripping nor tokenisation makes data anonymous under GDPR. Stripping is the stronger control — no key exists, no reversal is possible — and it is what the pipeline should default to. Tokenisation is a fallback for use cases that require entity consistency, and tokenised data remains personal data under GDPR, as covered in detail in the anonymisation article in this series. The pipeline reduces risk significantly. It does not eliminate the organisation's GDPR obligations for what reaches the model. The core architecture: a processing layer the LLM never bypasses The pipeline has five logical components. All personal data must pass through all five before reaching the model. The LLM has no direct access to raw data — only to what the pipeline has processed and approved. Data ingestion: Documents, database records, API responses, or user prompts enter the pipeline. This is the boundary where classification begins — what type of data is this, what sensitivity level does it carry, and what processing rules apply. PII detection: A detection layer identifies personal data in the ingested content. This covers direct identifiers (names, email addresses, phone numbers, national identification numbers, passport numbers, IBAN and financial account numbers, IP addresses, biometric identifiers) and indirect identifiers that could contribute to re-identification (age combined with location, job title combined with department size, medical codes combined with demographic data). Detection uses a combination of Named Entity Recognition (NER) models and pattern-matching rules. NER models identify contextual entities — a name in a sentence, a location in a paragraph. Pattern-matching rules catch structured identifiers — a correctly formatted IBAN, an email address, a German Steuernummer. PII removal and tokenisation layer: This is the core control, and it operates in two modes depending on what the task requires. Stripping (default): Detected identifiers are removed entirely. A name disappears. An email address is deleted. No key is created, no mapping is stored, no reversal is possible. This is the stronger control and should be the default for use cases where specific identity is not required — summarisation, document classification, general drafting, sentiment analysis. Stripped data is the closest the pipeline can bring data to the GDPR anonymisation standard, though it still may not meet the full legal threshold if indirect re-identification from remaining attributes remains possible. Tokenisation (exception): Where the AI task genuinely requires consistent entity references across a document — contract analysis referencing the same counterparty throughout, an HR review tracking a specific employee's record — detected identifiers are replaced with consistent tokens (PERSON_001, EMAIL_001, ORG_002). The token mapping is stored in a key vault that lives entirely on-premise and is never transmitted. The model sees consistent references without knowing the real identity. This is pseudonymisation. Pseudonymised data is still personal data under GDPR. The pipeline's GDPR obligations — transfer mechanisms, DPA requirements, breach notification — apply to any tokenised content that reaches the external model. LLM processing: The stripped or tokenised content is sent to the external model. No original personal data is present. The model has no access to the key vault. Output de-tokenisation (tokenisation mode only): The model's response is returned to the pipeline. Where tokenisation was applied, the pipeline replaces tokens with the original identifiers using the on-premise key vault before delivery to the user. The key vault never left the organisation. What the vendor receives, what any external disclosure would expose, and what any CLOUD Act order could compel them to produce is the stripped or tokenised version. The reduction in exposure is real. It is not, however, a GDPR exemption. Building the PII detection layer The detection layer is where most implementations have their weakest point. NER models trained on general text perform well on common identifiers — names, locations, organisations in standard formats. They perform poorly on domain-specific identifiers, non-standard formats, and contextual indirect identifiers that require understanding of the data environment. For enterprise deployments, the detection layer needs to be tuned for the organisation's specific data. A healthcare provider will have medical record identifiers that a general NER model won't recognise. A financial services firm will have account reference formats specific to its systems. A logistics company will have shipment identifiers tied to individual customers. Microsoft Presidio is an open-source PII detection and anonymisation framework that provides a base NER capability with extensible recogniser architecture — custom recognisers can be added for organisation-specific identifier formats. spaCy provides the underlying NER model that can be fine-tuned on domain data. GLiNER is a more recent generalist NER model that performs well on zero-shot entity recognition, useful for catching entity types that weren't anticipated in the original configuration. The detection layer should be treated as imperfect by design. False negatives — personal data that is not detected and therefore passes through to the model as-is — will occur. The residual risk from false negatives needs to be part of the organisation's data protection impact assessment. For high-sensitivity data categories (health data, financial data, data on minors), a higher detection threshold or additional validation step is warranted. For lower-sensitivity data, the residual risk from occasional false negatives may be acceptable. Logging is essential. Every item that passes through the pipeline should generate a log entry: what data type was detected, what action was taken (stripped, tokenised, blocked), what was transmitted to the model, and when. This log is the audit trail that GDPR and the June 2025 DSK guidance require. It also lets the organisation identify patterns where specific data types are consistently slipping through the detection layer. Synthetic data as an alternative for AI development For organisations building, fine-tuning, or testing AI models on internal data, the pipeline approach described above is a runtime control — it operates on data flowing to the model during inference. A different problem arises at development time: the organisation needs representative data to develop, test, and fine-tune AI systems, and real personal data carries risk even in development environments. Synthetic data generation creates datasets that have the same statistical properties as real data without containing records about real individuals. For a customer dataset, the synthetic version preserves the distribution of ages, purchase patterns, geographic spread, and behavioural signals — but every record is generated, not derived from an actual customer. The GDPR position on genuinely synthetic data is more favourable than on pseudonymised data, because data generated from statistical distributions rather than from individual records may qualify as anonymous if the generation process cannot be reversed to recover the underlying individuals. This requires that the generative process genuinely produces new data rather than perturbing or recombining existing records — a meaningful technical distinction that affects both the data's legal status and its usefulness as a development dataset. For model testing and development environments, synthetic data removes the need to apply runtime privacy controls and reduces the risk that development data leaks through version control, shared notebooks, or development tooling. For production fine-tuning, synthetic data has quality limitations — complex behavioural patterns and rare edge cases are difficult to preserve through synthesis — but it is a viable alternative for many use cases where the fine-tuning goal is improving general performance rather than capturing rare behaviours. What this architecture doesn't solve The pipeline addresses the most direct GDPR exposure in AI: personal data flowing to an external model without appropriate controls. It does not address all exposure. Usage metadata is not affected. The vendor still sees query frequency, document structure patterns, workflow characteristics, and feature usage — the competitive intelligence risk described in the first article in this series. The pipeline controls what is in the data; it doesn't control that data is being sent. Transfer mechanisms still apply. Pseudonymised data is still personal data under GDPR. The pipeline reduces the sensitivity and impact of a disclosure event, but it does not remove the organisation's obligation to have a valid cross-border transfer mechanism in place for the pseudonymised content that does reach the external model. Fine-tuning risk remains for the training phase. If the organisation fine-tunes a model on pseudonymised data, membership inference risk applies to the trained model — the model may encode patterns that allow indirect re-identification, as covered in the anonymisation article in this series. The runtime pipeline prevents personal data from reaching the model during inference. It does not address patterns the model has already learned from training data. Homomorphic encryption is the emerging technical approach that would allow data to be processed by the model while remaining encrypted throughout — the model would never see plaintext data at any point. This would represent a fundamentally stronger privacy guarantee than pseudonymisation. In practice, homomorphic encryption adds significant computational overhead and is not yet deployable at enterprise scale for LLM inference. It is being actively researched and is recognised in EU policy frameworks as a potential compliance tool for the medium term. It is not a practical option for most enterprise AI programmes today. What to take from thisBuild the privacy pipeline before deploying AI to end users. The decision to deploy an AI interface to employees without technical privacy controls in place is a decision to rely on policy enforcement alone — which consistently fails. The pipeline is the enforcement mechanism. Start PII detection with a general-purpose NER model and extend it with custom recognisers for your organisation's specific identifier formats. General models will miss domain-specific and proprietary identifiers. Tuning for your data is required, not optional. Default to stripping, not tokenisation. Stripping removes PII entirely — no key, no reversal, no personal data transmitted. Use tokenisation only when the AI task genuinely requires entity consistency across a document, and treat tokenised data as personal data throughout, because under GDPR it is. Where tokenisation is applied, the key vault must live on-premise, isolated from all network-facing components, and never transmitted. If the key vault is in the same cloud environment as the model API calls, the architectural separation is broken. Treat the detection layer as imperfect and log everything. Build the DPIA assumption that some personal data will pass through undetected. Log all pipeline activity so you can identify detection gaps and demonstrate the audit trail regulators will ask for. Use synthetic data for AI development and testing environments. Remove real personal data from development pipelines entirely where possible — synthetic data preserves enough statistical structure for most development use cases while eliminating the exposure. Understand what the pipeline does and doesn't cover. Transfer mechanisms, metadata exposure, and fine-tuning data risk sit outside what a runtime pseudonymisation pipeline addresses. The pipeline closes the most direct exposure. The other risks need separate controls.The gap between "we have a data protection policy for AI" and "our AI systems are technically enforcing that policy" is where most European enterprises currently sit. Policy documents don't intercept a prompt. Architecture does. The pipeline is not complex to build in its basic form — a PII detection layer, a tokeniser, and an on-premise key vault are available as open-source components or commercial services. What requires investment is tuning detection for your specific data environment, integrating the pipeline into every AI touchpoint in the organisation, and maintaining it as the AI programme grows. That investment is also what makes the compliance position defensible when examined — because a working technical control is a fundamentally different answer to a regulatory question than a policy that assumed people would follow it.
Read full article
- 29 Jun, 2026
EU AI Act Enforcement: What European Companies Must Already Comply With Right Now
Something changed on 2 February 2025 that most European executives missed entirely. A category of AI practices became prohibited and enforceable under the EU AI Act. Not in the future. Not subject to a transitional period. Enforceable on that date. Organisations running AI systems that fell into the prohibited category were non-compliant from that point forward — not approaching non-compliance, not at risk of future non-compliance, but in breach of an enforceable prohibition. Most European executives I work with treat the AI Act as a 2026 problem. Some treat it as a 2027 problem, given that the Digital Omnibus agreement of May 2026 deferred the full obligations for certain high-risk system categories to December 2027. The confusion is understandable — the regulation has a staggered enforcement schedule, and media coverage has focused primarily on the August 2026 milestone for high-risk systems. But the staggered schedule doesn't mean nothing is in force. It means different obligations apply at different points in time. And the first milestone passed eighteen months ago. The enforcement timeline — what is live, what is coming The AI Act entered into force on 1 August 2024. From that date, the regulation exists and creates obligations. The question is which obligations are enforceable when. 2 February 2025: Prohibited AI practices became enforceable. Organisations using AI systems in the prohibited categories have been in breach since this date. 2 August 2025: Rules governing General Purpose AI (GPAI) models became enforceable. This applies primarily to providers of foundation models — the companies building the underlying AI infrastructure. For most European enterprises that are deployers rather than builders, this milestone creates indirect obligations through the requirement to use GPAI providers who themselves comply. 2 August 2026: Full high-risk AI system obligations apply for most categories under the Act. 2 December 2027: Annex III high-risk systems, deferred under the Digital Omnibus agreement of May 2026. This is the deferred deadline that has led some organisations to extend their compliance timelines across the board — which is a misreading of the agreement. The deferral applies to Annex III systems specifically, not to the full regulation. For European enterprises that are AI deployers — using third-party AI systems rather than building foundation models — the immediate exposure is the prohibited practices that have been enforceable since February 2025, and the high-risk system obligations that arrive in full in August 2026. The prohibited practices — enforceable now The banned categories under the AI Act cover AI applications that the European legislature determined are fundamentally incompatible with European values and fundamental rights, regardless of the purpose they serve. Understanding them is important because enforcement has been live for over a year, and because some of these systems appear in enterprise AI programmes under different names. Social scoring systems are prohibited: AI that evaluates individuals based on their social behaviour or personal characteristics and applies consequences to them in contexts unrelated to where the data was collected. Enterprise "employee engagement scoring" or "social collaboration analytics" tools that rate individuals and feed those ratings into decisions about career progression, training access, or benefit eligibility need to be assessed against this definition. Systems that exploit psychological vulnerabilities to manipulate behaviour are prohibited. This includes AI that exploits age, disability, or specific psychological susceptibilities to influence individuals in ways that harm them. Marketing AI that targets individuals based on inferred vulnerability profiles sits in this category. Emotion recognition in the workplace is prohibited. AI systems that infer the emotional state of employees from facial expressions, voice tone, physiological signals, or other biometric data — for any employment-related purpose — are banned. This includes tools marketed as "engagement measurement," "meeting analytics," "performance monitoring," or "wellbeing tracking" if they incorporate emotional inference. Any European company currently running such systems has been operating a prohibited AI practice since February 2025. Real-time biometric identification in public spaces by law enforcement is prohibited, with narrow exceptions. This has limited direct application for most enterprises but is relevant for companies operating public-facing AI surveillance systems. The assessment question for any European enterprise is not "does this tool do what the label says?" It is "what does this tool actually do with the data, and does that function fall within a prohibited category?" The label is marketing. The function is what the Act governs. The high-risk obligations arriving in August 2026 High-risk AI categories under the Act include HR and recruitment AI, credit scoring and financial risk assessment, critical infrastructure management, educational AI that determines access or progression, and law enforcement tools. For a typical European enterprise, the most immediately relevant categories are HR AI and credit or risk scoring. A high-risk designation is not a prohibition. It is a full set of technical and governance obligations that must be in place before and during deployment: Technical documentation of the system must exist before deployment — a complete description of the system's purpose, design, training data, performance characteristics, limitations, and risk mitigation measures. A conformity assessment — either a self-assessment or, for certain categories, a third-party assessment — must be completed before the system goes into production. Human oversight mechanisms must be built into the system's operation. Automated decisions in high-risk categories cannot be final without a meaningful opportunity for human review. Meaningful means a human who has been given sufficient information to actually evaluate the decision — not a checkbox that routes through a queue. Post-market monitoring must be ongoing. The deployer must track how the system performs in real use, log outcomes, identify drift or unexpected behaviour, and have a process for reporting serious incidents. For European companies that have been deploying HR AI — tools that screen CVs, rank candidates, recommend candidates for rejection or progression, or support performance review — without this framework in place, the August 2026 deadline is not far off. Building the technical documentation, redesigning oversight mechanisms, and implementing post-market monitoring are not quick projects. They require architectural decisions made in advance. The four risk tiers and how to map an enterprise AI programme The Act organises AI systems into four tiers. Understanding which tier an AI tool sits in determines what obligations apply. Unacceptable risk (prohibited): Banned outright. Enforceable since February 2025. If any tool in your programme sits here, it needs to stop, not be redesigned. High risk: Full documentation, conformity assessment, human oversight, and post-market monitoring required. Most HR AI, credit scoring AI, and similar tools sit in this tier if they produce decisions with significant effects on individuals. Limited risk: Transparency obligations apply. Users must be informed when they are interacting with an AI system. Chatbots, customer service AI, and AI-generated content tools typically sit here. The obligation is disclosure, not redesign. Minimal risk: No specific obligations beyond general GDPR. Most AI tools — recommendation systems, spam filters, productivity tools that don't make decisions about individuals — sit here. A well-run mapping exercise across a European enterprise's AI programme will typically find: a large number of tools in the minimal or limited risk tier; some tools in or approaching the high-risk tier, particularly in HR, compliance, and finance functions; and — in organisations that have deployed AI broadly without systematic oversight — a reasonable chance of finding a tool that functions in ways that fall within or near the prohibited category. The mapping is not a one-time document. AI systems change. Vendors update their models. A tool that was minimal risk in 2024 may be operating differently in 2026. The mapping needs to be a living assessment. What to take from thisAudit your AI programme against the prohibited practices list immediately. If any system involves emotion recognition in the workplace, social scoring, or exploiting psychological vulnerabilities to influence behaviour, it has been operating in breach since February 2025. Stop first, then assess. Map every AI tool against the four risk tiers. This is a one-time exercise that becomes a living document. Do it across the full programme — not just the tools the CIO knows about, but the tools individual functions have procured and deployed without central oversight. For high-risk systems — particularly HR AI and any credit or risk scoring tools — start the documentation and oversight architecture now, not in 2026. Technical documentation, conformity assessment, and human oversight mechanisms require architectural decisions made before deployment, not added retrospectively. Put AI Act penalty exposure in the board-level risk register with realistic probability weightings. The €35 million or 7% ceiling for prohibited practices is enforceable today. Complaint-driven enforcement from affected individuals is the most likely trigger. Confirm that deployer obligations under Article 26 are addressed in your vendor contracts. Using a third-party model provider doesn't remove the deploying company's compliance obligations. The vendor's compliance doesn't substitute for the deployer's. Review the Digital Omnibus deferral carefully. The deferral to December 2027 applies to Annex III systems specifically. It does not apply to prohibited practices (enforceable since February 2025), GPAI rules (enforceable since August 2025), or the majority of high-risk system obligations that arrive in August 2026.The AI Act is a regulation with a staggered enforcement schedule, and that staggering has created the conditions for a comfortable deferral of the compliance conversation. The problem is that staggering means different things apply at different times — not that nothing applies yet. The organisations currently treating the AI Act as a future problem have already missed the February 2025 milestone. The question is how many will realise that before a complaint does.
Read full article
- 26 Jun, 2026
Data Anonymisation for AI: Why the GDPR Standard Is Harder Than European Companies Think
A European data team prepares a dataset for processing by an external AI model. Before it goes, they run the standard anonymisation process: names removed, email addresses stripped, identification numbers deleted. The team confirms the data is anonymous. The dataset is sent. It isn't anonymous. What most organisations call anonymisation is pseudonymisation. The distinction matters — not as a legal technicality, but because GDPR treats them entirely differently, the actual standard for true anonymisation is far higher than most teams have achieved, and AI models introduce a specific re-identification risk that didn't exist in conventional data processing. European companies making data flow decisions based on an anonymisation claim they haven't actually met are carrying GDPR exposure they believe they've eliminated. The gap between the two terms is where most enterprise AI programmes are currently sitting. The difference between pseudonymisation and anonymisation — and why it changes everything Pseudonymisation replaces identifying fields with artificial identifiers. A name becomes a token. An ID number becomes a different ID number. The original identity can be restored using a key held by the originating organisation. The data looks de-identified. Under GDPR, it is explicitly classified as personal data. GDPR Recital 26 makes this clear: pseudonymised data — data that could be attributed to a natural person by the use of additional information — remains within the scope of the regulation. Cross-border transfers of pseudonymised data still require a valid transfer mechanism. Data subject rights still apply to it. Breach notification obligations apply if it is compromised along with the re-identification key. Pseudonymisation is valuable risk reduction — GDPR Article 32 explicitly recommends it as a technical security measure — but it does not remove data from GDPR scope. True anonymisation is the irreversible kind. Data is genuinely anonymous when re-identification is not reasonably possible, taking into account all the means likely to be used by any party. That is the GDPR standard. It is an extremely high bar. Truly anonymous data falls entirely outside GDPR scope — it can be transferred to AI models outside the EU without triggering the transfer restrictions in Chapter V, it carries no breach notification obligations, and it is not subject to data subject access rights. The problem is that meeting this standard for real-world enterprise datasets is genuinely difficult. And most organisations claiming anonymisation have not met it. Why the re-identification standard is harder than it appears Remove a name, remove an email address, remove a national identification number. The dataset looks clean. Now consider what remains: age, job title, department, location, salary band, tenure, performance rating, and a timestamp on when the record was created. For a company with three hundred employees in a single city, how many people does that description fit? Often fewer than five. Sometimes one. This is the re-identification problem. It doesn't require the re-identifier to have the original data. It requires them to have other data — which may be publicly available, commercially available, or deducible from the context — that allows them to match records back to individuals. The more detailed and rich the dataset, the smaller the group that any given record could describe, and the easier the re-identification. Study after study has demonstrated this in practice. Researchers have re-identified individuals in supposedly anonymous medical datasets by cross-referencing with public records. Anonymous mobility data has been linked back to named individuals by correlating location patterns with public social media. Netflix user records, released as part of an anonymisation research dataset, were re-identified by cross-referencing with public film ratings platforms. The European Data Protection Board is working on updated guidelines on anonymisation — this was announced in its 2024-2025 work programme. The updated guidance hasn't resolved the standard; it has reinforced the position that organisations asserting anonymisation should be able to demonstrate it, not assert it as a processing decision taken internally without external validation. For European enterprises, the practical implication is this: if the dataset is complex, if the records are detailed, if the organisation holds a re-identification key anywhere in its systems, or if the data could be cross-referenced with external sources to recover identities — the data is pseudonymous at best, not anonymous. And pseudonymous data requires all the GDPR protections that come with personal data. How AI models add a re-identification risk that didn't exist before Conventional data processing treats data as data. An analytics platform runs queries. A reporting system aggregates records. The data doesn't learn from its own processing. AI models are different. A model trained on pseudonymised data can, under certain conditions, learn patterns correlated with the identity of specific individuals — patterns derived not from the identifiers (which were removed) but from the combination of other attributes in the records. This creates what is known as a membership inference risk: an attacker — or the model itself, under specific prompting conditions — can determine whether a specific individual was present in the training set and, in some cases, infer characteristics of that individual's record. The model hasn't stored the record. It has encoded patterns from it during training. Those patterns are queryable. This is not a theoretical attack vector — membership inference has been demonstrated on real models trained on real datasets, and it is an active area of regulatory concern at both the EDPB and national supervisory authority level. For European companies, this creates a specific additional problem. An organisation might correctly conclude that sending pseudonymised data to an external processor for conventional analytics is acceptable. Sending the same pseudonymised data to an AI model for training or fine-tuning is a materially different decision — because the model learns from the data in ways a conventional processor doesn't, and those learned patterns can be queried in ways that don't look like data access requests. The re-identification key doesn't need to be compromised. The model output can, under the right conditions, function as a partial re-identification mechanism — without any key ever leaving the organisation. What the November 2025 ECJ ruling actually says — and what it doesn't The European Court of Justice issued a ruling in November 2025 that clarified a specific scenario: pseudonymised data held by a recipient who does not have access to the re-identification key, and who cannot reasonably obtain it, may fall outside GDPR scope from that recipient's perspective. This ruling has been widely reported as a significant relaxation of the pseudonymisation standard. It is useful clarification. It is not the broad clearance it has sometimes been described as. The ruling applies to the recipient's position, not the sender's. The organisation sending pseudonymised data still holds the re-identification key. From the sending organisation's perspective, it is still processing personal data, and the transfer still requires a valid mechanism under GDPR Chapter V. The question of whether the recipient — the AI vendor — is processing personal data is a separate analysis that requires examining whether the vendor can reasonably obtain or derive the means to re-identify individuals. For major AI vendors with access to large external datasets, broad internet data, and models capable of cross-referencing patterns at scale, answering that question is not straightforward. A vendor that cannot access the specific organisation's re-identification key may nonetheless be capable of re-identifying individuals by cross-referencing the pseudonymised data with other sources in their possession. The ECJ ruling doesn't provide clearance in that scenario — it requires case-by-case analysis. The safe position for European companies is not to cite the November 2025 ruling as a blanket clearance for pseudonymised AI transfers. It is to conduct the case-by-case analysis the ruling prescribes — which requires understanding what external data the AI vendor has access to, what their re-identification capabilities are, and whether the specific pseudonymised dataset could reasonably be linked back to identifiable individuals by someone in the vendor's position. What this means in practice for European AI programmes The path forward is not to avoid AI models or to conclude that no data can be transferred. It is to be accurate about what has and hasn't been achieved. Pseudonymisation before data leaves the organisation is the correct technical control for reducing exposure when processing data on external AI models. It should be implemented properly — direct identifiers removed, re-identification key held on-premise and not transmitted, pseudonymisation applied consistently and auditably. It significantly reduces the impact of any disclosure event. It reduces but does not eliminate re-identification risk. What it does not do is remove the organisation's GDPR obligations. The transfer still requires a valid mechanism. The data still counts as personal data for GDPR purposes. The organisation still needs a lawful basis for the processing. These obligations persist alongside the pseudonymisation. For fine-tuning or training AI models on internal data, the additional membership inference risk needs to be part of the architecture decision — not assessed after the model is in production. The question of what personal data is in the training dataset, what patterns an attacker could derive from it, and what technical controls reduce the inference risk is an architecture question, not a post-deployment compliance question. What to take from thisAudit what your organisation actually does when it calls data "anonymous" before sending it to an AI model. In most cases, direct identifiers have been removed but the data remains pseudonymous. Update the data flow documentation accordingly — GDPR obligations follow from the accurate classification, not the assumed one. Treat pseudonymisation as risk reduction, not GDPR elimination. It is valuable and recommended under Article 32. It does not take data out of GDPR scope. Cross-border transfer mechanisms still apply. Do not cite the November 2025 ECJ ruling as a blanket clearance. It applies to recipients who genuinely cannot re-identify individuals. Establish whether your AI vendor meets that condition before relying on it — because a vendor with access to large external datasets may not. Include membership inference risk in AI architecture reviews. For any model trained or fine-tuned on internal data containing personal information, assess whether the model could be queried in ways that allow re-identification of individuals in the training set. Do this before the model is built, not after it's running. Keep the re-identification key on-premise. This sounds obvious. In practice, I have seen organisations pseudonymise data and then include a lookup table in the data package sent to the vendor — which converts pseudonymisation into a form of weak encryption rather than a genuine privacy control. Document the anonymisation or pseudonymisation process with enough detail to support a regulatory assessment. If a supervisory authority asks how the data was anonymised, the answer needs to be a documented process, not a verbal account of what the data team did six months ago.The word "anonymisation" has been in data governance policies for years as shorthand for "we processed the data carefully before sharing it." That shorthand has never fully survived GDPR scrutiny. It survives even less well when the sharing is with an AI model that learns from what it receives. The question isn't whether the organisation intended to comply with the regulation. The question is whether the data classification was accurate — and for most European enterprises running AI programmes, the honest answer is that it hasn't been tested against the standard that actually applies.
Read full article
- 25 Jun, 2026
Sending European Data to US AI Models: The Risks That Outlast the Contract
The legal team signed off. The data processing agreement is in place. Standard contractual clauses are attached to the vendor agreement. The European company is operating under the EU-US Data Privacy Framework. From a compliance documentation standpoint, the boxes are ticked. Three specific risks persist that none of those documents fix. This isn't an argument against using external AI models. European companies are running AI programmes on US-hosted infrastructure because those models are currently the most capable available, and the business case for using them is real. The argument here is narrower: the compliance documentation that most European companies have in place for cross-border AI data transfers is necessary but not sufficient. Understanding what it doesn't cover — and why — is the starting point for an honest risk assessment. What signing the contract actually achieves Standard contractual clauses (SCCs) create binding obligations on the data importer — the US-based AI vendor receiving European data. The vendor commits to handling the data in accordance with GDPR requirements, to notifying the European company of any requests from authorities, and to applying appropriate technical and organisational security measures. This is genuine legal protection. It creates enforceable obligations. It is also the minimum requirement for a cross-border transfer to be lawful under GDPR, not the complete solution. What SCCs do not do: they do not change the jurisdiction of the vendor's infrastructure. They do not change the legal obligations a US-incorporated company is subject to under US law. A US-incorporated AI vendor that receives a lawful demand from US law enforcement is subject to that demand whether or not it has signed an SCC with every European customer it serves. The SCC creates a contractual conflict — the vendor may be obliged to notify the European customer and resist disclosure where legally possible. It does not create a legal mechanism that overrides US law. The data processing agreement operates in the same space. It defines how the vendor handles the data, what it can do with it, what happens in a breach. It is a contract between two private parties. It cannot bind US law enforcement. The political fragility underneath current EU-US transfers The EU-US Data Privacy Framework (DPF) is the current mechanism that provides an adequacy basis for transfers of personal data from the EU to US companies that have self-certified under the framework. Most major US AI vendors have self-certified. This is what allows European companies to transfer data to those vendors without relying solely on SCCs. The DPF was adopted in July 2023. It is the third iteration of a transatlantic data transfer agreement. The first two — Safe Harbor and Privacy Shield — were both invalidated by the European Court of Justice, in 2015 and 2020 respectively. The conditions that produced those rulings — US surveillance laws, specifically FISA Section 702 and Executive Order 12333, which permit broad data collection on non-US persons — have not materially changed. The DPF was adopted alongside a US executive order (EO 14086) that introduced new redress mechanisms for EU data subjects, but whether those mechanisms are sufficient to satisfy the ECJ's adequacy requirements remains contested. A challenge to the DPF is already working through the European legal system. Privacy rights organisations have filed complaints in several EU member states. The pattern from Safe Harbor and Privacy Shield — challenge filed, referred to ECJ, adequacy decision struck down — is a known sequence. European companies whose entire cross-border transfer strategy depends on the DPF holding are carrying a risk they may not have quantified. The practical position: SCCs should be in place independently of the DPF, as a fallback mechanism. But SCCs alone don't resolve the structural issues the ECJ has identified in prior rulings — which brings us to the third problem. The US CLOUD Act: a structural gap no SCC closes The Clarifying Lawful Overseas Use of Data Act (CLOUD Act), enacted in 2018, allows US law enforcement to compel US-incorporated companies to produce data stored on their servers — regardless of where that data is physically stored. A US-based AI vendor with a data centre in Frankfurt is still subject to a CLOUD Act order for data held in that Frankfurt facility, because the obligation runs to the company, not to the data's location. This is not a theoretical risk or an unlikely scenario. It is the structural architecture of US data law. European companies that have negotiated data residency commitments from their US AI vendors — contractual guarantees that data stays in EU data centres — have obtained a meaningful operational control and a genuine security benefit. They have not obtained protection from CLOUD Act orders. SCCs require vendors to notify the European customer of government access requests and to resist such requests where legally possible. In practice, where US law enforcement has obtained a valid CLOUD Act order, the vendor's ability to resist is limited and the notification may be subject to a gag order that prevents it. The SCC provision exists; the real-world protection it provides against a valid CLOUD Act order is narrow. The only structural approach that avoids CLOUD Act exposure entirely is infrastructure that is not subject to US jurisdiction — EU-incorporated companies, with EU-domiciled infrastructure, and no US parent company with control over the entity. Several European AI infrastructure providers market this explicitly as a selling point for regulated industries. It is a legitimate differentiator for use cases where the CLOUD Act risk is assessed as unacceptable. What actually reduces exposure — and what just creates paperwork For European companies that need the capability of major US-hosted foundation models and have concluded the business case outweighs the residual risk, the realistic risk reduction available is a layered approach, not a single mechanism. The most practical technical control is pseudonymisation before data leaves the organisation. If the data reaching the AI model doesn't contain directly identifying information — names, identification numbers, email addresses, and other direct identifiers have been replaced with tokens, and the re-identification key stays on-premise — then what the vendor receives, and what any CLOUD Act order would compel them to produce, is pseudonymised data. Pseudonymised data is still personal data under GDPR, so this doesn't eliminate GDPR obligations. But it significantly reduces the impact of a disclosure event. Combined with this: SCCs in place as a fallback, a documented assessment of the DPF's continued validity as a transfer mechanism, a clear record of what data categories are permitted to enter the AI model and which are not, and a written acceptance of the residual CLOUD Act risk at the appropriate seniority level in the organisation. That last point matters. The risk assessment is not complete until someone with authority has reviewed and signed off on the residual risk — not because sign-off eliminates the risk, but because undocumented accepted risks are a different regulatory problem from documented ones. What doesn't reduce exposure: signing the SCC addendum in the vendor's standard contract pack without reading it; obtaining a data residency commitment without understanding it doesn't address CLOUD Act; and treating DPF certification as equivalent to a permanent adequacy finding. These create documentation. They don't close the exposure. What to take from thisHave SCCs in place as a baseline, but treat them as necessary rather than sufficient. Understand what they actually obligate the vendor to do and where US law creates limits on what those obligations can achieve. Assess the EU-US Data Privacy Framework as a risk position, not a permanent mechanism. Build SCCs as a fallback and document the contingency position for the scenario where the DPF is challenged successfully — which is not an unlikely scenario given the historical pattern. Understand that data residency commitments from US vendors do not provide CLOUD Act protection. A contractual guarantee that data stays in EU infrastructure is a meaningful operational control; it is not a jurisdictional change. Implement pseudonymisation before data exits the organisation for AI processing. This is the most practical technical control available for reducing the impact of any disclosure event — whether that's a CLOUD Act order, a vendor breach, or a DPF invalidation. Document the accepted residual risk at the appropriate executive level. A risk that has been assessed, quantified, and signed off by someone with authority is a different compliance position from a risk that has been ignored. Regulators distinguish between the two. For use cases where CLOUD Act exposure is genuinely unacceptable — particularly in regulated industries, defence-adjacent work, or cases involving highly sensitive personal data — evaluate EU-incorporated AI infrastructure providers as the structural answer, not a contractual workaround.Cross-border data transfers in AI are a live compliance question with no clean answers at present. The honest position is that European companies using US-hosted AI models are operating in a legal environment that is less stable than the documentation in their procurement files suggests. That's not a reason to stop — it's a reason to understand the actual risk clearly, document what has been accepted and why, and hold a position that survives scrutiny when examined. The organisations that have done that work are in a defensible position. The ones that assumed the vendor's standard agreement covered everything are not.
Read full article
- 24 Jun, 2026
European Company Data in US AI Models: What Your Executives Don't Know Is Already Out
Six months after a European CTO signed off on an AI programme, the DPO sent a single-line email: which data has been sent to the model, where is it stored, and how long is it kept there? Nobody had an answer. The programme had been running on a standard enterprise API agreement. The vendor was reputable. The legal team had reviewed the terms. But the question of what data was leaving the organisation — in what form, to which infrastructure, under what retention policy — had never been mapped. I've seen versions of this conversation more times than I'd like. The AI programme gets approved at the business case level. The vendor agreement goes through procurement. The DPO gets looped in late, or not at all. By the time anyone asks the data flow question seriously, the model has been running for months and the answer is complicated. This isn't a scare piece. It's a map of what actually leaves a European organisation when it runs AI workloads on external models — because the gap between what executives believe is happening and what the contract actually governs is wider than most teams realise. GDPR has been in force since 2018. The enforcement precedent in the AI space is now live. And the organisations that haven't done the data flow mapping are running programmes that haven't been assessed for compliance. What enters a prompt has left the organisation Every query sent to an AI model — every instruction, every document, every piece of text — is transmitted to and processed on the vendor's infrastructure. This is obvious in principle and consistently underestimated in practice. The problem isn't that employees use AI tools. The problem is that the data flowing through those tools hasn't been classified before it flows. When a salesperson pastes a customer contract into a prompt to generate a summary, that contract has left the organisation. When an HR manager uploads a performance review to ask for a draft response, that review has left the organisation. When a finance analyst feeds a quarterly forecast into a model to restructure the narrative, that forecast has left the organisation. None of this requires a policy violation. It happens in ordinary, productive use — which is exactly why it's difficult to control retrospectively. Most enterprise AI APIs retain prompt data for abuse monitoring purposes unless the customer has explicitly opted out — and opted out in writing, in a way the contract records. Retention periods vary by vendor. They are subject to change. I have reviewed enterprise AI agreements where the data retention section was three lines long and referred to a separate policy document that had been updated four times in eighteen months. The business owner who approved the tool had never read either version. Data classification before any AI tool goes into production is not optional — it is the prerequisite. Without knowing what data is sensitive, personal, confidential, and commercially restricted, there is no basis for deciding what can enter a prompt. Most European enterprises have data classification policies. Most of those policies were written before AI prompt data was a category that needed to be governed. Personal data in prompts — minimisation is not enough When the data entering a prompt includes names, email addresses, identification numbers, health records, or financial information, the transfer isn't just a data governance concern — it is a GDPR event. Personal data has moved to an external processor, potentially outside the EU, under whatever transfer mechanism (or lack of one) applies to that vendor relationship. The standard response inside most European enterprises is to invoke data minimisation: GDPR's Article 5(1)(c) principle that personal data should be adequate, relevant, and limited to what is necessary. In theory, this means sending only the fields required for the AI task. In practice, it almost never works that way. Teams send full customer records because the model produces better outputs with more context. HR workflows send complete personnel files because partial information leads to incomplete responses. Customer service AI is fed entire interaction histories because isolated records produce generic answers. The business logic for sending more is always compelling, and the minimisation principle gets applied at the policy level while the actual data flows ignore it. Even when minimisation is applied properly, it is not a compliance answer for personal data. A dataset that contains only a customer's name, email address, and account status is minimal — and it is still fully personal data under GDPR. Minimising what you send does not change the legal classification of what arrives at the model. The transfer still requires a lawful basis. The vendor still needs to be a compliant data processor. The cross-border transfer rules still apply. The only way personal data falls outside GDPR scope when it reaches an external AI model is if it has been genuinely anonymised — not minimised, not pseudonymised, but anonymised to the standard where re-identification is not reasonably possible for anyone, including the model provider. That standard is far harder to meet than most teams assume, and AI models introduce a specific re-identification risk through pattern learning that conventional processing does not. This is covered in full in a separate post in this series. The practical position for European enterprises: names, email addresses, and any other direct identifiers should not enter an external AI model unless a valid GDPR lawful basis, a compliant data processing agreement, and a valid cross-border transfer mechanism are all in place — and verified, not assumed. The metadata picture your vendor is accumulating Prompt content is the obvious exposure. Usage metadata is the one that organisations consistently miss. AI vendors collect usage patterns: query frequency, document types, feature usage, workflow structures, which teams use the model for what. Individually, these data points look low-sensitivity. Collectively, over time, they build a detailed picture of how the organisation operates — what it's working on, where it's directing AI resources, which functions are scaling their usage and in what direction. This is a competitive intelligence risk that exists independently of whether any personal data is involved. A vendor who can see that a company's legal team started using the model heavily in March, its M&A function spiked usage in April, and its HR team has been running high volumes of termination-related drafts since May has learned something meaningful — without any named individual appearing anywhere in the data. Enterprise agreements typically don't restrict vendors from using aggregated, anonymised usage data for product improvement and benchmarking. The data processing agreement and the main commercial contract often address prompt data and metadata under different provisions, negotiated at different points in the contracting process, with different protections attached to each. The metadata protections are almost always weaker. Read both documents carefully — not just the DPA summary the vendor sends as a compliance checkbox. The training data assumption most enterprise teams get wrong There is a persistent assumption in European enterprises that the enterprise version of an AI tool doesn't use customer data for training. Sometimes this is correct. Often it is an assumption that hasn't been verified against the current version of the data processing agreement. Consumer-grade AI tools — the non-enterprise versions of products your employees are also using — may by default use user inputs to improve the underlying model. This is processing for a new purpose. Under GDPR, processing for a new purpose requires a separate lawful basis. When an employee uses the consumer version of a tool on work data because it's faster or the enterprise licence hasn't been set up on their machine, the organisation has a problem it didn't create deliberately but owns entirely. Enterprise agreements typically include data processing agreements that prohibit training on customer data. Three things determine whether that prohibition is actually protecting the organisation: whether it applies to all data types or only certain categories; whether it covers all infrastructure the vendor uses, including subprocessors; and whether the DPA reflects the current model architecture. I have seen enterprise agreements where the DPA predated a major vendor platform change by over a year. The prohibition on training was still in the original terms — which no longer accurately described what was happening with the data. The check is straightforward. Request the current, signed DPA. Confirm it covers training, all subprocessors, and the current model version. Don't assume continuity because the vendor's name hasn't changed. When model outputs carry a breach risk This risk sits with the deployer — and it is the one most European organisations haven't included in their AI risk registers. If a model has been trained on data that included personal information, it can under certain conditions reproduce elements of that information in outputs. Not reliably, not on demand — but under specific prompting patterns, models have been shown to surface fragments of training data. If European employee or customer data was part of the training set, and the model surfaces it in a response, the deploying organisation has a data breach on its hands. Under GDPR, that breach sits with the deployer. The model provider may carry liability, but the deployer cannot transfer its own. For foundation models from major providers, the deployer typically has no visibility into training data composition. The risk is difficult to quantify and impossible to eliminate entirely. For systems where the deployer has fine-tuned or augmented a foundation model with internal data — retrieval-augmented generation, fine-tuning on internal documents, embedding proprietary content — the deployer has direct control over what went in. That control increases responsibility. Before augmenting or fine-tuning a model with internal data, the architecture review needs to answer: what personal data is in these documents, what happens if the model later surfaces fragments of it, and what technical controls reduce that risk? This is a question most teams are not asking at the design stage. It becomes the question they wish they had asked after an incident. What the enterprise agreement covers — and what it doesn't The enterprise agreement creates contractual obligations. It doesn't create infrastructure controls. The data processing agreement requires the vendor to handle data in certain ways. It doesn't prevent US law enforcement from issuing a lawful demand to a US-incorporated vendor under US law — a structural problem I address in the next article in this series, focused specifically on cross-border transfer risk. What the enterprise agreement also doesn't replace is the audit trail that European regulators now expect to exist. The June 2025 guidance from Germany's data protection conference (DSK) requires documentation across the AI system lifecycle. That documentation needs to exist before regulators ask for it — not in the forty-eight hours after they do. For European companies with German operations, this requirement is not optional and it is not the vendor's responsibility to create. The enforcement precedent is live. A major AI provider was fined €15 million by Italian regulators in December 2024 for GDPR violations — failure to establish a lawful processing basis and inadequate transparency around how user data was handled. European regulators now have precedent, process, and political pressure to continue in this space. The organisations currently running AI programmes without a data flow map are running programmes that have not been assessed for compliance. The gap between those two states is not a legal technicality — it is the gap between a defensible position and one that cannot be defended when examined. What to take from thisMap data flows before deployment, not after. Every AI tool processing external data should have a documented flow — what enters, what is retained, where it's stored, for how long — completed before the tool goes live, not reconstructed months later. Read the current DPA, not the version in the procurement file. Data processing agreements are updated. Request the current signed version and confirm it covers training, subprocessors, and the current model architecture. Classify data before it enters a prompt. Sensitive, personal, commercially restricted, and confidential data should be classified before it enters any AI tool — not assessed after six months of use. Do not treat data minimisation as a substitute for anonymisation. Sending fewer fields does not change the legal classification of what arrives at the model. Personal data that has been minimised is still personal data — names, email addresses, and identifiers stripped down to the "minimum necessary" still carry full GDPR obligations. If personal data cannot be genuinely anonymised before it reaches an external model, the transfer needs a lawful basis, a compliant DPA, and a valid cross-border mechanism — not a smaller dataset. Treat usage metadata as a competitive risk, not just a data protection issue. Understand what your vendor collects on usage patterns, what restrictions apply, and what they're permitted to do with aggregated data. Enforce a clear policy on consumer versus enterprise tool versions. If employees have access to both, the policy needs to specify which applies to which data — because the data handling terms are materially different. Include model output re-identification in your AI architecture review. For any system fine-tuned or augmented with internal data, answer the re-identification risk question before build — not after the system is in production.GDPR has been in force since 2018. The data protection gap in AI programmes isn't a new legal problem — it is an existing problem that the pace of AI adoption is making visible at scale. Most of what European companies need to do to close this gap doesn't require waiting for new regulation. It requires applying existing obligations to a class of tools that moved faster than the governance frameworks around them.
Read full article