The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 10.10

In this chapter · 6 sections
Term help

Data Governance, Privacy & the Training-Data Legal Regime

Data rights can force a retrain, a model withdrawal or an in-region deployment — or, as one 2025 discovery order did, hand twenty million customer conversations to opposing counsel — so design lineage and residency before a corpus or hall becomes an unusable investment.

GOODPUT

What you'll decide here

  1. Which provenance tier your training corpus is built on — fully licensed, lawfully-acquired-and-transformed, or scraped-and-hope — because that single choice sets your settlement exposure, your EU disclosure obligations, and whether a single court ruling can strand the weights you spent nine figures training; the tiers are a US fair-use risk screen, and the acquisition record behind them is what the EU Article 4 route and a court actually read.
  2. Choose shared global capacity or in-region processing, key custody and administration under the applicable data/contract boundary; fragmentation can strand regional capacity.
  3. Where erasure reaches a model, choose verified removal, retraining or withdrawal; output filtering mitigates disclosure while the underlying removal obligation is resolved.
  4. Where the controller/processor line sits for every tenant — what you may and may not do with customer prompts and outputs, who your sub-processors are, and how tenant data is isolated so one customer's data never trains on another's behalf.
  5. Which decisions are data-governance (this chapter) versus model/weight security (Chapter 11.8) — because conflating 'protect the data' with 'protect the weights' produces controls that satisfy neither auditor.

An AI data center is a machine for turning data into a model and then turning user data into outputs. Every other part of this guide treats data as a payload — bytes to store, move, and compute on. This chapter treats data as a legal and regulated object: something with a provenance, an owner, a jurisdiction, a retention clock, and a set of rights attached to it whose applicability follows the operator’s actual purpose, role and jurisdiction. In 2025 these decisions became the difference between a defensible asset and a stranded one: a model trained on the wrong corpus can be the subject of a settlement-scale bill such as the $1.5 billion fund approved in Bartz v. Anthropic on 20 July 2026; a customer-data hall sited in the wrong country can be uncontractable; a deletion request you cannot honor can be a regulatory finding.

We work through five governance surfaces — training-data provenance and the copyright/fair-use frontier; the PII and privacy program; data-residency-by-jurisdiction and cross-border transfer; retention, deletion and the right-to-erasure plumbing; and the customer-facing contract stack (DPAs, sub-processors, tenant isolation). Each legal determination needs an architecture that can enforce it. You cannot bolt residency onto a global fleet, or erasure onto a model with no lineage graph, after the concrete is poured and the weights are trained. This chapter governs the data; the companion problem of protecting the model and its weights — at rest, in transit, in confidential execution — is a different threat model handled in Chapter 11.8.

The legal regime around training data crystallized in 2025 from speculation into case law, and the shape it took is jurisdiction-specific in a way that breaks any single global policy. Four illustrative major regimes now apply, and the corpus has to survive the most demanding one its model will be deployed under.

The US: fair use, but the acquisition matters. The watershed is Bartz v. Anthropic. Judge Alsup's June 2025 summary-judgment ruling split the question in two. Using lawfully-acquired books to train an LLM that returns new text was held "among the most transformative" uses imaginable — fair use under §107. But Anthropic had also downloaded millions of books from shadow libraries (LibGen, PiLiMi) to build a permanent "central library," and that acquisition was not excused by the downstream transformative training. The settlement — a minimum of $1.5 billion across roughly 500,000 works (approximately $3,000 gross per covered work before deductions, not a net claimant payout), plus destruction of the covered pirated LibGen/PiLiMi copies subject to legal preservation — received final approval and entered judgment on 2026-07-20. The release is deliberately narrow: past acquisition and copying (inputs) through 2025-08-25 only — output-based claims and future conduct are not released, and an appeal filed at the deadline has delayed author payouts past end-2026. It remains the most expensive lesson in the industry: fair use can protect the training act while the data-acquisition act remains independently infringing. And the content classes keep widening: Round Hill Music sued both Anthropic and Suno in the Northern District of California on 2026-08-17, pleading lyrics from at least 500 songs in the Anthropic complaint and saying it will amend toward 10,000, seeking up to $1B per case with a stated intent to go to trial — a complaint, not a ruling, but the first major music-publisher action against a foundation-model lab. The consequence for an operator is concrete: provenance is per-document, not per-corpus, and the acquisition method is a load-bearing fact you must be able to prove for every source.

The UK: copies, not weights. In Getty Images v. Stability AI (UK High Court, November 2025), Getty abandoned its primary training-copy claim because it had no evidence that Stable Diffusion training and development occurred in the UK, and lost on secondary infringement: on the evidence for the Stable Diffusion versions at issue, the court held that those final model weights were not infringing copies for the CDPA secondary-infringement claim. The narrow win that survived was a trademark finding for generated images bearing Getty's watermark. The lesson: the UK ruling resolved those weights and that secondary-infringement record, while primary training-copy exposure still turns on where the copies were physically made — a siting question as much as a legal one.

The EU: a statutory regime, not a fair-use doctrine. The EU has no fair use. For the Article 4 route, the 2019 DSM Directive’s text-and-data-mining (TDM) exception allows effective rights reservations, including appropriate machine-readable reservations for online content (robots.txt, TDM Reservation Protocol, metadata). Qualifying Article 3 research mining and licenses have different conditions. Article 4 permits mining of lawfully accessible works whose rightsholders have not reserved TDM rights — honor the reservation, keep copies only as long as the mining requires, and treat any personal data in the corpus under GDPR. Layered on top, the EU AI Act's GPAI obligations — applicable to new models since 2 August 2025 — require providers within the applicable GPAI scope to (a) maintain a copyright policy that respects opt-outs, and (b) publish a public summary of training content using the Commission's mandatory template (finalized July 2025). Detail varies by source: the template requires detailed descriptions of public and private datasets, identification of large public datasets, and additional crawler, domain, and collection details for scraped online sources while protecting trade secrets. This EU transparency obligation makes "we don't disclose our training data" a non-option for any model placed on the EU market.

EU AI Act: identify the role before assigning the artifact
Role / caseArtifact in this chapterStatus (September 2026)
Provider of a new GPAI model placed on market from 2 August 2025Copyright-compliance policy and the required public training-content summary; maintain the provider’s evidence recordGPAI obligations apply; Commission enforcement powers apply from 2 August 2026.
Provider of a GPAI model placed on market before 2 August 2025Dated model identity and the same applicable compliance workArticle 111(3) transition reaches 2 August 2027.
Infrastructure processor hosting another provider’s modelProcessing instructions, data flows, retention and assistance evidenceHosting alone does not assign every GPAI-provider obligation; determine any additional role from the actual activity.
Downstream deployer / system providerRole and model-version record; refer model/system obligations to Part 11Do not infer a GPAI exemption or a high-risk-system deadline from the hosting contract.
AI Act consolidated 27 July 2026; Commission GPAI guidance. The voluntary Code of Practice is a compliance route, not a replacement for mandatory obligations.
The training-data legal regime by jurisdiction (the fork that breaks a single global policy)
JurisdictionGoverning doctrineWhat it permitsWhat it forbids / requiresOperator consequence
United StatesFair use (§107), fact-specificTransformative training on lawfully-acquired works (Bartz v. Anthropic, 2025)No fair-use shield for the acquisition of pirated copies; per-document provenanceProve lawful acquisition for every source; Bartz approved a $1.5B fund and discussed ~$3k gross per covered work before deductions; establish this corpus’s rights separately
European UnionDSM TDM exception + AI Act GPAI rulesArticle 4: lawful access without effective rights reservation; Article 3 research mining and licenses follow separate conditionsArticle 4: honor effective reservations; GPAI providers publish the required training-content summary on their applicable timelineSelect Article 3, Article 4 or a license; apply the selected route and the actual GPAI-provider obligations
United KingdomCDPA; TDM only for non-commercialNarrow research TDM; Getty v. Stability (2025): on the evidence, the Stable Diffusion weights at issue were not infringing 'copies' under the CDPA secondary-infringement claimNo commercial TDM exception; physical location of copy-making mattersProve where training copies are made and whether the specific final model stores or reproduces protected works
JapanArt. 30-4 Copyright ActArticle 30-4: uses without enjoying the expressed thoughts or sentiments, including qualifying information analysisLimited where use unreasonably prejudices the rightsholder's interestsTest the non-enjoyment purpose and unreasonable-prejudice conditions before choosing a Japan corpus-assembly site
2025-2026 state of the law. The right column is the operator consequence, not legal advice. US fair use is fact-specific and unsettled above the district-court level; EU obligations phase in through 2027.

The defensive artifact that ties this together is dataset documentation — a per-corpus record of source, acquisition method, license terms, opt-out status, collection date, and any filtering applied. Datasheets-for-datasets and data cards are no longer academic hygiene; they are the evidence you produce when a plaintiff's expert asks how a specific book entered your weights, and they are the raw material for the EU training-content summary. An operator who cannot answer "where did this document come from and were we allowed to use it" for an arbitrary training sample has, in effect, chosen Tier 3 by omission.

PII handling, minimization, and the privacy program

Copyright governs creative works in the corpus; privacy law governs personal data — and the two regimes are orthogonal, so a corpus can be copyright-clean and privacy-dirty at the same time. The binding question under GDPR and its global cousins (CCPA/CPRA, LGPD, PIPL, India's DPDP) is whether you have a lawful basis to process personal data for training, and whether the model that results still "contains" that personal data.

The lawful-basis fork. The EDPB’s Opinion 28/2024 (17 December 2024) set the European frame: where a controller relies on legitimate interest (Art. 6(1)(f)) for training personal data, that basis must survive a three-step test — a legitimate interest, the necessity of the processing, and a balancing against data-subject rights. The Opinion sets a deliberately high bar for claiming the trained model is "anonymous": you must show that extracting personal data from the model, or regurgitating it, is not reasonably likely. The sting in the tail: if a model was trained on unlawfully-processed personal data, that can taint the lawfulness of deploying the model — unless it has been genuinely anonymized. This is the privacy analogue of the Bartz acquisition problem: an upstream data sin can follow the model downstream.

The enforcement is already here. Italy's Garante fined OpenAI €15 million in December 2024 for training ChatGPT without an adequate legal basis and for inadequate transparency, and ordered a six-month public-awareness campaign. The Court of Rome annulled that decision on 18 March 2026 — but on jurisdictional grounds only, holding that the Irish DPC had become OpenAI's lead supervisory authority in February 2024 before the Garante's final ruling; the court never reached the merits, and the decision is appealable. Regulators are not waiting for the AI Act to mature before applying ordinary GDPR principles to training pipelines.

The engineering response is minimization at ingest, not cleanup at audit. The cheapest place to handle PII is before it enters the corpus: PII detection and redaction in the data pipeline, de-duplication (which both improves the model and reduces the memorization that drives privacy risk), and aggressive filtering of high-sensitivity categories. The expensive place is after training, when removing a person's data may mean retraining. Build the privacy program as a pipeline stage with a documented lawful basis, a record of processing activities (ROPA) where Article 30 requires it, and a data-protection impact assessment (DPIA) where Article 35’s likely-high-risk test is met. One assessment can cover similar processing operations with similar risks; retain per-run evidence of which assessment and lawful basis cover this run — the same artifacts a DPA will ask for, produced before the run rather than after the complaint.

Data residency, cross-border transfer, and the controller/processor line

Residency is where data governance becomes a fleet-architecture decision. The question a customer or regulator asks is deceptively simple — "where does our data live and who can reach it?" — and the answer determines whether a workload can run on shared global silicon or must be pinned to in-region hardware.

The fork is logical vs physical residency. Logical residency is a contractual and software-enforced promise: data is tagged with a region and the control plane routes it to in-region storage and compute, but the underlying fleet is global and a misconfiguration can leak it across a border. Physical residency places processing and storage on silicon in the jurisdiction; separate clusters, key custody and operational staff must also constrain remote administration, replication, support and backups. The hall’s address alone does not keep bytes or access in-region. Physical residency is dramatically more expensive — it fragments your fleet, strands capacity in low-utilization regions, and forecloses the global load-balancing that makes inference economics work — but it is what sovereign, government, and regulated-industry customers increasingly require. Choosing physical residency for a workload is choosing a smaller, more expensive, less elastic deployment, and that choice cascades into siting (you need a hall in that jurisdiction) and capacity planning (you cannot pool demand across borders).

Cross-border transfer mechanisms. Where personal data subject to GDPR is transferred to a third country or international organization, Chapter V requires an applicable transfer route: an adequacy decision (the EU-US Data Privacy Framework, for participating US importers), Standard Contractual Clauses with a transfer impact assessment, or Binding Corporate Rules. China's PIPL requires qualifying handlers to store covered personal information in China and pass a security assessment before exporting it; India's DPDP §16, operationalised by Rule 15 of the DPDP Rules 2025 with effect from 13 May 2027, instead permits transfers unless the government restricts a destination, while stricter sectoral rules continue to apply. The operator consequence is that "the model is served from a US region" can be a compliance event for an EU customer's prompts, and the cross-border posture must be designed per-data-class, not per-company. The deeper sovereignty stack — export controls, control-of-stack, and the geopolitics of where compute is allowed to sit — is a siting problem treated in Chapter 3.12; this chapter handles the customer-data residency obligations that sit on top of it.

The controller/processor line governs what you may do with the data at all. When you serve another company's users, you are typically a processor: you may process their data only on documented instructions, and — critically — you may not reuse it for your own training merely because the contract is silent. Your own training purpose needs a controller-role assessment, lawful basis and transparency as well as the necessary permissions. When you collect data for your own model, you are a controller with the full weight of lawful-basis, transparency, and data-subject-rights obligations. Misplacing this line is the most common governance failure in multi-tenant AI: silently training the foundation model on tenant prompts is a controller act dressed up as a processor convenience, and it is exactly the kind of thing that turns a customer's data into a competitor's model.

$1.5B settlement fund
Bartz settlement fund approved 20 July 2026; not a completed-payout figure
Scope & caveats

Settlement fund approved in the named litigation, not proof of completed distribution, a per-work net payout, appellate finality or a general permission to train.

Aug 2, 2025
EU AI Act GPAI obligations apply to new models; training-content summary template (Commission) mandatory
Aug 2, 2027
Deadline for pre-existing (placed before Aug 2025) GPAI models to publish their training-content summary
3-step test
EDPB legitimate-interest assessment for AI training (interest, necessity, balancing); high bar to claim a trained model is anonymous
Scope & caveats

Purpose/interest, necessity and balancing where legitimate interests is relied upon; not an instruction that AI training generally uses that basis.

at-issue Stable Diffusion weights: not infringing copies
Getty: the at-issue Stable Diffusion weights and secondary-infringement claim, November 2025
Article 4: rights reservations matter
DSM Article 4 route: lawful access and no effective rights reservation; Article 3 has different conditions
Scope & caveats

Article 4 route only; qualifying scientific-research mining under Article 3 has different conditions, and licenses are a separate route. Copyright permission does not displace data-protection obligations.

2 August 2026
Commission GPAI enforcement timeline; determine the provider role first
Scope & caveats

Role-specific GPAI timeline; infrastructure hosting alone does not determine model-provider status. Consolidated AI Act dated 27 July 2026 supplies the binding text.

Retention, deletion, and the right-to-erasure plumbing

The right to erasure (GDPR Art. 17, mirrored by CCPA deletion rights and others) is the governance obligation that AI architecture handles worst, because the naive implementation — "delete the row" — does not reach the place the data actually lives. Personal data in an AI system exists in at least four locations, and a deletion request must be reasoned about for each: (1) the operational store (databases, prompt logs, vector indexes), (2) the training corpus (the dataset snapshot the model was trained from), (3) the model weights (where the data may have been memorized), and (4) downstream artifacts (caches, backups, derived datasets, fine-tunes).

Locations (1) and (2) are tractable with disciplined plumbing. Location (3) is the hard one. You cannot surgically delete an individual's contribution from a trained weight tensor — the data is distributed across billions of parameters in a way that is not addressable. The response starts with an applicability decision: does a valid erasure ground apply, is the material personal data, and is it present in or inferable from the model? If erasure reaches the weights, separate immediate mitigation from removal:

  • Immediate output mitigation. Use filters and guardrails to suppress regurgitation while the request is assessed. This is cheap and fast, but it is not erasure because the data remains in the model.
  • Machine unlearning. Apply an algorithm that approximately removes a data point's influence without a full retrain. An active research area as of 2026, not yet a turnkey production control, and hard to prove to an auditor's satisfaction.
  • Verified removal, retraining, or deletion. Exclude the data from the corpus and verify that the affected model no longer contains or permits inference of it; if that cannot be demonstrated, retrain without the data or delete the model. This is the costly response when erasure reaches the weights. This is why erasure obligations push toward not memorizing in the first place (de-duplication, minimization) and toward retention windows that bound how long raw data is kept before the corpus is frozen.

What makes any of this possible is data lineage — a graph that records, for every model and dataset, which sources fed it, when, and under what basis. That graph is what answers "is this person's data in this model?" and therefore "what is the cheapest valid response to their erasure request?" Without lineage, every erasure request is unanswerable and every audit is a fishing expedition. With it, you can scope the blast radius of a deletion to the smallest set of artifacts that must change.

DPAs, sub-processors, and tenant data isolation

The contract stack is where governance becomes enforceable against the operator. Where GDPR applies and you process personal data on a controller’s behalf, a Data Processing Agreement (GDPR Art. 28) defines the controller/processor relationship: the scope and purpose of processing, the prohibition on using data outside instructions (the no-training clause again), the security measures, the data-subject-rights support you must provide, and the breach-notification timeline. The DPA is the document that makes "we will not train on your data" legally binding rather than a marketing claim; a processor must follow documented controller instructions under Article 28(3)(a). Training for the operator’s own purpose requires its own controller-role and Article 6 lawful-basis assessment; permission in a contract alone does not supply that basis.

Sub-processor management is the part operators underestimate. A downstream party processing that personal data on your behalf — the cloud region, the GPU neocloud you burst to, the vector-database SaaS, the content-moderation API — can be a sub-processor; assess independent controller roles separately. Article 28(2) requires prior specific or general written authorization; under general authorization, give notice of intended additions or replacements and an opportunity to object. Article 28(4) requires equivalent data-protection obligations in the downstream contract and leaves the initial processor liable to the controller for that sub-processor’s performance. A published sub-processor list is a contractual requirement where the DPA makes it one. An AI serving stack assembled from a dozen vendors needs a role and data-flow determination for each, and a tenant's right to object to a new one can constrain your own procurement. The governance consequence: your supply chain is part of your compliance surface, and a neocloud burst-out (see Chapter 1.8) is a sub-processor event that must be papered before the traffic flows.

Tenant data isolation is where data governance meets the multi-tenancy architecture. The governance requirement is simple to state and hard to guarantee: one tenant's data must never be readable by another, never leak through a shared cache or vector index, and never enter cross-tenant model improvement without its own valid purpose, permissions, lawful basis and transparency. The enforcement mechanisms — namespace isolation, per-tenant encryption keys, isolated KV-cache and embedding stores, and the strict separation of any data used for model improvement — are the data-plane half of a problem whose compute-plane half (MIG, vGPU, confidential computing as security boundaries) lives in Chapter 11.6. Keep the two objectives apart: isolating tenants means the data does not cross the boundary; protecting the model weights from extraction is a different objective, handled in Chapter 11.8. An operator who treats these as one problem builds controls that satisfy neither the privacy auditor nor the security auditor.

Deep dive: why 'we'll just delete it later' fails for memorized training data

The most consequential misunderstanding in AI data governance is the belief that erasure is a database operation. It is not, and the reason is architectural. When a person's data is in your operational store, deletion is a DELETE statement and a backup-expiry policy. When that same data has been included in a training corpus and the model has been trained, the data has undergone a one-way transformation: it has been smeared across billions of floating-point parameters via gradient descent, in a representation that is not indexed by individual, not addressable, and not reversible. There is no row to delete. The information may not even be recoverable from the model — but "may not be recoverable" is not the same as "has been erased," and a regulator applying the EDPB's high anonymity bar will ask you to demonstrate the former.

This is why the governance burden shifts upstream. The cheapest erasure is the one you never have to perform on the weights, which means: minimize what enters the corpus, de-duplicate aggressively (de-duplication is the single most effective memorization suppressant), keep raw personal data only inside a bounded retention window before the corpus is frozen, and maintain lineage so you can scope any future request. The production posture starts with the applicability decision — confirm that erasure applies, identify whether the material is personal data, and determine whether it is present in or inferable from the model. Use output filtering only for immediate harm mitigation; when erasure reaches the weights, apply a documented, verifiable removal method or retrain or delete the affected model. What you cannot credibly commit to is on-demand surgical deletion from a frozen model, and a DPA that promises it is a liability you cannot honor. Design to the statutory clock, not to your release calendar: where GDPR erasure applies, Article 12(3) requires you to inform the data subject of the action taken without undue delay and within one month, extendable by two further months only with timely notice and reasons, and Article 17 requires the erasure itself without undue delay. A retrain cadence is not a lawful substitute for either, and disclosing the cadence does not make it one. Set the retention windows, the removal method, and the retrain-or-withdraw decision so the obligation is one you can actually execute inside those limits.

Deep dive: building the data-governance control plane as infrastructure, not policy

The recurring failure mode is treating governance as a set of PDFs reviewed by a legal team, when the obligations are enforceable only if they are wired into the data plane. The artifacts that actually hold up under audit are systems, not documents:

  • A provenance ledger at corpus-assembly time: every source tagged with acquisition method, license, opt-out status, and collection date — the evidence for a Bartz-style acquisition question and the raw material for the EU training-content summary.
  • PII handling as a pipeline stage: detection, redaction, and de-duplication enforced in the ingest path, with a per-run record of lawful basis and coverage by the risk-triggered DPIA, where Article 35 requires one.
  • A residency-aware control plane: data tagged with a jurisdiction, routing that honors it, and a transfer-impact register for every cross-border path — so "logical residency" is a software guarantee with a failure alarm, not a promise.
  • A lineage graph connecting data subjects to datasets to models, so an erasure request resolves to a bounded set of artifacts instead of a shrug.
  • A retention engine that executes deletion schedules by default and applies legal holds as documented exceptions — reconciling the privacy program's pull toward deletion with litigation's pull toward preservation.
  • A sub-processor and DPA registry that flows obligations down the supply chain and gates procurement on contractual coverage.

Each of these is an engineering deliverable with an owner and an SLO, sitting alongside the orchestration and storage planes (Chapter 10.1), not a quarterly compliance review. Governance that lives only in policy is governance that fails the first time it is tested by a discovery request, a DPA inquiry, or an erasure demand.

Anti-patterns

The same governance mis-scopes recur, each one from treating data as a payload instead of a regulated object:

  • Tier-3 by omission. Crawling and training without a provenance ledger, discovering only under subpoena that you cannot prove lawful acquisition for the sources a plaintiff cares about. The fix is upstream: tag provenance at assembly, not at audit.
  • The silent no-training violation. Training the foundation model on tenant prompts without documented controller instructions and a lawful basis for the operator’s own training purpose — a controller act performed under the cover of a processor role, and the fastest way to turn a customer's data into a competitor's model and a regulatory finding.
  • Promising surgical erasure from frozen weights. A DPA or privacy notice that commits to on-demand deletion of memorized data the architecture cannot deliver. Determine where the applicable erasure obligation reaches, then execute verified removal, retraining or model withdrawal within that obligation. Track Article 12(3)’s response separately; an internal retraining cadence cannot postpone action without undue delay under Article 17.
  • Keep-everything logging. Retaining prompt and output logs indefinitely "for debugging," then discovering they are a 20-million-record discoverable liability. Minimal retention by default; legal holds as exceptions.
  • Conflating data governance with weight security. Building one control set for "protect the data" that an auditor reads as also covering "protect the weights," satisfying neither. They are different threat models — this chapter and Chapter 11.8 respectively.

Admit a source only when its intended use has a defensible rights and personal-data basis, then carry that decision through every derived artifact. Quarantine delays training; an undocumented corpus can strand the weights, embeddings and releases built from it. Choose erasure mechanisms that can execute the obligation, rather than a promise to forget on a schedule the architecture cannot meet.

This chapter governs the data; the model and its weights are protected as a distinct asset in Chapter 11.8, and the multi-tenant compute-isolation boundary (MIG, vGPU, confidential computing) is engineered in Chapter 11.6. The sovereignty, export-control, and geopolitical layer beneath data residency is sited in Chapter 3.12, with the market-cluster and site-scoring mechanics in Chapter 3.13. The orchestration plane the governance control plane plugs into is Chapter 10.1, and the object-storage and data-lake tier that carries the retention and deletion path is Chapter 9.6; the KV-cache and embedding stores that must be tenant-isolated are the memory hierarchy of Chapter 9.7. The economics of the neocloud burst-out that creates sub-processor obligations live in Chapter 1.8; the metric and vocabulary backbone is Chapter 0.3; and the SLA framing that DPAs sit alongside is Chapter 12.4.
Cite this chapter
Fehn, J. (2026). Data Governance, Privacy & the Training-Data Legal Regime (Chapter 10.10). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-10-data-governance-privacy-and-the-training-data-legal-regime (accessed 2026-09-29).
@misc{aidc-10-10,
  author       = {Fehn, Jacob},
  title        = {Data Governance, Privacy & the Training-Data Legal Regime (Chapter 10.10)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-10-data-governance-privacy-and-the-training-data-legal-regime},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit