Data Compliance

What Does the EU AI Act Actually Require About Training Data?

15 min read · 2026-08-17
Aug 2, 2025date GPAI training-data-summary and copyright-policy duties (Article 53) took effect
Dec 2, 2027new deadline for Article 10 data-governance duties on Annex III high-risk systems, delayed from Aug 2, 2026
3% / EUR 15Mmaximum fine (whichever is greater) for GPAI non-compliance, enforceable from Aug 2, 2026
Aug 2, 2027compliance deadline for GPAI models already on the market before Aug 2, 2025
Key takeaways
  • 1Two articles govern training data specifically: Article 53 for general-purpose AI (GPAI) models, in force since August 2, 2025, and Article 10 for high-risk AI systems, which requires data to be relevant, representative, error-free, and complete, with examination for bias.
  • 2GPAI providers must publish a public training data summary using the AI Office's official template and maintain a copyright policy that honors text-and-data-mining opt-outs under Article 4(3) of Directive (EU) 2019/790, both since August 2, 2025.
  • 3The Digital Omnibus (Regulation (EU) 2026/1744), in force since July 27, 2026, pushed the Article 10 high-risk data governance deadline for Annex III systems from August 2, 2026 to December 2, 2027, and for Annex I product-embedded systems to August 2, 2028.
  • 4The Act applies regardless of where a company is headquartered: Article 2 reaches any provider placing a GPAI model or AI system on the EU market, or whose system's output is used in the EU, U.S. and other non-EU sellers of training data included.
  • 5GPAI enforcement carries real teeth starting August 2, 2026: fines up to the greater of 15 million euros or 3% of global annual turnover, which is why documented provenance and licensing paper trails are now a purchasing requirement, not a nice-to-have.
The short version

Most explainers of the EU AI Act stay vague about training data because the actual text is scattered across two articles that apply on two different tracks. Article 53 covers general-purpose AI models and has been in force since August 2, 2025: providers must publish a detailed public summary of training content using the AI Office's official template, keep technical documentation, and maintain a copyright policy that respects opt-outs under the EU's text-and-data-mining exception. Article 10 covers high-risk AI systems and requires that training, validation, and testing data be relevant, representative, error-free, complete, and examined for bias, but its main compliance deadline was pushed from August 2, 2026 to December 2, 2027 by the Digital Omnibus regulation that entered into force on July 27, 2026. Both articles apply to any company placing a covered model or system on the EU market, regardless of headquarters. For anyone buying or selling AI training data, the practical result is that documented provenance, a clean licensing record, and bias/representativeness testing have moved from best practice to a purchasing requirement enforced with fines of up to 3% of global turnover.

On this page ▾

The EU AI Act's training data rules live in two articles, not one

Search for "EU AI Act training data requirements" and most results describe the Act in general terms: risk tiers, prohibited practices, a vague nod to transparency. That's not useful if you're sourcing, selling, or licensing a dataset that might end up training a model sold into the EU. The Act's training-data-specific obligations sit in exactly two places, and they apply on two different timelines to two different kinds of AI system.

  • Article 53 governs providers of general-purpose AI models (GPAI): the foundation models that power chatbots, coding assistants, and most commercial LLM products. It has been in force since August 2, 2025. [8]
  • Article 10 governs high-risk AI systems: narrower, specific-purpose systems listed in Annex III (biometrics, employment, credit scoring, law enforcement, migration, education, and similar categories) and safety components covered by Annex I. Its main compliance date, originally August 2, 2026, was pushed to December 2, 2027 by an amendment that took effect in July 2026, covered in detail below. [8][10]

Regulation (EU) 2024/1689, the AI Act's formal name, entered into force on August 1, 2024, and applies on a staggered schedule set out in Article 113: prohibited practices from February 2, 2025; governance rules and GPAI obligations (Chapter V, which includes Article 53) from August 2, 2025; the general application date of August 2, 2026; and originally Article 6(1) high-risk classification rules for Annex I products from August 2, 2027. [2] Keep that structure in mind, because almost every date claim below traces back to it or to the 2026 amendment that adjusted part of it.

Why this matters if you're not building an AI model yourself
If you're a startup selling a dataset, or a lab or enterprise buying one, you don't file compliance paperwork with Brussels. But your buyer or your buyer's buyer might, and Articles 10 and 53 both flow downhill: a GPAI provider can't write a truthful training-content summary or a defensible copyright policy without knowing what's in the data it licensed, and a high-risk system builder can't document data governance without provenance records from whoever supplied the training set.

What GPAI providers must do about training data, and by when (Article 53)

Article 53(1) lists four obligations for any provider of a general-purpose AI model. Two of them are about training data directly, and a third, the copyright policy, is functionally a training-data-sourcing obligation even though it's framed around copyright compliance. [2]

  • 1Technical documentation (Art. 53(1)(a)). Providers must draw up and keep current technical documentation of the model, including its training and testing process, in the level of detail specified in Annex XI, and make it available to the AI Office and national authorities on request.
  • 2Downstream documentation (Art. 53(1)(b)). Providers must give AI system developers who build on top of the model enough information (per Annex XII) to understand its capabilities and limitations and integrate it in compliance with their own obligations.
  • 3Copyright policy (Art. 53(1)(c)). Providers must put in place a policy to comply with EU copyright law, specifically to identify and honor a reservation of rights that a rightsholder has expressed under Article 4(3) of the Copyright in the Digital Single Market Directive (EU) 2019/790, the EU's text-and-data-mining opt-out mechanism.
  • 4Public training data summary (Art. 53(1)(d)). Providers must draw up and publish a sufficiently detailed public summary of the content used to train the model, using the standardized template the AI Office published.

The training-content summary requirement is the one that gets the least accurate coverage online, so it's worth being precise about the mechanics. The European Commission adopted the official template on July 24, 2025, nine days before the GPAI rules took effect, and tied it explicitly to Article 53(1)(d). [2] It requires providers to give a comprehensive overview of the data used to train the model: the main data collections, other sources drawn on, and enough detail for a party with a legitimate interest, most notably a copyright holder, to understand whether their material was likely used. [4] It does not require disclosing a full dataset manifest or trade secrets; it is a narrative-level public disclosure, not a technical audit trail.

Providers of general-purpose AI models shall draw up and make publicly available a sufficiently detailed summary about the content used for training of the general-purpose AI model, according to a template provided by the AI Office.
European Commission, AI Act Service Desk: Article 53: Obligations for providers of general-purpose AI models
The copyright policy duty is a training-data sourcing control in disguise
Article 53(1)(c) doesn't just ask for a policy document. It requires the provider to actually identify content that rightsholders have opted out of text-and-data-mining under Article 4(3) of Directive (EU) 2019/790, using state-of-the-art technical means, and exclude it. [2][4] In practice this pushes GPAI providers to demand clean provenance and licensing records from anyone supplying training data, because a provider can't honor an opt-out for a dataset whose source chain it can't trace.

A GPAI model, in the Act's terms, is a model trained with a large amount of data using self-supervision at scale, that displays significant generality, and that can competently perform a wide range of distinct tasks, the definition covers most large foundation models, whether or not the provider labels them as "general-purpose." The Article 53 obligations applied to any such model placed on the EU market from August 2, 2025 onward. [4][6]

Models that were already on the market before that date got a grace period: providers of pre-existing GPAI models have until August 2, 2027 to bring their models and documentation into compliance, a two-year runway rather than an exemption. [4] Enforcement, meanwhile, follows its own clock. The Commission and AI Office gained the power to request documentation, run technical evaluations, and demand risk-mitigation measures from August 2, 2025, but the ability to actually fine a GPAI provider, up to the greater of 3% of global annual turnover or 15 million euros, only activated on August 2, 2026, alongside the general application of most of the rest of the Act. [10]

A voluntary General-Purpose AI Code of Practice, finalized in mid-2025, gives providers a way to demonstrate good-faith compliance; signing it isn't mandatory and doesn't grant immunity, but the AI Office weighs adherence to it when deciding penalty amounts. [2] None of that changes the underlying legal duty: whether or not a provider signs the Code, Article 53's documentation, copyright-policy, and training-summary obligations apply on the dates above.

Article 10: data governance for high-risk AI systems

Article 10 is more prescriptive than Article 53 and reads much more like a data quality standard than a disclosure rule. It applies to high-risk AI systems, the category defined in Article 6 and enumerated mainly in Annex III (biometric identification, critical infrastructure, education and vocational training, employment and worker management, access to essential services, law enforcement, migration and border control, and administration of justice) plus Annex I product-safety-embedded systems. [4] Where such a system involves training an AI model, Article 10 requires the training, validation, and testing datasets to satisfy specific quality criteria before the system can be placed on the market or put into service.

Article 10 requirementWhat it actually means for a datasetSub-paragraph
RelevantData must actually bear on the task the system performs; no dumping in unrelated data to pad volumeArt. 10(3)
Sufficiently representativeCoverage must reflect the populations and contexts the system will be used on, not just the easiest-to-collect subsetArt. 10(3)
Free of errors, to the best extent possibleLabeling and data-entry error rates must be actively managed, not just tolerated as noiseArt. 10(3)
CompleteNo systematic gaps that would bias the model against particular groups or scenariosArt. 10(3)
Appropriate statistical propertiesDistributions must suit the specific geographic, contextual, behavioral, or functional deployment settingArt. 10(4)
Examined for biasProviders must actively examine data for biases likely to affect health, safety, fundamental rights, or lead to discrimination, and take mitigation measuresArt. 10(2)(f), 10(2)(g)
Documented governanceData collection processes, origin of data, and preparation steps (annotation, labeling, cleaning, enrichment, aggregation) must be documentedArt. 10(2)(a)-(c)

Article 10(5) carves out a narrow, tightly-conditioned allowance: providers may process special-category personal data specifically for bias detection and correction, subject to safeguards including that the data can't be transferred elsewhere and must be deleted once the bias assessment is complete. [2] Article 10(6) extends the quality requirements to testing datasets even for high-risk systems that don't involve training an AI model at all, so the article isn't limited to machine-learning-based systems narrowly defined.

This is the closest thing in EU law to a training-data quality statute
Article 10 is unusual among data regulations in specifying substantive data quality criteria rather than just procedural consent or disclosure requirements. For a buyer of training data whose end use might land in a high-risk category, representativeness and documented bias testing are not optional add-ons Dayda recommends as best practice; they are what Article 10 requires the deploying system to be able to prove.

The July 2026 Digital Omnibus changed the high-risk timeline, not the substance

This is the fact most existing explainers of the AI Act will get wrong if they were written before mid-2026: the August 2, 2026 deadline for high-risk AI system obligations no longer applies as originally written. The European Parliament and Council adopted Regulation (EU) 2026/1744, commonly called the Digital Omnibus on AI, on July 8, 2026. It was published in the Official Journal on July 24, 2026, and entered into force on July 27, 2026, roughly three weeks before this article's publication. [6]

Recital 40 of the amending regulation sets the new date of application for Sections 1 through 3 of Chapter III, the sections covering high-risk system classification, requirements, and obligations, including Article 10, at December 2, 2027 for systems classified as high-risk under Article 6(2) and Annex III (the standalone, non-product-embedded category: biometrics, employment, education, and the rest). [4] Annex I systems, those embedded as safety components in products already regulated under separate EU product-safety law, get a further deferral to August 2, 2028, a full year later than the original August 2, 2027 date. [8]

ObligationArticle / AnnexOriginal dateCurrent date (post-Omnibus)
Prohibited AI practicesArt. 5Feb 2, 2025Feb 2, 2025 (unchanged)
GPAI technical documentation, copyright policy, training-data summaryArt. 53Aug 2, 2025 (new models)Aug 2, 2025; Aug 2, 2027 for pre-existing models [2]
GPAI enforcement and fining powersChapter XII / Art. 113Aug 2, 2026Aug 2, 2026 (unchanged) [2]
New prohibition: non-consensual intimate imagery / CSAM-generation systemsArt. 5 (added)n/a (new)Dec 2, 2026 [2]
Data governance for standalone high-risk systems (biometrics, employment, education, etc.)Art. 10, Annex IIIAug 2, 2026Dec 2, 2027 [4]
Data governance for product-embedded high-risk systemsArt. 10, Annex IAug 2, 2027Aug 2, 2028 [4]
A delayed deadline is not a repealed requirement
The Digital Omnibus moved dates; it did not remove Article 10's substantive text-quality and bias-examination requirements, and it did not touch Article 53 at all. Treat the new 2027/2028 dates as the current binding deadlines, not as a sign that data governance for high-risk systems has become optional. The Commission's own stated rationale for the delay was implementation readiness (slow designation of national authorities and conformity-assessment bodies, and a lack of finalized harmonized standards), not a policy retreat from the underlying obligations.

Who this actually applies to (it is not an EU-companies-only rule)

Article 2 sets the Act's territorial scope, and it is deliberately broad. It applies to providers placing an AI system or GPAI model on the EU market or putting it into service in the EU, irrespective of whether those providers are established or located within the Union or in a third country. [4] A separate clause reaches providers and deployers located in a third country whose AI system's output is used in the Union. [6] In practice: a U.S.-headquartered lab training a foundation model that any EU customer can access has Article 53 obligations, full stop, regardless of where its servers, staff, or incorporation sit.

For a training-data marketplace and its participants, that scope collapses a distinction some sellers assume protects them: "my buyer is American" is not a compliance argument if the buyer's model, or the buyer's buyer's model, ends up serving EU users. Once a GPAI model or high-risk system touches the EU market, the entire training-data supply chain feeding it becomes relevant to the provider's Article 53 or Article 10 obligations, and the provider will push documentation requirements back down that chain to whoever it licensed data from.

A worked example: a startup selling clinical intake data to two different buyers

Take a health-tech startup winding down with 600,000 de-identified patient intake conversations. It has two term sheets on the table: one from a U.S. enterprise building an internal triage tool used only in its domestic clinics, and one from an AI lab building a general-purpose clinical assistant it plans to offer through an EU-accessible API. The data itself is identical. The compliance path is not.

  • 1Buyer A (domestic-only triage tool). If the deployed system is used exclusively outside the EU and its output never reaches EU users, Article 2 scope isn't triggered by this deal alone. [4] The startup's obligations remain what they were already: HIPAA de-identification and standard provenance documentation, covered in Dayda's provenance and consent guide.
  • 2Buyer B (EU-accessible GPAI clinical assistant). The moment this model is placed on the EU market, Buyer B has Article 53 duties: a training-content summary, technical documentation, and a copyright policy. Buyer B's diligence team will ask the startup for exactly the source chain and licensing terms that let Buyer B write a truthful public summary, and if any of the intake data includes content subject to a text-and-data-mining opt-out under Article 4(3) of Directive (EU) 2019/790, Buyer B needs to know before training, not after publishing the summary. [4][6]
  • 3If either buyer's tool is later reclassified as high-risk (a clinical triage assistant that influences access to healthcare plausibly falls under an Annex III category), Article 10's data governance duties attach on top: the startup's representativeness across patient demographics and its bias-examination records become part of the deploying system's own compliance file, due by December 2, 2027 under the current schedule. [4][6]
The practical takeaway
The seller's own compliance burden doesn't automatically change; the buyer's does, and the buyer's compliance burden becomes the seller's due-diligence questionnaire. A startup that can answer provenance, licensing, and representativeness questions before being asked closes the EU-bound deal faster and usually at a better price, because it removes the diligence bottleneck that kills deals with buyers who have their own Article 53 or Article 10 exposure.

A practical compliance checklist for buyers and sellers of training data

Whether you're preparing a dataset for sale or evaluating one to buy, this is the checklist that maps directly onto what a counterparty's legal team will ask for once Article 53 or Article 10 is in play.

  • 1Establish whether the buyer's end product has any EU touchpoint. If the model or system will be placed on the EU market, or its output used in the EU, assume Article 2 scope applies and prepare accordingly. [4]
  • 2For GPAI-bound data, document source-level provenance well enough to support a public summary. The buyer needs to state, at a narrative level, where training content came from; a seller who can't name the source chain for a corpus makes that summary impossible to write truthfully. [4]
  • 3Flag any content subject to a text-and-data-mining opt-out. If any part of the corpus originated from sources where a rightsholder reserved rights under Article 4(3) of Directive (EU) 2019/790, disclose it. A buyer's copyright policy has to identify and exclude that content, and finding it after training is far more expensive than flagging it before the deal closes. [4]
  • 4Keep data-preparation records: collection process, annotation, cleaning, enrichment. Article 10(2) requires exactly this documentation from any high-risk system builder, and they will not accept a dataset that arrives with no record of how it was labeled or cleaned. [4]
  • 5Run and document a representativeness and bias check before listing data for a high-risk-adjacent use case. Demographic coverage gaps and known bias patterns are Article 10(3) territory; disclosing them upfront is a diligence asset, not a liability, as long as they're documented and, where feasible, mitigated. [4]
  • 6Track both timelines separately. GPAI obligations (Article 53) are already live and enforceable; high-risk obligations (Article 10) are not enforceable until December 2, 2027 for Annex III and August 2, 2028 for Annex I, under the current Digital Omnibus schedule. Don't over- or under-prioritize compliance work based on the wrong deadline. [4]
  • 7Re-check the timeline before every deal. The high-risk deadline already moved once, by 16 months, in mid-2026. Verify the current applicable date against EUR-Lex or the AI Office rather than relying on an older article, including this one, without a re-check.

What's still unsettled as of this writing

Some parts of this framework are genuinely still in motion, and treating them as settled would be dishonest. The General-Purpose AI Code of Practice remains voluntary; it was finalized in mid-2025 as a way to demonstrate good-faith compliance with Article 53, and while the AI Office has said it will weigh adherence when assessing penalties, it is guidance for demonstrating compliance, not a substitute for the statutory text itself. [2] Harmonized technical standards that would give Article 10 compliance a presumption-of-conformity path are, as of the Digital Omnibus's own stated rationale for delay, still incomplete; the Commission cited the absence of finalized harmonized standards and slow designation of national conformity-assessment bodies as reasons for pushing the high-risk deadline back. [4] And the Digital Omnibus itself is recent enough, in force less than a month at the time of this article's publication, that secondary guidance interpreting its exact interplay with the earlier high-risk classification and registration rules is still being written by national authorities and the AI Office. Anyone relying on a specific high-risk compliance date beyond what's cited here should verify it against EUR-Lex or the Commission's AI Act service desk directly, not against an older summary.

The rule of thumb: two articles, two clocks, one underlying standard

If you remember nothing else from this article, remember this: Article 53 for GPAI models is live now and asks for a public training-content summary plus a copyright policy that respects text-and-data-mining opt-outs; Article 10 for high-risk systems asks for relevant, representative, error-free, complete, bias-examined data, and its main deadline is December 2, 2027 after the July 2026 Digital Omnibus delay. Both apply to any company reaching the EU market, wherever it's headquartered. Neither article cares whether the underlying data changed hands through a formal marketplace or a handshake deal, only whether the provenance and quality can be documented when a regulator, or a buyer's legal team, asks.

That's the shift worth internalizing if you're winding down a company and sitting on a proprietary dataset, or shopping for one to license. Documented provenance, a clean licensing paper trail, and representativeness testing used to be a way to command a premium price. Under the EU AI Act, for any deal with an EU-market buyer anywhere downstream, they're the price of admission.

Sell data that clears EU AI Act diligence, not just a price negotiation.

Get a free, no-obligation provenance and compliance review from the team that brokers AI training data deals every day. We'll flag what an EU-bound buyer's legal team will ask for before they ask.

List your data on Dayda