Regulation explainerCustom AI Model Development
Can you train on this dataset? Copyright, GDPR and EU AI Act checks
Whether you may train a model on a dataset turns on separate questions: who owns the content and on what terms, whether it contains personal data and on what lawful basis, and whether the model brings provider duties under the EU AI Act. This page maps the instruments behind each question, explains when a narrow custom model falls outside the general-purpose AI rules and ends with a review to finish before training starts.
On this page
- Scope and limits of this training data rights explainer
- Why data rights shape the dataset, not just the sign-off
- Instruments that decide whether you may train on a dataset
- Common data sources and the rights questions each raises
- Does your model bring EU AI Act provider obligations?
- Opt-outs, provenance and deletion requests in practice
- Rights review to complete before training starts
- Questions and answers
- Sources
Scope and limits of this training data rights explainer
Why data rights shape the dataset, not just the sign-off
Rights questions are cheapest to answer before collection, because the answers change what you collect. A license that excludes machine learning removes a source; a data protection assessment may require dropping free-text fields or pseudonymizing identifiers. Found after training, the same issues can mean retraining, which is why ColdAI's custom model work treats rights-cleared task data as a readiness condition1.
Give each question an owner. Copyright asks whether you may copy and analyze the content at all. Data protection asks whether training on personal data has a lawful basis compatible with its original purpose. Contract asks whether your agreements permit this use. For some projects the EU AI Act adds a fourth: whether the model makes you a provider with duties of your own.
Instruments that decide whether you may train on a dataset
The main EU instruments, plus the UK and US positions that most often arise for organizations working across those markets.
Directive (EU) 2019/790 on copyright in the Digital Single Market, Articles 3 and 4[^2]
European Union, through national implementing lawsApplies whenYou reproduce or extract protected works or databases to analyze them computationally, including to train a model.
- Article 3 permits text and data mining for scientific research by research organizations and cultural heritage institutions with lawful access2.
- Article 4 permits mining for any purpose, including commercial, with lawful access and where the rightsholder has not reserved that use2.
- Under Article 4(3), reservations for content online may be made by machine-readable means; Article 4(2) allows copies to be kept only as long as the mining requires2.
GDPR, Regulation (EU) 2016/679 on the protection of personal data[^3]
European Union and EEA; the UK applies its own UK GDPRApplies whenThe training data relates to identified or identifiable people, including through free text, images or metadata.
- Identify a lawful basis under Article 6; the EDPB has said legitimate interest needs a case-by-case balancing test4.
- Respect purpose limitation under Article 5(1)(b), with compatibility of reuse assessed under Article 6(4)3.
- Process special-category data, such as health data, only under a condition in Article 9(2)3.
- Carry out an impact assessment under Article 35 where high risk to individuals is likely3.
EU AI Act (Regulation (EU) 2024/1689), Article 53[^5]
European Union, for providers placing models on the EU market wherever they are establishedApplies whenYou place a general-purpose AI model on the EU market; these duties apply from 2 August 2025, and models already on the market before then have until 2 August 20276.
- Keep technical documentation and inform downstream providers under Article 53(1)(a) and (b); open-source releases are exempt unless the model has systemic risk5.
- Maintain a policy to comply with EU copyright law, including honoring reservations under Article 4(3) of Directive (EU) 2019/7905.
- Publish a sufficiently detailed summary of training content on the AI Office template5.
Copyright, Designs and Patents Act 1988, section 29A[^7]
United KingdomApplies whenYou copy works in the UK to carry out computational analysis of them.
US Copyright Act, fair use (17 U.S.C. § 107)[^8]
United StatesApplies whenYou copy protected works in the US to train a model without a license.
- Fair use is decided case by case on statutory factors; the US Copyright Office's training report concludes some training uses are likely fair and others are not, depending on purpose and market effect8.
Common data sources and the rights questions each raises
| Data source | Copyright and license | Personal data | Contract and confidentiality |
|---|---|---|---|
| Your own operational records | Usually yours; check embedded third-party material | Often present; check the original purpose and notices | Internal policies or employee agreements may limit reuse |
| Customer data you hold as a supplier | The content may belong to the customer | You may be a processor acting only on the customer's instructions | The service agreement must permit training, and say for whose benefit |
| Licensed third-party datasets | License scope decides; look for terms on machine learning and derived models | Ask how the licensor collected personal data | Check sublicensing, territory and what happens to trained models at expiry |
| Public web content | EU commercial mining relies on Article 4 and must respect opt-outs; elsewhere it varies | Public availability does not remove data protection duties | Site terms may prohibit scraping or automated collection |
| Outputs of another AI model | Protection for generated output is unsettled in many jurisdictions | Generated text can still reproduce real personal data | The model provider's terms may restrict training competing models on outputs |
The table lists questions, not answers: the same source can be cleared for one purpose and not for another.
Does your model bring EU AI Act provider obligations?
Article 53 applies to providers of general-purpose AI models, not to every organization that trains one. The Commission's guidelines give indicative criteria for telling the difference6.
- If
You are training a narrow model for one task, such as classifying documents, detecting defects or transcribing speech.
ThenArticle 53 is unlikely to apply, because the guidelines treat models limited to a narrow task as outside the general-purpose definition6. Check instead whether the system it sits in is high-risk, which brings data governance duties under Article 105.
Obligations follow what a model can do and how it is used.
- If
You are training a generative model whose training compute exceeds 10^23 floating-point operations and that can generate text, audio, images or video6.
ThenTreat it as presumptively a general-purpose AI model and plan the documentation, copyright policy and public training-content summary from the start.
The summary needs provenance recorded at collection.
- If
You are fine-tuning or otherwise modifying someone else's general-purpose model.
ThenYou do not normally become its provider; the guidelines expect that only where the modification uses more than a third of the original model's training compute6.
You still need rights to your own fine-tuning data.
- If
You plan to release the weights under a free and open-source license.
Opt-outs, provenance and deletion requests in practice
Honoring rights reservations is engineering work. For web content, check machine-readable signals such as robots.txt rules, plus terms or metadata expressing a reservation, at collection time; store what you found with each document and exclude reserved content before training. The Commission's code of practice for general-purpose model providers describes measures of this kind9.
Provenance records, covering source, license or legal basis, collection date and the dataset versions that include each item, let you answer a rightsholder complaint or an erasure request without guessing.
Deletion is harder for models than for databases: removing a person's influence from trained weights generally means retraining from the cleaned dataset at the next release. The EDPB has said that whether a model trained on personal data is anonymous must be assessed case by case, considering both extraction of training data and what queries reveal4. Minimizing personal data before training makes the question rarer.
Rights review to complete before training starts
Questions and answers
Can we train on content scraped from public websites?
Possibly, but public does not mean free to use. In the EU, commercial training relies on the Article 4 text and data mining exception, which requires lawful access and yields to machine-readable opt-outs, and site terms may prohibit scraping. Personal data in scraped content still needs a lawful basis under the GDPR. Elsewhere the position differs, so check each jurisdiction involved.
Does anonymizing training data remove GDPR obligations?
Truly anonymous data falls outside the GDPR, but the bar is high: re-identification must not be reasonably likely with the means available to you or anyone else. Pseudonymized data, where identifiers are replaced but could be linked back, is still personal data. Anonymize or pseudonymize as early as possible, document the method, and remember that the trained model may need its own anonymity assessment.
Must we retrain a model after a deletion request?
Not always immediately, but you need a defensible approach. Remove the person's data from the dataset at once, assess whether the model could reproduce or reveal it, and plan to retrain from the cleaned dataset at the next release if it could. Recording which dataset versions contain which records makes this manageable; without provenance you cannot show the request was honored.
Can we use customer data to improve our own model?
Only if your contracts and your role allow it. If you process data on a customer's behalf, you may act only on their instructions, so training a model for your wider business needs their agreement and a clear contractual basis. Check confidentiality clauses too, because trained models can sometimes reproduce fragments of their training data.
Sources
- Custom AI Model Development: readiness conditions and delivery approach — ColdAI
- Directive (EU) 2019/790 on copyright and related rights in the Digital Single Market — EUR-Lex · checked 10 October 2026
- Regulation (EU) 2016/679 (General Data Protection Regulation) — EUR-Lex · checked 10 October 2026
- Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models — European Data Protection Board · checked 10 October 2026
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) — EUR-Lex · checked 10 October 2026
- Guidelines for providers of general-purpose AI models — European Commission · checked 10 October 2026
- Copyright, Designs and Patents Act 1988, section 29A: copies for text and data analysis for non-commercial research — legislation.gov.uk · checked 10 October 2026
- Copyright and Artificial Intelligence, Part 3: Generative AI Training (pre-publication version) — US Copyright Office · checked 10 October 2026
- The General-Purpose AI Code of Practice — European Commission · checked 10 October 2026