Using personal data to train artificial intelligence (AI) models is increasingly common, but complying with data protection law is legally and practically challenging. Under GDPR Article 6 the three primary bases to consider are consent, performance of a contract with the data subject, and the legitimate interests of the controller or a third party. Each raises specific problems when training models on data collected without the data subject’s direct awareness.
The typical situation: ancillary data processing
In most cases the interaction between the data subject and controller concerns a service (for example using social media, buying in a webshop, a phone call with customer service, or prompts given to a chatbot), and the use of personal data for model training is incidental. Because training is usually not required by law, not necessary to protect a data subject’s vital interests, nor performed in the exercise of public authority, the viable legal bases to examine are consent, contract performance and legitimate interests.
a) Consent
Consent can provide strong legal grounds: GDPR guidance requires consent to be voluntary, specific and informed. Mechanisms to collect and later evidence consent (for example web popups) are technically feasible. However, wide‑scale web scraping complicates valid consent collection: organisations that scrape many websites often have no direct contact with data subjects and therefore cannot provide the specific information or obtain a demonstrable, informed consent.
There are also economic reasons against mandatory consent: AI training benefits from large datasets, so making consent opt‑in will likely reduce the amount of lawfully usable personal data and may degrade model performance. This may create a competitive incentive for AI developers to operate outside the EU where fewer constraints apply.
b) Performance of a contract with the data subject
Because training is often ancillary, it is usually not objectively necessary to perform the main contract between the parties. Some controllers may seek to include training‑related processing in their general terms (a “take it or leave it” approach) and claim Article 6(1)(b) as legal basis. The Court of Justice of the EU in Meta v. Bundeskartellamt (C‑252/21, 4 July 2023) ruled that processing is only covered by contract performance if it is objectively indispensable for a purpose integral to the contractual obligation. EU rules such as the Digital Markets Act similarly restrict how gatekeeper platforms can reuse data across services without offering a real choice to end users.
Analyses commissioned by the European Parliament and other legal commentary suggest that relying on contract performance for further uses like business analytics or later incorporation into predictive decision models is unlikely to be acceptable. The author of this article thus considers it improbable that controllers could lawfully convert broad training uses into a contractual legal basis merely by adding them to terms of service.
c) Legitimate interests
If consent is impractical and contract performance is doubtful, many controllers consider the legitimate interests basis. Interpretations across authorities differ. On 7 May 2024 X introduced an opt‑out for using posts and interactions to train its Grok chatbot and stated it relied on legitimate interests; this prompted regulatory scrutiny. The Irish Data Protection Commission brought the matter before Ireland’s Supreme Court arguing potential infringements of fundamental rights; X agreed to suspend the practice and proceedings were discontinued.
National authorities and the European Data Protection Board (EDPB) reach divergent conclusions. The French CNIL accepts that legitimate interests can be appropriate for some web scraping‑based AI training if strict data‑minimisation and safeguards are in place; the Dutch authority views commercial web scraping as rarely constituting a legitimate interest. The EDPB’s ChatGPT Taskforce warns that web scraping poses fundamental risks, and that controllers must perform and document a detailed legitimate‑interest assessment considering lawfulness, necessity and the balance with data subjects’ rights.
At present the legitimate interests basis is legally unsettled; an EU position may clarify matters, but in any case controllers will bear heavy documentation duties for case‑by‑case interest balancing.
Special categories of data
GDPR Article 9 generally prohibits processing of special categories of personal data. The AI Act (MI Rendelet) allows limited derogations for bias detection and correction in high‑risk systems under strict safeguards, but Article 9 remains the general rule. The EDPB Taskforce highlights that even the Article 9(2)(e) exception — processing data that the data subject has manifestly made public — is of limited practical use for web scraping, since intent to make data public is hard to establish at scale.
Both CNIL and the EDPB recommend designing training pipelines to exclude special categories from the input data and to promptly delete or anonymise any such data that is inadvertently collected. A further concern is that AI techniques can re‑personalise supposedly anonymous datasets, increasing the re‑identification risk and magnifying privacy harms.
Compatibility with original purpose (Article 6(4))
If training uses differ from the original purpose of collection, controllers may attempt to rely on the Article 6(4) compatibility test. The author argues that AI training often goes beyond statistical uses contemplated in the GDPR preamble and is unlikely to be found compatible in many cases, particularly where the controller did not collect the data directly and no clear original purpose exists. Thus an independent lawful basis for training will generally be required.
Technical mitigation: federated learning and data minimisation
Technical approaches can reduce reliance on personal data. Federated learning trains models locally on user devices and uploads only model updates to a central server, reducing centralised personal data collection. While promising for data minimisation, federated learning also poses risks (parameter leakage, poisoning attacks) and does not eliminate the need for legal and governance measures. Research anticipates progress in privacy‑preserving techniques, but those methods introduce new technical and legal challenges.
Conclusions and next steps
There is currently no uniform legal solution for using personal data to train AI models without data subjects’ knowledge. Consent is often impractical for large‑scale scraping, relying on contract performance appears legally narrow, and legitimate interests is contested across jurisdictions. Special‑category data remain tightly constrained and re‑identification risks must be taken seriously.
Until clearer jurisprudence or EU guidance is established, controllers should proceed cautiously: perform and document thorough legal assessments, apply strict data‑minimisation and governance measures, and consider privacy‑preserving technical alternatives. The next part of this series will examine privacy issues arising when AI systems operate in production.
Footnotes and sources referenced in this article correspond to the legal instruments and documents cited in the original analysis (GDPR Article 6 and 9, Meta v. Bundeskartellamt C‑252/21, EDPB and CNIL positions, AI Act provisions, etc.).


