Regulation

The black-box challenge: AI development, transparency and data-protection duties

The ‘black box’ problem in artificial intelligence complicates legal compliance during both training and deployment of AI systems, especially under the GDPR and the EU AI Act (Regulation (EU) 2024/1689).

The so‑called "black box" phenomenon in artificial intelligence describes systems whose inputs and outputs are observable, but whose internal decision‑making processes are opaque to outside observers. This opacity raises not only technical questions but also legal ones, notably under the European Union’s data protection framework — the (EU) 2016/679 Regulation (GDPR) — and the EU’s AI Act, Regulation (EU) 2024/1689 (the "AI Act").

Cultural and literary touchstones: Blade Runner, Turing, Dune

To illustrate the issue, commentators often cite Ridley Scott’s 1982 film Blade Runner and Philip K. Dick’s novel Do Androids Dream of Electric Sheep? The Void‑Kampff test depicted there uses physiological responses (pulse, respiration, pupil dilation, flushing) to try to distinguish humans from human‑like androids. The test draws part of its inspiration from Alan Turing’s Turing Test.

Frank Herbert’s Dune cycle similarly explores the risks of inscrutable machines: the human struggle against the Omnius collective superintelligence — which can predict outcomes with overwhelming accuracy because it has access to all past battle data — helps explain the canonical injunction against creating machines in the image of the human mind. These cultural references underline the risks that arise when a system’s inner workings cannot be understood or controlled.

How the EU frames AI

Article 3 of the AI Act defines AI systems as machine‑based systems capable of producing outputs such as predictions, content, recommendations or decisions by deriving models or algorithms from input data. From that perspective, the human‑like androids in the cited works fit the definition: they act autonomously, consume sensory input, and produce outputs that affect their environment.

Why opacity matters for data protection

The black box problem matters from a data‑protection standpoint for two linked reasons. First, opacity prevents transparent explanation of how an AI reaches specific conclusions or learns over time — sometimes even for developers. Second, AI training typically involves personal data, so opacity can conflict with GDPR principles such as transparency, purpose limitation, and lawful processing, and it complicates meaningful notice to data subjects.

Difficulties in specifying the processing purpose

Under the GDPR, an appropriate processing purpose must be specified, clear and lawful. When personal data are used to train AI, satisfying these three requirements is often challenging:

  • "Specified": In practice, collected datasets are frequently reused for ongoing or future development rather than a single training instance. If future uses, improved models or subsequent AI solutions cannot be precisely described at the time of collection, it becomes questionable whether the original purpose was truly "specified".

  • "Clear": If the purpose is vague, the information provided to data subjects will not be clear. The European Data Protection Board (and its predecessor, the Article 29 Working Party) has emphasised that transparency must enable individuals to foresee the scope and consequences of processing; generic statements such as "we may use your data to develop new services" or "for research purposes" are inadequate for complex or technical processing.

  • "Lawful": Training AI may engage other legal fields and sectoral rules. If an AI’s eventual use is unlawful — for example if it produces manipulative content, targets minors in breach of separate rules, or otherwise violates sectoral law — this may retroactively call into question whether the processing purpose was lawful at the outset. Whether unlawful application of an AI automatically renders the training data processing unlawful, or whether this is a matter for evidentiary and expert assessment, is likely to be contested in courts and by regulators.

Transparency in both regulatory regimes

Both the GDPR and the AI Act elevate transparency as a core requirement. The AI Act enumerates ethical principles including human agency and oversight; technical robustness and safety; privacy and data governance; transparency; diversity, non‑discrimination and fairness; societal and environmental well‑being; and accountability. Ensuring AI development and deployment adhere to these principles implies interpreting AI rules together with data‑protection obligations.

Conclusions and next steps

The black box phenomenon — simultaneously a technical, legal and societal issue — complicates efforts to make AI development GDPR‑compliant. Clearly defined purposes, meaningful notice to data subjects, and assessing lawfulness in a cross‑disciplinary way are essential. This article serves as the first part of a series: subsequent pieces will examine privacy concerns specific to AI training datasets and the data‑protection challenges that arise during operational use of deployed AI systems.