Researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) warn that most AI-generated images cannot be reliably traced back to a single image in the training dataset. The team labels this phenomenon “attribution decay.”
Their analysis shows that as generative models are trained on larger datasets, it becomes increasingly difficult to attribute a generated image to an individual training example. The authors note that even if an artist’s works — for instance, Pablo Picasso’s — were removed from the training set, an AI-generated image might still resemble Picasso’s style, meaning removing particular works may not eliminate stylistic similarity.
Legal implications
AI developers have frequently used books, articles, photographs and other copyrighted works to train models, often without consulting or obtaining permission from rights holders. That practice has spawned several high-profile copyright lawsuits: last year Disney, NBCUniversal and DreamWorks filed an intellectual property suit against the image generator MidJourney; The New York Times sued OpenAI and Microsoft in 2023 for similar reasons.
According to the CSAIL study, attribution decay complicates legal questions about training data and fair use. If a generated image cannot be clearly linked to a single original work, proving specific infringement becomes more difficult. Zheng Dai, a former MIT researcher and the study’s lead author, told Semafor: “We might have to rethink what intellectual property means. You can’t just assume it, and the attribution link sort of vanishes.”
Consequences and open questions
The findings carry regulatory and judicial implications: courts and policymakers will need to consider how to interpret similarity produced by learning models in the context of copyright law. The study also highlights the challenge of balancing transparency about training datasets and protecting authors’ rights, without unduly hindering technological development.
The authors’ observations suggest that further dialogue between technologists, lawmakers and rights holders will be necessary to develop practical approaches to data provenance, accountability and usage rights in generative AI.



