Research

Study finds about one-third of new websites by 2025 may be AI-generated

Researchers from Stanford University, Imperial College London and the Internet Archive analyzed Internet Archive data and estimate that roughly 35% of webpages published by mid-2025 are either AI-generated or AI-assisted.

Researchers from Stanford University, Imperial College London and the Internet Archive analyzed how generative artificial intelligence that emerged in 2022 has reshaped the World Wide Web. Their work used Internet Archive data and covered a 33-month period from August 2022 through May 2025.

Key findings

  • The team estimates that by mid-2025 roughly 35 percent of newly published web pages can be classified as AI-generated or AI-assisted. By the end of 2022 that share was effectively near 0 percent.

  • The study began from the so-called “dead internet” theory, which posits that a substantial portion of online communication is now attributable to bots and automated actors.

  • The researchers evaluated multiple AI-detection tools; among the detectors they tested, Pangram v3 produced the most accurate results for identifying AI involvement in webpages.

Research questions and results

The authors posed six specific questions about the consequences of AI-generated text — for example, whether viewpoints would narrow, whether hallucinations would trigger a surge in disinformation, and whether unique authorial voices would disappear. Analyzing 33 months of data, they found clear evidence for only two effects:

  • A downward trend in the semantic density of online texts.
  • An overall shift toward a more positive tone in web writing.

Other hypothesized effects, such as a sharp increase in disinformation or the dominance of a single stylistic voice, were not clearly supported by this analysis.

Expert comments

Jonáš Doležal, a Stanford researcher, told 404 Media that the speed of AI's adoption on the web is striking: after decades in which humans shaped most online content, generative AI achieved noticeable presence within roughly three years.

Németh Gábor, lead AI advisor at Stylers Group, commented in HVG’s recent AI publication that the real problem with much AI-produced content is not that it is maliciously bad but that it is mediocre — not intentionally harmful, simply not very good.

Implications and next steps

The study’s authors note that the internet has not definitively become a less truthful place, but they qualify that judgment: current detection algorithms may still miss sophisticated AI-generated falsehoods, and it is also possible the internet never had truth as a dominant attribute.

The research team plans to continue by examining in greater detail how AI has altered the qualitative and stylistic characteristics of texts available online.