News
One in Three Websites Created Since ChatGPT's Launch Is Written by AI, Pew Research Finds

An analysis of nearly half a million English-language websites by Pew Research Center finds that 35 percent of pages published after ChatGPT's debut show signs of AI authorship, with .com domains generating up to ten times more such content than .edu or .gov sites.
Contents
The American research organization Pew Research Center has published an analysis showing how deeply artificial intelligence has changed the way content is created online. Among websites published after ChatGPT's launch in November 2022, fully 35 percent show clear signs of being written or co-written by AI.
The study was conducted by a Pew Research Center team led by Samuel Bestvater, a senior data scientist at the organization, along with Aaron Smith, director of Data Labs research, and colleagues Carson TerBush, Chris Baronavski and Janakee Chavda. The analysts drew on data from Common Crawl, a nonprofit organization that maintains a web archive dating back to 2008, and used the Open Pangram tool to detect text generated by language models.
Scale of the phenomenon
The sample of nearly 500,000 English-language pages collected in July 2026 shows that 10 percent of all analyzed pages, regardless of publication date, display significant signs of AI authorship. That share rises sharply when looking only at content created after ChatGPT's debut, where it reaches 35 percent. The upward trend began in late 2022 and has held steady ever since, with no sign of slowing.
The authors note that the phenomenon is distributed very unevenly depending on the type of domain. About one in ten .com websites shows traits of AI-generated text. That is twice the rate seen on .org domains, where the figure stands at 4.6 percent, and roughly ten times the rate on .edu and .gov domains, where it hovers around 1 percent.
How AI text was identified
The researchers did not rely on a single detection tool alone. They also analyzed stylistic features that statistically appear more often in text generated by language models. These included more frequent use of connecting dashes, a higher frequency of the Oxford comma in lists, characteristic vocabulary such as delve, interplay, testament, bolstered and pivotal, and parallel constructions along the lines of not only X, but also Y.
The scale of the shift is striking. The frequency of connecting dashes rose from 5.79 per 10,000 words in January 2023 to 11.19 in January 2026, nearly double. Use of the Oxford comma increased by 63 percent over the same period, and the frequency of AI-typical vocabulary more than doubled, from 11.94 to 28.15 occurrences per 10,000 words. The steepest rise, nearly threefold, was in the frequency of negative-parallelism constructions.
What this means for web quality
The authors point out that AI's impact on the internet is not limited to the sheer number of pages. Language models trained on increasing amounts of text generated by other language models can entrench repetitive stylistic patterns and narrow the web's linguistic diversity. This phenomenon is sometimes called a training-data feedback loop: the more AI content flows into the internet, the more of it ends up in the training sets of subsequent models.
Pew Research Center's findings feed into a broader debate about the reliability of online content, one unfolding alongside AI's growing role in search and recommendations. Commercial sites, which feel the strongest pressure to churn out content for SEO, turn out to be the most prone to heavy use of text generators, while academic and government institutions keep a much greater distance.
Implications for Polish users and businesses
The study covers only English-language content, but the trend matters for the Polish market too. Companies running websites, newsrooms and marketing agencies are increasingly using text generators to scale up publishing quickly, raising questions about the quality of information reaching readers and about how search engines and AI assistants will treat content written by other AI models in the future. The growing share of machine-generated text also makes it harder for readers to tell human-written material apart from automatically generated content.
The authors caution that the detection method is reliable when analyzing large datasets, but individual documents can be misclassified. Despite that caveat, the scale and consistency of the trend, unbroken since late 2022, suggest that AI's share of online content creation will keep growing in the years ahead.


