IA al Día
the efficient way to stay informed
Models June 15, 2026 explainer 11 min read

When AI eats its own data: the synthetic content crisis and model collapse

The internet is filling with artificial intelligence-generated content at an unprecedented pace. This not only degrades the quality of the information we consume: it threatens to poison the very training data that future models depend on, creating a cycle of degradation that researchers have already named "model collapse."

By IA al Día

AI models are eating their own training data, and the result is a cycle of degradation that researchers call “model collapse.” Imagine an ecosystem where the waste of some organisms becomes the food of others — but that waste is increasingly toxic. That is what is happening with generative AI and the synthetic content flooding the internet.

The phenomenon has two inseparable faces. On one hand, an avalanche of synthetic content — articles, images, videos, comments, academic research — is inundating the web at a scale researchers are only beginning to quantify. On the other, AI models trained on that same synthetic data are beginning to show signs of progressive degradation, a process the academic literature has baptized as model collapse. The paradox is as unsettling as it is inevitable: the more AI we use to generate content, the worse the models we train on that content could become.

This article explores the available evidence, the mechanisms behind the phenomenon, and the implications for the future of digital information.

The silent flood

The magnitude of the problem is difficult to grasp. In 2024, a team of researchers at Amazon Web Services AI Labs analyzed a sample of more than six billion sentences extracted from Common Crawl — the repository that serves as the backbone of many AI training datasets — and found that more than 57% of the sentences had been automatically translated, with particularly low quality in those that had passed through multiple languages. This is not invisible traffic: it is the raw material with which the most advanced language models on the planet are being trained.

The phenomenon is not limited to automatic translations. In 2023, University College London estimated that more than 60,000 academic articles — exceeding 1% of all global publications that year — were probably written with assistance from language models. Stanford’s Institute for Human-Centered AI (Stanford HAI) raised the figure even higher: approximately 17.5% of computer science articles and 16.9% of peer reviews incorporate AI-generated content.

In journalism, the Pravda network — a constellation of disinformation sites — published up to 10,000 articles per day, many of them AI-generated, according to a 2025 report from the American Sunlight Project. The stated objective: to infiltrate pro-Russian narratives into the training data of the world’s language models.

Perhaps the clearest signal that the problem has reached critical mass is Robyn Speer’s decision, creator of wordfreq — a word frequency database used by computational linguists — to abandon the project in September 2024. Her explanation was brutally direct: “generative AI has contaminated the data.”

The era of “slop”

The phenomenon is so ubiquitous that it already has a name of its own. The word “slop” was selected as the 2025 Word of the Year by both Merriam-Webster and the American Dialect Society. The New York Times defines it as the digital equivalent of spam: “sloppy or unwanted AI content in social media, art, books, and search results.”

The examples are hard to forget. In 2024, AI-generated images of so-called “Shrimp Jesus” — depictions of Christ fused with shrimp — went viral on Facebook, accumulating hundreds of thousands of interactions. The engine behind this phenomenon is not ideological but economic: creators in developing countries produce low-effort images aimed at American audiences to maximize advertising revenue. A medical student in India told the press that he earned thousands of dollars per month with AI-generated images on Instagram.

Model collapse: when AI poisons itself

If synthetic content were just background noise, it would be a minor problem. But it is not: that same content is re-entering the training pipelines of the models that generated it, and the consequences are concerning.

In 2024, a team led by Ilia Shumailov published an article in the journal Nature that coined the term “model collapse.” The central finding is counterintuitive but deeply logical: when an AI model is trained exclusively on data generated by another AI model, its quality degrades. And if that process is repeated — each new model trained on the output of the previous one — the degradation accelerates until complete collapse.

Shumailov and his colleagues described two stages. The first is subtle: the model loses information about the tails of the distribution — minority data, exceptions, the infrequent. This loss is difficult to detect because overall performance can even appear stable or improve. But in the second stage, the damage becomes evident: the model begins to confuse concepts, loses variance, and its output becomes indistinguishable from repetitive noise.

The mechanisms behind collapse are three: functional approximation errors (the model cannot perfectly represent the underlying distribution), sampling errors (synthetic samples do not capture the true diversity of the original data), and learning errors (the model reinforces its own hallucinations). In complex models like large language models, these errors combine and accelerate each other.

Is it inevitable?

Not everyone agrees that model collapse is an imminent existential threat. More recent research suggests that if synthetic data accumulates alongside human-generated data (rather than completely replacing it), collapse can be avoided. These researchers argue that this mixed scenario is more realistic than Shumailov’s controlled experiment.

However, the warning is clear: the Wikipedia article on generative AI summarizes it unambiguously: “Training an AI model exclusively on the output of another AI model produces a lower-quality model. Repeating this process leads to progressive degradation and, eventually, complete model collapse after multiple iterations.”

Proposed solutions include AI-generated content detectors and watermarking systems that allow identifying and filtering synthetic content from training sets. But both approaches face considerable technical and practical challenges: detectors can be fooled, and watermarks can be removed.

No platform better embodies the contradiction of this crisis than Google. On the one hand, the company was one of the first to publicly acknowledge the problem. In 2024, a Google spokesperson admitted that search results were being “flooded with websites that seem created for search engines rather than people” and explicitly pointed to the role of generative AI in the proliferation of that content. The company launched its “Helpful Content Update” specifically to combat low-quality AI-generated content.

But at the same time, Google has deeply integrated generative AI into its flagship product. AI Overviews (originally launched as Search Generative Experience in May 2023) appeared in more than 48% of search queries in March 2026, a 58% year-over-year increase. When AI Overviews and featured snippets appear together, they occupy approximately 67% of the screen on desktop and 76% on mobile, pushing organic results far below the fold.

The feature was widely criticized for its inaccuracies. In May 2024, Google was forced to temporarily restrict the tool after it recommended users “eat rocks” and “apply glue to pizza.” In 2025, Chegg sued Alphabet over AI Overviews, alleging copyright infringement and loss of traffic. That same year, the European Commission announced an investigation to determine whether AI Overviews violated competition legislation.

The situation has led critics to describe Google as a digital “Potemkin village”: the indexable web is much smaller than it seems, and AI-generated content reduces the ability to find quality information. A 2025 study confirmed that synthetic content hinders the retrieval of high-quality information and degrades search engine indexes.

Dead Internet Theory: from conspiracy theory to documented phenomenon

One of the most unexpected twists in this story is that a theory that until recently was considered fringe — the Dead Internet Theory — has been gaining credibility in academic and technology spaces.

The original theory postulated that the internet consists mainly of automated activity and bot-generated content, possibly coordinated by government agencies. The conspiratorial version has never been verified, but a growing number of researchers have proposed a leaner version that removes the conspiratorial elements and focuses on observable evidence: algorithmically generated content, difficulty distinguishing between human and automated interactions, and a progressive erosion of trust in the digital ecosystem.

The supporting data is blunt. Imperva’s 2023 Bot Traffic Report found that 49.6% of all internet traffic was automated, 2% more than the previous year, attributed in part to data scraping for AI training.

But perhaps the most revealing aspect is who is recognizing the phenomenon. In September 2025, Sam Altman, CEO of OpenAI, tweeted: “I never took Dead Internet Theory seriously, but it seems like there really are a lot of Twitter accounts run by LLMs now.” The tweet went viral. A month later, Alexis Ohanian, co-founder of Reddit, declared at TechCrunch Disrupt that “the Dead Internet Theory is real.”

In January 2025, Connor Hayes, vice president of generative AI product at Meta, stated that the company expected AI accounts “to exist on our platforms, much like human accounts do… They will have bios and profile pictures.” The accounts were quickly removed after public backlash, but the strategic direction was clear.

The most dramatic case was that of Digg. The platform was relaunched in January 2026 by Alexis Ohanian and Kevin Rose, but closed in March 2026 due to an “unprecedented bot problem.” In May 2026, Digg was relaunched again… as an AI news aggregator.

Even academia has taken notice. Hal Berghel, in a 2026 article in the journal Computer, has developed an academic framework for the leaner version of the theory. Yoshija Walter published a 2024 academic article titled “Artificial influencers and the dead internet theory.” Linguist Adam Aleksic told Time in 2025 that the theory “used to be a lunatic conspiracy theory, but it looks more and more real every day.”

The vicious cycle

The most concerning aspect of this crisis is not any of these phenomena separately, but their interconnection. The four processes reinforce each other in a feedback loop:

  1. AI-generated content floods the internet, contaminating training data, search indexes, and social platforms.
  2. Search quality degrades, as synthetic content competes with legitimate content and generative interfaces (AI Overviews) cannibalize organic traffic.
  3. Model collapse threatens the future of AI, as new models are trained on increasingly synthetic data.
  4. Dead Internet Theory gains credibility, because the combination of automated traffic, junk content, and algorithmic curation makes it increasingly difficult to find genuine human interaction.

The central concern is that this cycle could become self-sustaining: more AI content → worse models → even worse AI content → faster model collapse → even more synthetic data → even harder for humans to find quality information.

Is there a way out?

Solutions exist in theory but are elusive in practice. The most immediate is the detection and filtering of synthetic content in training sets, combined with watermarking systems that allow tracing the origin of data. However, none of these mechanisms is infallible, and their effectiveness depends on widespread adoption that has not yet occurred.

Another line of research bets on preserving and prioritizing human-generated content, creating curated repositories that serve as safe havens for model training. Initiatives like the Common Corpus — a massive public domain dataset — point in this direction.

But perhaps the most important solution is the least technical: recognizing that the problem exists. For years, the AI industry has operated under the implicit assumption that training data was an infinite and renewable resource. The evidence suggests the opposite. Human-generated content — real conversations, texts written with genuine communicative intention, knowledge produced with rigor — is a finite and valuable resource. Treating it as inexhaustible is the surest recipe for exhausting it.


Learn more

  • Shumailov, I. et al. (2024). “AI models collapse when trained on recursively generated data.” Nature, 631, 755–759. — The foundational article on model collapse.
  • Imperva. (2023). Bad Bot Report 2023. — 49.6% of internet traffic is automated.
  • Walter, Y. (2024). “Artificial influencers and the dead internet theory.” AI & Society. — Academic framework for DIT.
  • Berghel, H. (2026). “The Leaner Dead Internet Theory.” Computer, 59(1). — Non-conspiratorial version of the theory.
  • Amazon Web Services AI Labs. (2024). Common Crawl analysis: 57%+ of sentences automatically translated.
  • University College London / Stanford HAI. (2023–2024). Estimates of AI use in academic publications.
  • American Sunlight Project. (2025). Report on the Pravda network and AI-generated content.