The usual explanation is that people got lazy. It is the wrong answer, and believing it means you will keep producing slop while thinking you are one of the good ones.
Four separate mechanisms produce it. Only one of them involves anyone being lazy. The other three would fill the internet with this stuff even if every writer using a model were conscientious.
1. The cost of a page fell to nothing
For thirty years, publishing a mediocre article cost somebody two hours. That floor did the quality control. Nobody wrote 400 pages of filler about pet insurance because nobody could afford to.
The floor is gone. A page now costs a fraction of a cent and about eight seconds, which changes the arithmetic of every content strategy that was ever throttled by writer capacity.
When the marginal cost of a bad page approaches zero, the rational number of bad pages approaches infinity. That is not a moral failure. It is what happens to any market when a production constraint disappears, and it explains the volume better than anything about individual writers.
Google named the resulting behaviour in its spam policies: scaled content abuse, generating many pages mainly to manipulate rankings. Worth noting what the policy does not ask. It does not ask who wrote the page. It asks whether the page says anything.
2. Models were trained to write this way on purpose
This is the part most explanations miss, and it is the one that matters if you use a model yourself.
A base language model does not write in the slop register. That register is added during alignment, when human raters score candidate responses and the model learns which style earns approval.
Raters, working quickly and at volume, reliably prefer answers that run longer, signpost their own structure, hedge, sound enthusiastic and close with a summary. Those preferences become the reward signal. The model learns the register that scores well.
Juzek and Ward tested this directly in their COLING 2025 paper. They checked whether the vocabulary spike could be explained by model architecture, by algorithm choice, or by the training data itself. None of those accounted for it. Comparing preference-tuned models against their base versions did, which points at alignment as a driver rather than the raw corpus.
The practical consequence is uncomfortable. The habits you dislike are not a bug that a better model will fix. They are the thing the model was optimised to produce, which is also why asking for "more natural writing" fails. Natural is not a token-level instruction. A ban-list is.
3. The corpus is eating itself
Models trained after 2023 learn from an internet that already contains a great deal of model output.
Kobak and colleagues measured the penetration in one corpus. Analysing 15 million PubMed abstracts for Science Advances, they estimated that at least 13.5 per cent of 2024 biomedical abstracts showed signs of LLM processing. In some subcorpora the figure reached 40 per cent.
That is peer-reviewed academic writing, the most heavily gatekept text humans produce. Whatever the number is for marketing blogs, it is higher.
So the tells reinforce themselves. A model that overused "delves" in 2023 produced text that became training data for the model that overuses it more in 2025. The vocabulary spike is not decaying, which is what you would expect if this were a passing fashion.
4. Distribution pays for volume, not quality
Every major platform pays out on engagement, and engagement is measured in impressions before anyone assesses whether the impression was worth having.
Upload a hundred videos and a few will catch. The other ninety-odd cost almost nothing to make, so the expected value of flooding the queue stays positive even when the hit rate is dismal. The same arithmetic drives the pages and the books.
Platforms have started pushing back, and the timeline tells you when the volume became unmanageable:
- Amazon capped KDP uploads at three titles per day and now requires authors to declare AI-generated content
- YouTube tightened monetisation rules around mass-produced and repetitious uploads
- LinkedIn added a native "seems like AI slop" report control
- Google wrote scaled content abuse into its spam policies as a named violation
Each of these is a platform admitting the incentive it created worked too well.
Why Facebook looks worse than everywhere else
Facebook comes up more than any other platform in this question, and the reason is structural rather than generational.
Facebook's recommendation surface pulls heavily from accounts you do not follow, which removes the social cost of posting garbage. On a network where distribution depends on your existing followers, posting forty AI images a day costs you those followers. On a discovery-driven feed it costs nothing and occasionally pays.
Add an audience less practised at recognising generated images, and image slop finds its cheapest distribution there. The pattern is about the feed algorithm rather than about who is using it.
When did this start
The word arrived before the flood. "Slop" appeared on 4chan and Hacker News around 2022, aimed at AI images from DALL-E and Midjourney. Simon Willison argued in May 2024 that it should become the standard term the way "spam" did, and it did. By 2025 both Merriam-Webster and the American Dialect Society had named "slop" their Word of the Year.
The text flood tracks ChatGPT's public release in late 2022, with the measurable vocabulary shift appearing in 2023 corpora and becoming unmissable through 2024.
Why it actually matters
Three consequences, none of which are aesthetic.
Search degrades. When a query returns eight pages that restate the question, the answer is somewhere on page four. That cost is paid by everyone, including people who never touch a model.
Trust transfers. Readers who get burned start assuming machine authorship everywhere, and they are wrong often. Careful non-native English speakers get accused constantly, because clear writing in a second language produces exactly the statistical profile detectors flag. The Stanford study in Patterns found seven commercial detectors misclassifying 61.3 per cent of TOEFL essays as AI-generated.
Good writing gets harder to sell. If a client can get eight hundred passable words for nothing, the market for eight hundred good words has to explain what the difference is. That explanation is now part of the job.
When does it stop
It does not, and expecting it to is the wrong frame.
Spam did not stop. Filtering got good enough that most people forgot how much of it arrives. The same thing is happening here, on two fronts. Platforms are building detection and demotion into ranking, which is what the Amazon caps and the Google policy are. Readers are building their own filters, which is why a stock phrase now reads as a warning rather than as competent prose.
That second front is the one that concerns anyone publishing. The threshold for "this was not written by a person who cared" keeps dropping. Text that passed in 2023 gets clocked in 2026, because the audience has read a great deal more of it since.
Which leaves a practical question rather than a philosophical one: does your writing carry the signatures your readers have learned to spot?
The free workbench on our homepage answers it directly. Paste a post or a chapter and it flags the vocabulary, measures sentence variance against the human threshold, then scores the result. If the number comes back low, you are already writing past the filter. If it comes back high, you now know which of the four mechanisms above is showing up in your drafts.