← Field Notes

What Is AI Slop? The Measurable Definition

AI slop is machine-written text that is fluent, confident and says nothing. Here is where the word came from, the four signatures that identify it, and how to check whether your own writing has them.

AI slop is machine-generated content that reads fluently and carries almost no information. It is grammatically clean. It is confidently written. And when you finish a paragraph of it, you cannot say what you learned.

That last part is the whole definition. Bad writing is hard to read. Slop is easy to read and empty, which is why it slips past editors and past the person who wrote it.

The word is now measurable rather than a matter of taste. Researchers have counted which words spiked after ChatGPT shipped, and by how much. This piece covers where the term came from, the four signatures that identify slop in text, what it is often confused with, and how to check your own drafts.

Where the word came from

"Slop" started as forum slang. It appeared on 4chan and Hacker News around 2022, aimed at the flood of AI images arriving from DALL-E and Midjourney.

It went mainstream on 8 May 2024, when programmer Simon Willison argued in a blog post that the internet needed an ugly, widely understood word for unwanted AI output, in the same way "spam" became the word for unwanted email. The comparison stuck because it does real work. Spam is not defined by being badly written. It is defined by being unrequested and worthless to the reader. Slop is the same idea applied to content.

Google's rollout of AI Overviews in the second quarter of 2024 accelerated adoption. People who had never read a word about language models suddenly had a word for the thing at the top of their search results.

By 2025 the dictionaries had caught up. Merriam-Webster and the American Dialect Society both named "slop" their Word of the Year. Australia's Macquarie Dictionary picked "AI slop" specifically.

The four signatures

Calling something slop is an accusation. These four signatures are what make it a checkable one.

1. Lexical spikes

Language models do not pick words evenly. They over-select a small set, and the effect is large enough to show up in corpus statistics.

Kobak and colleagues analysed 15 million PubMed abstracts and published the result in Science Advances in 2025. The word "delves" appeared at roughly 28 times its expected rate after 2022. Juzek and Ward measured the same shift as a percentage across scientific abstracts between 2020 and 2024 and reported "delves" up 6,697 per cent, with "underscores" up 904 per cent, in their COLING 2025 paper.

Phrases behave the same way. Pangram Labs clocked "vibrant tapestry" at about 17,000 times its human baseline frequency.

The list runs to a few hundred terms. A representative sample:

Verbs Modifiers Metaphor props
delve, underscore, showcase robust, seamless, multifaceted tapestry, beacon, testament
leverage, harness, foster comprehensive, pivotal, meticulous landscape, realm, cornerstone
unlock, elevate, empower ever-evolving, cutting-edge, transformative journey, frontier, ecosystem

None of these words is bad. "Robust" is the correct word in a statistics paper. The signal is density: when six of them appear in one paragraph, a model chose them, not a person.

2. Flat rhythm

Human writing varies. A twenty-eight word sentence carrying an argument, then four words that land it.

Unedited model output does not do this. Sentences cluster between 14 and 18 words and stay there for pages. Statisticians call the measure burstiness, calculated as the standard deviation of sentence length divided by the mean. Edited human prose usually sits above 0.60. Raw model output tends to land near 0.30.

This is the signature most readers feel without being able to name it. The prose is competent and somehow exhausting, because nothing in the rhythm tells you which sentence matters.

3. Structural tics

Three patterns repeat so reliably they function as a fingerprint.

The rule of three. Models bundle items into triads whether or not three items exist. "Speed, security and scalability." "Fast, efficient and reliable." Watch for the third item that adds nothing, invented to complete a rhythm.

Negative parallelism. The "not X, it's Y" construction. It's not just a tool, it's a catalyst for transformation. The shape promises escalation and delivers a restatement.

The wrap-up. A closing paragraph that summarises what you just read and congratulates you for reading it. No human editor asks for this. Models produce it because reward models liked it.

4. Zero proposition density

The deepest signature, and the hardest to fake your way past.

Take any paragraph and count the claims a reader could check, disagree with, or act on. Slop scores zero. It restates the question and gestures at importance. It commits to nothing.

Test it on any paragraph. Ask what it asserts that a competitor could not have written verbatim about their own product. If the answer is nothing, the paragraph is slop regardless of which words it used.

Compare:

Our platform helps modern businesses streamline their workflows and unlock new efficiencies across the organisation.

Our tool watches your Jira board and assigns pull request reviewers in under ten seconds. In a sixty-day trial with forty companies, code review time fell by 35 per cent.

The first is not badly written. It is well-formed and correctly punctuated. It carries no information. That is slop.

What AI slop is not

Two confusions cause real damage, so both are worth stating plainly.

Slop is not "text a detector flagged." Commercial AI detectors are unreliable in a specific and unfair direction. In a Stanford study published in Patterns in 2023, Liang and colleagues ran seven detectors against essays written by non-native English speakers under exam conditions. The detectors misclassified 61.3 per cent of those essays as AI-generated, while scoring native-speaker essays almost perfectly.

The reason is mechanical. Detectors look for low perplexity, meaning text that is statistically predictable. Someone writing clearly in a second language produces exactly that. So does a careful technical writer. A detector score is not evidence, and treating it as evidence has already cost students their grades.

Slop is not "written with AI." A model can draft something and a person can then cut the padding, break the rhythm, add claims. What comes out is not slop. The failure is publishing the first draft unedited, which is a process problem rather than a tool problem.

Where you run into it

YouTube. The tell is retention rather than vocabulary. An AI script opens by restating the title and promising what is coming. Viewers leave during the intro, the retention curve drops in the first thirty seconds, and the algorithm stops recommending the video. YouTube has also tightened monetisation rules around mass-produced and repetitious uploads.

Search results. Google's spam policies name scaled content abuse directly: generating many pages mainly to manipulate rankings. The policy does not ask who wrote the page. It asks whether the page says anything.

LinkedIn. The platform added a native "seems like AI slop" report control, which tells you how common the complaint became.

Books. Amazon capped KDP uploads at three titles per day and now requires authors to declare AI-generated content. Readers spot it without a detector, usually through repeated stock gestures and chapters that end by explaining themselves.

Code. "AI slop in coding" usually means one of two things. Either generated code that compiles and quietly does the wrong thing, or the docstring disease: a comment explaining that parse_config "plays a pivotal role in facilitating robust configuration management" instead of saying what it returns when the file is missing.

How to check your own writing

Reading for slop in your own draft is hard, because it reads fine. That is the trap. Four checks, in the order that catches most:

  1. Count sentence lengths for one paragraph. If every sentence lands between 14 and 18 words, the rhythm is machine-flat regardless of vocabulary.
  2. Count em dashes. More than one per 500 words is conspicuous. Models drop four to eight per page.
  3. Circle the banned vocabulary. Any of the terms in the table above, plus their variants.
  4. Run the proposition test on your opening paragraph. What does it claim that a reader could check?

The free workbench on our homepage does the first three automatically. Paste a blog post or a chapter. It highlights the terms it finds, measures sentence variance against the human threshold, and scores the result out of 100. It runs about half our full lexicon and it is a diagnostic rather than a rewriter. It tells you what is wrong. The judgment about what to cut stays yours.

How to stop producing it

Editing slop out of a finished draft costs about 45 minutes per piece. Most people do it once, find it tedious, then stop.

The alternative is to condition the model before it writes. That means giving it an explicit ban-list rather than asking it to "write naturally", a hard cap on em dashes, an instruction to vary sentence length deliberately, and a requirement that each section carry at least one checkable claim. Models follow constraints of that shape reliably. They ignore vague requests for better prose, because "better" is not a token-level instruction.

Set the constraints once, in Claude Projects or ChatGPT Custom Instructions or a Cursor rules file, and the drafts arrive clean. That is the difference between fixing slop and not generating it.