← Field Notes

How to Spot AI Slop: 12 Tells, Ranked by Reliability

Not every AI tell is worth trusting. These twelve are ranked by how often they are right, starting with the two that survive editing and ending with the ones that get innocent writers accused.

Most guides to spotting AI writing hand you a word list and wish you luck. That fails twice over: it misses machine output that somebody bothered to edit, and it convicts humans whose only crime is writing clearly.

These twelve tells are ranked by how often they are right. The first four survive editing and rarely produce a false accusation, which matters because most machine-written text now passes through at least one round of human cleanup before anyone sees it. The last three appear on every listicle and deserve the least trust. Where a number exists, it is cited.

One rule before the list. No single tell is evidence. A paragraph carrying six of them is a different matter, and that is how the ranking should be used.

The reliable four

1. Flat sentence rhythm

The strongest signal, and the one almost nobody checks.

Human writing varies its sentence length constantly, because a person writing with intent lengthens a sentence to carry an argument and then cuts one short to land it. Unedited model output does not. Sentences cluster between 14 and 18 words and stay there for pages.

Measure it with burstiness: standard deviation of sentence length divided by mean sentence length. Edited human prose usually sits above 0.60. Raw model output lands near 0.30.

Do it by hand in thirty seconds. Take one paragraph. Count the words in each sentence and write the numbers in a row. Human prose gives you something like 8, 24, 5, 31, 12. Slop gives you 16, 15, 17, 16, 18.

This tell survives light editing because most people edit vocabulary, not rhythm. Someone who swaps out "delve" still leaves every sentence the same length.

2. Zero proposition density

Count the claims in a paragraph that a reader could check, disagree with, or act on.

Slop scores zero. It restates the topic, asserts that the topic matters, then moves on. The words are fine. The paragraph is load-bearing for nothing.

The test that catches it: ask whether a competitor could publish the same paragraph verbatim about their own product. If yes, it says nothing about yours.

Our platform helps modern businesses streamline operations and unlock efficiencies across the organisation.

Any company on earth could publish that sentence. Now compare:

Our tool assigns pull request reviewers in under ten seconds by reading your Jira board. Across a sixty-day trial with forty companies, review time fell 35 per cent.

Second one commits to numbers you could prove wrong. That is the difference.

3. Lexical density, not lexical presence

The word list is real. Kobak and colleagues analysed 15 million PubMed abstracts for Science Advances and found "delves" running at roughly 28 times its expected rate after 2022. Juzek and Ward put the same shift at 6,697 per cent across scientific abstracts between 2020 and 2024 in their COLING 2025 paper.

But presence proves nothing. "Robust" is the correct word in a statistics paper describing an estimator that tolerates outliers, and "underscores" is the right verb in a report about emphasis.

What matters is density. One flagged word in 300 is noise. Six in a single paragraph is a model.

A working sample of the list:

Verbs Modifiers Metaphor props
delve, underscore, showcase robust, seamless, multifaceted tapestry, beacon, testament
leverage, harness, foster comprehensive, pivotal, meticulous landscape, realm, cornerstone
unlock, elevate, empower ever-evolving, cutting-edge, transformative journey, frontier, ecosystem

4. The wrap-up that nobody asked for

A closing paragraph that summarises what you just read, then congratulates you for reading it.

"In conclusion, by understanding these key principles, you can take your writing to the next level."

No editor requests this. It appears because reward models rated it highly during training, so the habit is baked into the weights rather than the prompt. It is easy to spot and easy to delete, which is why it survives only in text nobody touched after generation. When you see one, the draft was published as it came out.

The strong middle four

5. Em-dash saturation

Models drop four to eight em dashes per page, using them where a full stop belongs.

The tell is rate, not presence. Above roughly one per 500 words is conspicuous. Plenty of good writers use em dashes deliberately, so treat this as corroborating rather than deciding.

6. The invented third item

Models bundle things into triads whether or not three things exist. Speed, security and scalability. Fast, efficient and reliable.

Look at the third item specifically. If it restates the second or could be deleted without loss, it was generated to complete a rhythm.

7. Negative parallelism

It's not just a tool, it's a catalyst for transformation.

The construction promises escalation and delivers a restatement, because both halves of the sentence carry the same claim with the second one dressed in bigger words. Once you notice it you will not stop seeing it in launch copy.

8. Throat-clearing openers

In today's fast-paced world. In an era where. When it comes to.

The first sentence contains no information and exists to warm up. Human writers under an editor lose these in the first pass, because an editor deletes them. A model has no editor.

The weak four, and why they get people accused

These appear in every "how to spot AI" listicle. Each one produces false accusations, so they belong at the bottom.

9. Perfect grammar

Clean grammar is not evidence of anything. Professional writers produce clean grammar. So do careful non-native speakers, who often write more correctly than natives because they learned the rules explicitly rather than by absorption.

10. Bullet points and bold text

Formatting is a house style, not a fingerprint. Technical documentation has looked like this since long before language models existed.

11. Being "too polished"

This is a vibe, not a tell. It is also the most common thing people say when they mean "I have a feeling", and feelings are exactly where false accusations come from.

12. A detector score

The weakest signal on this list, and the one treated most often as proof.

In a Stanford study published in Patterns in 2023, Liang and colleagues ran seven commercial detectors against TOEFL essays written by non-native English speakers under exam conditions. The detectors misclassified 61.3 per cent of those essays as AI-generated. Native-speaker essays scored almost perfectly.

The mechanism explains the bias. Detectors look for low perplexity, which is a formal way of saying text whose next word is easy to guess, and someone writing carefully in a second language produces exactly that. So does a technical writer. So does anyone ever taught to write plainly.

A detector score is a probability estimate from a tool with a documented bias against a specific group of people. Using one to accuse a student is not a defensible position.

Using the ranking

Score a passage rather than judging it. Three tells from the top four is a strong signal, six drawn from anywhere is a stronger one, and two from the bottom four is nothing at all. Weight what you find rather than counting it.

The reason for ranking rather than listing: tells 1 and 2 survive editing, while tells 5 through 8 vanish the moment someone runs a search-and-replace. That means the top of the list catches machine-written text that has been cleaned up, which is now most of it.

Checking your own writing

Reading for these in your own draft is harder, because your draft reads fine to you. That is what the tells are made of.

The free workbench on our homepage automates the mechanical ones. Paste a post or a chapter and it highlights flagged vocabulary in context, calculates burstiness against the 0.60 threshold, counts em dashes per thousand words, checks your opening line against known stock openers, and returns a score out of 100.

It runs about half our full lexicon and it does not rewrite anything. Diagnosis only. What you cut is your call.

Fixing what it finds

Tells 5 through 8 are search-and-replace work. Tells 1 and 2 are not, and those are the ones that matter.

You cannot fix flat rhythm by swapping words, because rhythm is a property of how the sentences were planned. You fix it by telling the model to vary sentence length before it writes, capping em dashes, banning the vocabulary outright, and requiring each section to carry a checkable claim. Constraints in that shape work. A request to "write more naturally" does not, because it gives the model nothing to act on at the token level.

Set those constraints once in Claude Projects, ChatGPT Custom Instructions, or a Cursor rules file, and the drafts stop arriving with these twelve tells in them.