← Field Notes

AI Detectors Flagged 61.3% of Non-Native English Essays as AI

A Stanford study ran seven commercial detectors against TOEFL essays written under exam conditions. More than three in five were wrongly called machine-written. Native-speaker essays scored almost perfectly.

Seven commercial AI detectors were run against essays written by non-native English speakers under supervised exam conditions. The detectors called 61.3 per cent of them machine-written.

The same detectors scored essays by native-speaking US students almost perfectly.

The study is Liang and colleagues at Stanford, published in Patterns in 2023. The essays were TOEFL submissions, written by hand under invigilation, in a year that made machine authorship impossible for most of the sample. Every one of those flags was wrong.

Why the error lands on one group

The mechanism is not mysterious and it is not a bug anyone forgot to fix. It follows directly from what these tools measure.

Detectors score perplexity, which is a formal way of asking how surprising each next word is given the words before it, and machine-generated text scores low on that measure because a model samples near the most probable continuation rather than reaching for the unexpected one. So the tools reason backwards. Low perplexity, therefore probably a machine.

Now consider who else writes text that is easy to predict.

Someone working in a second language draws on a smaller and more consciously learned vocabulary, chooses standard constructions because those are the ones they are confident in, and avoids idiom because idiom is where the mistakes live. Every one of those choices is a good one. Every one of them lowers perplexity. The result is correct, clear, and exactly what a model would have predicted. That is the whole problem.

That is the same signature the detector is hunting for. The tool is not biased against non-native speakers by accident of training data. It is measuring a property that careful second-language writing has.

The uncomfortable part. Writing more clearly makes your detector score worse. Every plain-language guideline ever issued to government departments, medical publishers and technical writers pushes prose toward exactly the profile these tools flag.

Who else gets caught

The same logic sweeps up several groups who have nothing in common except a predictable prose style.

Technical writers. Documentation is supposed to be unsurprising. Consistent terminology and simple constructions are the job.

Anyone trained in plain language. Government and healthcare style guides mandate short sentences and common words, which is to say low perplexity by design.

Neurodivergent writers. Highly structured, consistent prose is a documented pattern for some autistic writers, and it reads to a detector exactly like sampling near the mode.

Students who were taught a formula. Five-paragraph essay structure was drilled into a generation. A model trained on the internet produces the same shape.

What the vendors say, and what happened next

OpenAI shipped an AI Text Classifier in January 2023 and withdrew it that July, citing low accuracy. The company with the most direct access to how its own models generate text could not build a reliable detector for them.

Most vendors now publish accuracy figures alongside caveats, and the caveats are the interesting part. The typical language recommends against using scores as the sole basis for an accusation. That recommendation exists because the vendors know what the false positive rate does to individuals.

Institutions read the number and skip the caveat. That is where the harm happens.

The base rate problem nobody mentions

Even an accurate detector produces mostly false accusations in a population where cheating is rare.

Suppose a detector is 95 per cent accurate in both directions, which is better than anything demonstrated in independent testing. Apply it to 1,000 essays where 50 were machine-written.

It catches about 48 of the 50 cheats, which sounds like a success until you work out what it does to the other 950 essays, where a 5 per cent error rate produces roughly 48 more flags. Same number. Innocent people.

So roughly half of everyone flagged did nothing wrong. That is with a tool performing better than any of them do, and it gets worse as cheating gets rarer. The arithmetic is not a criticism of any particular product. It is what happens when you apply an imperfect test to a population where the thing you are testing for is uncommon.

What to do if you are accused

Four things, in order.

  1. Ask which tool produced the score, and what its documented false positive rate is. Every vendor publishes one. Many institutions using these tools have never read it.
  2. Cite the Stanford figure. 61.3 per cent of TOEFL essays under exam conditions, published in Patterns, peer reviewed. It is a specific, checkable number about a specific, comparable situation.
  3. Produce your process, not your prose. Version history, drafts, notes, search history, the document's own revision log. Google Docs and Word both keep one, and a real draft has a mess in it that no accusation survives.
  4. Ask what other evidence exists. If the answer is only the detector score, the accusation rests entirely on a tool whose own maker recommends against using it that way.

What this means for the "AI slop" conversation

There is a real problem here, and none of the above says otherwise. Machine-written filler is flooding search results, book listings and video platforms at a scale that makes the complaint reasonable, and learning to recognise it is worth the effort.

But recognising it and detecting it are different activities, and only one of them works.

Slop is identifiable by properties you can point at. Sentence lengths that never vary. Vocabulary drawn from a measurable set of over-selected words. Paragraphs containing no claim a reader could check. Those are observable in the text, and a person can verify each one by reading.

A detector score is not observable. It is a probability produced by a model with a documented bias, and it cannot show you its reasoning.

That distinction matters if you write for a living. You cannot control what a detector says about your work, and chasing a better score means making your writing worse, since the tools reward unpredictability rather than clarity. What you can control is whether your writing carries the actual markers of unedited machine output.

The free workbench checks those markers rather than guessing at authorship. It flags the over-selected vocabulary in context, measures your sentence-length variance against the 0.60 human benchmark, counts your em dashes per thousand words, and scores what it finds. It does not tell you who wrote something. It tells you which habits are in the text, which is the only part you can do anything about.