← Field Notes

Why ChatGPT Uses So Many Em Dashes (And How to Make It Stop)

Models drop four to eight em dashes per page, which is roughly ten times the rate of edited human prose. Here is where the habit comes from, what the safe threshold is, and the instruction that actually removes it.

The em dash became the tell everyone knows, and it happened faster than any other marker on the list. Ask a room of editors how they spot machine writing in 2026 and this is the first answer.

The frequency is real and measurable. Unedited model output runs roughly four to eight em dashes per page, while edited prose from professional publications sits closer to one per 500 words. That is about a tenth of the rate.

What follows is where the habit comes from, what threshold actually reads as human, and the instruction that removes it. There is also a section defending the em dash, because the backlash has gone further than the evidence supports.

Where the habit comes from

Nobody trained a model to love this particular mark. It arrives through two doors.

The training corpus leans that way. Long-form web writing is heavily represented in training data, and it uses em dashes far more freely than print editing does, because blogs and Substack posts and marketing copy reach for them constantly while nobody stands behind the writer asking whether a full stop would work. Books are different. A copy editor gets involved, and they mostly come out.

Alignment rewards the shape they create. An em dash lets a sentence add a clause without committing to how that clause relates to the rest of it, which means the writer never has to decide whether the new material is subordinate, parenthetical or a separate thought entirely. It postpones a decision. During preference tuning, human raters working quickly tend to reward text that flows and elaborates, and the em dash is the cheapest way to keep elaborating without restructuring anything.

That second point explains why the habit is so persistent. The mark is not decoration. It is what a model reaches for when it wants to extend a thought and has not planned where the sentence ends.

The threshold that reads as human

One per 500 words. That is the number our lexicon enforces, and it comes from the rate at which edited prose actually uses them.

Some numbers for calibration. A 1,000 word blog post with two em dashes reads normally. The same post with nine reads as machine output, even to someone who could not tell you why. At fifteen, readers start commenting.

The tell is rate rather than presence, which matters because the internet has spent two years telling people that any em dash is proof of AI. That claim is wrong and it has consequences.

In defence of the em dash

The backlash overshot. Em dashes have been in English prose for centuries, and several writers built an entire voice on them, which is worth remembering before you strip every one out of a draft you wrote yourself.

Emily Dickinson used them structurally rather than decoratively, and Vonnegut used them for timing. Plenty of careful writers still reach for one when a parenthesis would be too quiet and a full stop too final. That is a legitimate choice.

People are now editing em dashes out of writing they did entirely themselves, because a reader accused them. Students are being marked down. That is the same failure mode as detector-based accusations, and it comes from treating a probabilistic signal as proof.

The honest version of this tell. A high em-dash rate is corroborating evidence, not a verdict. If a passage also has flat sentence rhythm and no checkable claims in it, the dashes are part of a pattern. On their own they mean nothing.

Why the model reaches for it, sentence by sentence

Look at what the mark is doing in typical output:

Our platform helps teams work faster — and it does this by removing manual steps — which means less time on admin.

Two dashes, and both are covering for a sentence the model had not finished planning. Rewritten with the decisions made:

Our platform removes manual steps, so teams spend less time on admin.

Shorter, and the relationship between the clauses is now explicit rather than implied by a mark that means whatever the reader decides it means. That is the general pattern. An em dash in machine writing marks the spot where a choice was avoided.

The four replacements

When you remove one, something has to take its place. Which one depends on what the dash was hiding.

A full stop when the clause is a separate thought. This is the right answer most of the time, and it also improves sentence-length variance, which is the deeper tell.

A comma when the clause is clearly subordinate and short.

Parentheses when the material is a genuine aside that the sentence would survive without.

A colon when what follows explains or lists what came before.

Run through a draft applying that decision to each dash in turn and you will find that perhaps one in ten survives the question, because the rest were standing in for punctuation that carries an actual meaning. Replace them.

Making it stop at generation

Editing dashes out afterwards works and takes about ten minutes per piece. Stopping the model producing them is faster and it holds.

The instruction has to be specific about the rate, because a request to "use fewer em dashes" gives the model nothing to count. This works:

Use at most one em dash per 500 words. Where you would use one,
      choose a full stop, a comma, parentheses or a colon based on the
      actual relationship between the clauses.
      

Two things make that instruction work where vaguer ones fail. It gives a countable limit rather than a preference, and it names the replacements, so the model has somewhere to go instead of the dash it would otherwise reach for. Specificity is the whole trick.

The same principle applies to every other marker. Models comply with constraints they can evaluate while generating, and ignore judgments about the finished product, which is why "write more naturally" produces nothing. A decoder cannot check naturalness. It can count dashes.

Where this sits in the bigger problem

Em dashes are the easiest tell to fix and the least important one.

They are visible and removable with a search. That is exactly why they became the popular tell, and also why fixing them alone changes almost nothing. A draft with zero em dashes and 16-word sentences throughout still reads as machine output, because rhythm is the signal readers actually register.

The free workbench measures both at once. Paste a draft and it counts your em dashes per thousand words against the threshold, calculates your sentence-length variance against the 0.60 human benchmark, flags the stock vocabulary in context, and scores the result out of 100. Most people find their dash count is fine and their variance is not.