← Field Notes

Words ChatGPT Overuses (With the Replacements That Actually Work)

Thirty of the most over-selected words in machine writing, each with the measured multiplier and a concrete replacement. Plus why swapping words alone will not fix your drafts.

Language models do not pick words evenly. They over-select a small set, and the effect is large enough that researchers can measure it against pre-2022 baselines and tell you the multiplier.

Below are thirty of the worst offenders, each with a replacement that carries more information than the word it replaces. The multipliers come from published corpus work, cited at the bottom.

Then the part most word lists leave out: why swapping these words will not, on its own, fix your writing.

The verbs

Verbs are where models do the most damage, because a vague verb hides the absence of a mechanism.

Word Measured Use instead
delve / delves into 28x expected rate (Kobak); +6,697% (Juzek) examine, dig into, or just say what you found
underscore / underscores 13.8x (Kobak); +904% (Juzek) show, prove, highlight
showcase / showcasing 10.7x (Kobak); +1,396% (Juzek) show, demonstrate, display
leverage / leveraging corporate register marker use, apply, exploit
harness same family as leverage use, channel, run on
foster / fostering vague non-action verb build, encourage, cause, support
unlock / unleash infomercial register enable, start, allow, release
elevate status inflation raise, improve, lift
empower hollow in most business use let, allow, equip, pay for
streamline process cliché simplify, cut steps from, speed up
facilitate almost always deletable help, run, host, arrange
navigate metaphor doing no work handle, cross, deal with, get through
embark on melodramatic start, begin
shed light on worn metaphor explain, reveal, show

How to use this table. Read the replacement column and notice what changes. "Facilitate a meeting" becomes "run a meeting", and now you know who was in charge. "Leverage our platform" becomes "use our platform", and the sentence has lost nothing except a suggestion of sophistication that was doing no work.

The modifiers

Adjectives are where slop hides its lack of evidence. Each of these asserts a quality without measuring it.

Word Why it flags Use instead
robust rarely means anything outside statistics reliable, or state the failure rate
seamless / seamlessly almost always untrue quick, automatic, or name the step removed
multifaceted says "complicated" at four syllables complex, or list the facets
comprehensive claims completeness without proof complete, or say what is covered
cutting-edge / state-of-the-art dates badly, proves nothing new, or name the version
transformative / groundbreaking superlative without evidence say what changed and by how much
pivotal / crucial asserts importance important, or explain the consequence
meticulous self-congratulation careful, or describe the check performed
invaluable / unparalleled unfalsifiable useful, or give the number
ever-evolving filler modifier changing, or say what changed
vibrant / dynamic atmosphere words busy, fast, growing
nuanced often means "I did not explain it" complicated, or state the distinction

The metaphor props

These are the ones readers mock, because they arrive in contexts where no metaphor was needed.

Word Measured Use instead
tapestry / rich tapestry "vibrant tapestry" at ~17,000x baseline (Pangram) mix, range, or just list the things
testament / stands as a testament "serves as a testament" ~4,000x (Pangram) shows, proves
beacon faux-reverent example, model, guide
landscape metaphorical crutch market, field, industry
realm pompous framing field, area

What a fixed paragraph looks like

Take a real sentence carrying six of these:

Our robust platform leverages cutting-edge technology to seamlessly empower teams to navigate the ever-evolving digital landscape.

Word-swapping alone gets you here:

Our reliable platform uses new technology to quickly let teams handle the changing digital market.

Grammatically fine. Still says nothing, because the problem was never the vocabulary. Every noun in that sentence is unspecified.

The actual fix requires knowing what the thing does:

Our tool reads your Jira board and assigns pull request reviewers in under ten seconds. In a sixty-day trial with forty companies, review time fell 35 per cent.

That is the difference between a word list and an editing method. A list tells you which words to suspect, which is useful for about ten minutes and then stops being the bottleneck, because the sentence above failed on evidence rather than vocabulary. No list can tell you what your product does.

The test that matters. After you swap the words, ask whether a competitor could publish your paragraph verbatim about their own product. If they could, the paragraph is still slop, whatever vocabulary it now uses.

Why swapping words is not enough

Three reasons the list on its own will disappoint you.

Rhythm survives the swap. Replacing "delve" with "examine" leaves your sentence the same length. Machine writing clusters between 14 and 18 words per sentence and stays there, which readers register as monotony before they notice a single word. Edited human prose scores above 0.60 on sentence-length variance. Raw model output sits near 0.30. No amount of vocabulary editing moves that number.

The structures survive too. The invented third item in a list. The "not just X, it's Y" construction. The closing paragraph that summarises what you just read. None of those are words, so none of them appear on a word list.

Density is untouched. A paragraph with zero checkable claims still has zero after you improve its vocabulary. This is the deepest tell and the only one that cannot be faked, because it requires you to know something.

Doing it before the draft instead of after

Editing these out afterwards costs about 45 minutes per piece. The cheaper order is to stop the model producing them.

Models follow constraints of a specific shape. Give one an explicit list of banned tokens and it complies. Tell it to vary sentence length deliberately, cap em dashes at one per 500 words, and require each section to carry a claim a reader could check, and it complies with those too.

Ask it to "write more naturally" and nothing happens, because natural is not something a model can act on at the token level. It is a judgment about output, not an instruction about generation.

The thirty words above are a working subset. Our full lexicon runs to 311 terms with replacements for each, and the free workbench runs about half of them against your own writing, flags what it finds in context, measures your sentence variance against the human threshold, and returns a score. Paste in a post you have already published. The result is usually informative.