← Field Notes

The AI Script Retention Cliff: Why Your Video Dies in 30 Seconds

An AI-written script fails before anyone judges the writing. It opens by restating the title and promising what is coming, which is exactly when viewers leave and the algorithm stops recommending you.

A written article survives a slow opening. A reader who is unimpressed keeps scanning, because scanning is free.

Video does not work like that. Leaving is free instead, and the first thirty seconds decide whether the algorithm ever shows the video to anyone else. That makes the opening of a script the only part that has to be right.

Which is a problem, because the opening is exactly where a language model does its worst work.

What the model writes, and why it kills you

Ask a model for a video script and the first fifteen seconds come back as some arrangement of these:

Hey everyone and welcome back to the channel. In today's video we're going to be diving into seven productivity habits that will completely transform the way you work. This is something a lot of people struggle with, so make sure you stick around to the end, and if you find this helpful, don't forget to like and subscribe.

Count what happened. The video restated its own title, promised value, established that the topic is common, and requested engagement.

Zero of the seven habits were mentioned. A viewer who clicked because they wanted the habits has now waited fifteen seconds and received nothing except confirmation that they clicked the right thing.

The model is not being lazy. It is doing what its training rewards. Signposting scored well with human raters, so a model opens by explaining what it will do, which works in a document with a table of contents and fails completely in a medium where the audience can leave.

What retention looks like when this happens

The curve has a distinct shape and every creator recognises it.

It starts at 100 per cent, drops steeply through the first twenty to thirty seconds, then flattens into a long slow decline. Typical numbers on an unedited AI script: retention at 0:30 somewhere near 30 per cent, average view duration under a minute on a twelve minute video, and click-through falling on subsequent uploads as the channel's recent performance drags impressions down.

The cliff is not spread across the video. It is concentrated exactly where the introduction runs.

Check yours in one minute. Open Studio, go to any video's audience retention graph, and look at where 0:30 falls. If you have lost more than half your viewers before the first real piece of content, the script is the problem and no amount of editing later in the video will recover it.

The cold open

The fix is structural and it is one rule: start inside the content.

Not a greeting. Not a description of the video. The first sentence should be something the viewer came for, or a specific claim that makes the rest necessary.

Here is the same video, opened cold:

Habit four is the one that got me back four hours a week, and it is the one nobody talks about, because it makes you look unresponsive. I turned off every notification between nine and one. Here is what happened to my output, and the three that came before it.

Now the viewer is thirteen seconds in and holding two specific things: a number and an unresolved question. The greeting can happen at 0:40, once they have a reason to stay for it.

This is not a new idea. It is how documentaries and news packages have opened for decades. It is unfamiliar to models because their training corpus is mostly written text, where a slow introduction is normal.

The other three script tells

The opening does the most damage. Three more markers do the rest.

Spoken cadence that never varies. Text with uniform sentence lengths reads flat. Spoken aloud it becomes hypnotic in the bad way, because the listener loses the signal that tells them which sentence matters. Human speech varies wildly. A long explanatory run, then three words.

The section summary. A model closes each segment by restating what the segment covered. On the page it is padding. On video it is the point where the viewer decides they have got the idea and leaves.

Vocabulary that nobody says out loud. "Delve into." "A myriad of." "It is worth noting that." These read as neutral and sound like a press release, because nobody talks that way. Read any script aloud and the borrowed register becomes obvious in a way it never is on screen.

The platform side

YouTube tightened its monetisation rules around mass-produced and repetitious content, which is the policy language for the channels uploading generated videos at volume.

That policy matters less than the algorithm for most creators, because the algorithm gets you first. A channel does not need to be demonetised to fail. It needs its retention to drop below the threshold where the recommendation system stops taking a risk on the next upload, and that happens quietly, without a notification.

Fixing a script before you record

Read it aloud. That single step catches most of it, because the ear rejects borrowed register faster than the eye does, and because you will hear the flat cadence immediately.

For the mechanical markers, paste the script into the free workbench. It flags the over-selected vocabulary in context, measures sentence-length variance against the 0.60 human threshold, and scores the result. Spoken scripts should sit well above that number, since speech varies more than prose.

The upstream version is a constraint set rather than a request. Tell the model to open on a specific claim with no greeting and no preview of the video, to vary sentence length deliberately because the script will be spoken, to ban the section summaries, and to keep the vocabulary to words people say out loud. Models follow instructions like those because each one is checkable while generating. They ignore "make it engaging", which is a judgment about the finished thing.