Generated video has improved to the point where individual frames are frequently indistinguishable from filmed footage. The clips are still recognisable almost immediately. The reasons are worth naming precisely, because they determine what the technology is currently good for.
**Temporal consistency is the hard part, and it is a different problem from image quality.** A model that produces a beautiful frame has to produce the next twenty-nine consistent with it, and the one after that. Faces drift. Hands change count between shots. A logo on a shirt reads differently at second three than at second one. The eye is extraordinarily sensitive to this — far more than to a static image being slightly wrong — because tracking objects through time is the thing vision evolved to do.
**Physics is approximated, not simulated.** Cloth settles wrong. Liquid behaves like a slightly thick gas. Weight is the most common tell: things move as if they mass less than they should, which is why generated walking so often reads as floating. Some models are noticeably better at this than others, and the improvement is real, but none of them are computing forces — they are predicting what footage of forces tends to look like, which fails in exactly the cases you have not seen much footage of.
**Clip length is a hard constraint disguised as a soft one.** Most models produce a handful of seconds. You can chain clips, and the chain is where the seams show: the lighting shifts, the character is subtly different, the camera loses its logic. This is why generated video appears overwhelmingly in advertising and title sequences, where a cut every two seconds is a style rather than an admission.
Partner
Decktopus AI
DesignDecktopus is an AI presentation maker, that will create amazing presentations in seconds. You only need to type the presentation title and your presentation is…
We earn a commission if you sign up — at no extra cost to you.
Partner
Decktopus AI
DesignDecktopus is an AI presentation maker, that will create amazing presentations in seconds. You only need to type the presentation title and your presentation is…
We earn a commission if you sign up — at no extra cost to you.
**The two that are getting worse, not better:** first, the aesthetic convergence. Models trained on overlapping data produce overlapping looks — the same shallow depth of field, the same slow push-in, the same amber-and-teal grade — and as more output is generated, more of it is recognisable as belonging to the same family. Second, audience calibration. Two years ago people could not spot it; now a great many can, and the threshold keeps dropping. The technology has to improve faster than the audience learns, and it is not obvious that it is.
**What this makes it good for.** Short shots where a cut is expected. B-roll that supports rather than carries. Stylised material where the uncanny quality is a choice — animation, abstraction, dream sequences. Previsualisation, where the point is to communicate an intention to a crew who will then film it properly. Anything where the alternative is stock footage that does not quite match.
**What it is not good for yet.** Anything with a face held on screen long enough to study. Anything where physical plausibility carries the meaning — a product demonstration, a safety video. Anything documentary, for reasons that are ethical before they are technical.
The gap will close. It has closed faster than most predictions. But it will close unevenly, and the last thing to arrive will be the thing that matters most: a shot you can hold for thirty seconds on a person's face without the audience noticing anything at all.