top of page

Why People Can Tell Your Content Is AI, in Text and in Video

Aug 27
6 min read
Why People Can Tell Your Content Is AI, in Text and in Video

Short answer: in both cases the tell is structural, not cosmetic. In writing it is rhythm, the uniform shape of sentences and paragraphs, rather than word choice. In video it is where the camera is standing, not the resolution of the face. Neither problem is solved by asking for higher quality, which is why so much obviously-AI content is also technically excellent.

I should disclose up front that I build two of the tools mentioned here, so treat those parts as claims rather than reviews. The findings underneath them are checkable and I have tried to include the ones that make my own products look worse as well as better.

What actually gives AI writing away?

Not vocabulary. Everyone hunts for words like "delve" and "landscape", swaps them out, and is surprised when the piece still reads as machine-written. The giveaway is rhythm.

Machine-written prose is unnaturally even. Paragraphs come out the same length. Sentences cluster around the same length. Headings are perfectly parallel to each other. Lists arrive in threes far more often than chance would produce. Every section ends by restating what it just said. None of these is wrong in isolation, and all of them together are a fingerprint.

Here is the part that surprised me. I ran a piece through an AI detector that I had written by hand, start to finish, and it came back at roughly 37% AI. Useless as a verdict. But the passage-level breakdown was genuinely useful: almost every sentence it flagged was a short punchy declarative. "The product works." "It is evidence." "A boring first second is."

Which is a problem, because "vary your sentence length, drop in short punchy ones" is the single most repeated piece of writing advice on the internet, and models have absorbed it completely. The short dramatic fragment used to read as a human writer with a pulse. It now reads as a language model performing a human writer with a pulse.

What machine prose does

What people actually do

Paragraphs of near-identical length

Paragraphs that run long and then stop dead

Short punchy fragments used for rhythm at regular intervals

Fragments used rarely, and usually in anger

Perfectly parallel headings

Headings that change grammatical shape halfway down

A summary sentence closing every section

Sections that stop as soon as the point is made

Balanced hedging on every claim

One strong opinion, then an unhedged aside

Examples that could apply to any company

A specific number, a date, and a name

The last row is the one that matters most and the one no rewriting tool can fix for you. Specificity is the thing models cannot fake, because they do not know what happened in your Tuesday.

Does rewriting AI text actually fix it?

Partly, and it depends entirely on whether the rewriting is structural or cosmetic.

Swapping synonyms changes nothing, because the fingerprint is in the shape rather than the words. Rewriting that restructures sentences, breaks the even rhythm and varies paragraph length does change the read. That is the useful category, and it is what TextToHuman, a free AI text humanizer, is built to do: it rewrites the draft structurally rather than substituting words, and scores it as it goes, in any of 25 languages.

Two honest caveats, because I would rather say this than have you find out later.

Detector scores are probabilistic signals, not proof. My own hand-written piece scored 37%. Treat any number as a hint about which passages read oddly, never as a verdict on authorship, and never chase a score into worse writing.

No tool adds the specifics. It can fix your rhythm. It cannot add the date, the number, or the thing your customer said on the phone last week. That part is yours, and it is the part that actually convinces anyone.

What gives AI video away?

Not resolution, and not skin texture. Composition.

Real selfie video is shot with the phone about thirty to forty centimetres from the face. The head fills most of the frame width, the hair grazes or crosses the top edge, the lens sits slightly above or below eye level with a few degrees of tilt, and the background is meaningless fragments of a room. A door frame. A wall meeting a ceiling. A blown-out window edge.

Generated footage defaults to none of that. It frames shoulders to mid-torso, which is portrait distance. It puts the lens level and centred. It composes a tidy room behind the subject. The result is photoreal and obviously fake, for reasons the viewer feels a second before they could explain them.

What real selfie video does

What generated footage defaults to

Phone 30 to 40 cm away, head filling the frame

Shoulders to mid-torso, which is portrait distance

No camera arm in shot, because at that range it is out of frame

An arm reaching in from a bottom corner, which is itself a synthetic signature

Lens off axis, a few degrees of tilt

Level, centred, eye height, like a webcam on a stand

Background fragments you cannot identify

A composed interior: bookshelf, plant, monitor, all readable

Mid-moment expression, caught between words

A held, camera-aware half smile, repeated across every angle

We found this by pulling mid-clip frames from licensed footage of real people we had already bought, seven creators of them, and comparing frame against frame with our own generated output. The uncomfortable discovery was that our own prompt had been causing four of the five. It asked for shoulders-to-torso framing. It demanded a visible camera arm in every shot. We were paying a model to produce the exact signature we were trying to avoid.

The bookshelf one is worth dwelling on, because it is the most common and least obvious. A phone held close and tilted physically cannot see a composed, eye-level wall of shelves. When the background is well arranged, the brain flags the shot before it flags the person.

Can you prompt your way out of it?

Some of it. Two things fight back.

The mouth. Generate a person and the model will animate speech whether you asked for it or not. You can write "not speaking" six different ways and still get lips moving against no audio, which is fatal in a feed watched on mute where the caption is doing the talking. We stopped rewriting the instruction and started generating extra takes and discarding the ones where the mouth moves. Generate and filter beats rewrite.

Contradiction. Ask for two things that cannot both be true in one shot, a close-up and a visible camera arm, or a tilted phone and a composed background, and the model does not choose. It attempts both, and that is where extra fingers and impossible geometry come from. Most "the AI is broken" moments are a prompt asking for something that cannot physically exist.

What does respond well is the list in the table above, stated as physics rather than as style. Close range, arm out of shot, uneven light from a single source, shine on the forehead, a fragment of a room. Models follow camera facts far more reliably than they follow adjectives like "authentic".

This is the thinking behind ClipMyApp, an AI UGC video generator for apps and SaaS. And the honest footnote is that filtering alone did not get us across the line: roughly half of the opener library we ship, 94 clips of the 190 published, is licensed footage of real people rather than generated video.

What does getting it wrong actually cost?

Credibility, and faster than people expect.

The failure most worth guarding against is not stylistic. Generated copy will happily state a number your source material does not contain: a percentage, a user count, a saving. It reads perfectly well and it is a lie, and one screenshot from an unimpressed reader makes it a public one. The rule worth enforcing, whatever tooling you use, is that a claim may only assert what the underlying material actually shows. If you cannot ground it, cut the line rather than soften it.

Everything else on this page is about rhythm and framing. That one is about whether people can trust you, which is the only reason any of the rest matters.

Quick answers

Why does AI writing sound like AI? Mostly rhythm rather than vocabulary. Machine prose produces paragraphs and sentences of unnaturally even length, perfectly parallel headings, frequent groups of three, and a summarising sentence at the end of each section.

Do AI humanizers work? Structural rewriting that changes sentence shape and paragraph length does change how a piece reads. Synonym substitution does not, because the fingerprint is in the structure rather than the words.

Are AI detector scores reliable? They are probabilistic signals rather than proof of authorship. A piece written entirely by hand can score in the 30s, so use the passage-level breakdown to find odd-reading sections rather than treating the percentage as a verdict.

Why does AI generated video look fake? Composition, not image quality. Real selfie footage is shot 30 to 40 cm from the face, off axis, with fragments of a room behind it, while generated footage defaults to portrait distance, level with the eyes, against a composed background.

How far should the camera be for realistic UGC video? About thirty to forty centimetres, close enough that the head fills most of the frame width and the holding arm falls outside the shot.

Why do AI generated people have extra fingers? Usually because the prompt asked for two things that cannot both be true in a single shot, so the model attempts to satisfy both instead of choosing one.


Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page