Alle Beiträge

How to spot a deepfake video

4 Min. Lesezeit

First, work out which kind you are looking at

'Deepfake' covers two quite different things, and conflating them makes you worse at spotting either. A face swap is a real recording of a real person with someone else's face composited onto it. The body, the room and the camera motion are all genuine; only a region of the frame is synthetic. A fully generated clip — from Sora, Veo, Runway, Pika and similar — has no source footage at all. Everything in it is synthesised. Face swaps fail at the boundary between real and fake. Generated clips fail at physics and continuity. Check for the wrong one and you will miss it.

Face swaps: watch the boundary

A swap has to blend a synthetic face into a real head, and that seam is the weak point. Play the clip at quarter speed and watch the perimeter of the face. The jawline and hairline are where it shows: a faint shimmer, a slight mismatch in skin tone between face and neck, hair that flickers as it crosses the boundary. Also watch what happens during fast head turns and when a hand or object passes in front of the face — occlusion is one of the hardest cases and often produces a visible tear or a brief flicker.

Face swaps: lip-sync and micro-expression

Synthesised mouths tend to be approximately right rather than exactly right. Watch consonants that require full lip closure — b, p, m. If the lips do not fully meet on those sounds, or meet a frame or two late, that is a strong signal. Micro-expressions are the other tell. Real faces do small involuntary things constantly: asymmetric brow movement, a flicker at the corner of the mouth, blinks at irregular intervals. Swapped faces are often smoother and more symmetric than a real face ever is, and blink patterns can be unnaturally regular or oddly rare.

Generated clips: physics and continuity

A fully generated clip has no real scene behind it, so it has no obligation to be consistent over time. That is where it breaks. Watch background objects across the whole clip: a parked car that changes shape, a pattern on a wall that drifts, a person in the background whose clothing shifts colour. Watch physics: fabric that moves without inertia, liquid that does not behave like liquid, feet that slide slightly against the ground during a walk. Watch hands interacting with objects — grips that pass through surfaces or reform between frames. Length matters. Under five seconds, generated clips are very hard to fault. Longer clips accumulate continuity errors.

Check the audio separately

Voice cloning and video synthesis are separate systems, and the audio is often the weaker one. Listen with the video hidden. Synthetic speech tends to have flat or oddly-placed emphasis, breathing that does not match the sentence rhythm, and a room tone that stays constant even when the speaker moves relative to the camera. If the acoustics do not match the visible space, that is worth as much as anything in the picture.

Where the clip came from matters more than any of this

Provenance beats forensics. Before analysing frames, ask: who posted this, when, and is it anywhere else? A clip of a public figure saying something newsworthy would, if genuine, exist in more than one place — a broadcast, a press pool, another attendee's recording. A single account posting a single copy with no corroboration is the strongest signal available, and it requires no technical skill to check.

  1. 1Search for the claim in text — a real event has reporting attached
  2. 2Check the posting account's age and history
  3. 3Reverse-search a distinctive frame to find earlier copies
  4. 4Look for a second angle; consequential events usually have one

What a detector adds

Video detectors analyse things you cannot see: per-frame compression inconsistency, temporal artefacts across the sequence, blending residue at face boundaries. They can flag a specific stretch of frames as anomalous, which is genuinely useful — it narrows where to look. They are also the least reliable category of AI detection. Social platforms re-encode everything they touch, and re-encoding destroys exactly the frame-level signals detectors depend on. A clip that has been through TikTok twice is much harder to assess than the original file. So read a video score as a pointer, not an answer: 'these frames are worth your attention', not 'this is fake'.

The realistic standard

You are not going to reach certainty on a forwarded clip, and you should stop trying. The achievable goal is calibrated doubt: enough signal to decide whether to act on something, share it, or set it aside. In practice that usually means provenance first, physics and continuity second, detector third. If the clip has no traceable origin and the background does not hold together, you have what you need — you do not need the detector to agree with you.

Weiterlesen

Quellen

videodeepfakesguide