Production sample · Tidenza

AI UGC talking head, built to survive a close look

Eleven seconds, one continuous take, no cuts and no captions. Captions would cover the mouth, and the mouth is the part worth judging.

Everything in the frame is generated. The room, the person and the voice do not exist.

How it was made

Start image
Soul 2, single generation, no inpainting
Motion
Wan 2.7, image to video, 1080p
Lip sync
Driven by the voice track, not by the model's own audio
Voice
ElevenLabs, written for uneven delivery with a hesitation
Finishing
Phone-mic treatment, room tone, loudness pass, no beauty retouch
Build time
Under two hours, start to export

Where these usually fail. Lettering on clothing comes out garbled, so the shirt is plain. Skin goes waxy, so the prompt asks for pores, redness and a blemish and refuses any beauty filter. The arm is cropped at the frame edge, because that is the geometry you get when someone holds their own phone.

In motion, the giveaway is a head that talks in a smooth arc while the room sits frozen behind it. So the camera keeps a live micro-shake, the delivery runs uneven, and there is a blink and a glance away where a real person would put one.

When the mouth corners go rubbery, the fix is a better starting frame. Not a filter downstream.

The audio is half the tell. A voice recorded in a bathroom on a phone is band-limited, bounces off tile and sits on a floor of room noise. Studio-clean speech over a handheld shot reads as fake even when the lip sync is perfect, so this track is cut to phone bandwidth, given short tile reflections and a quiet room floor underneath.