Skip to main content

One portrait, one voice track
and your character speaks

Bring any speech track — a recording, a dub, a voice you generated — and the person in your photo delivers it. Lip-synced, in your photo's own framing.

Natural speech, your framing, ready to cut

Speech that reads as performance

Speech that reads as performance

Lip-sync alone looks like a puppet. The stronger mode drives head turns and hand movement from the same audio, so the take holds up in a cut instead of sitting there.

Your frame, not a template

Your frame, not a template

The output inherits the proportions of the photo you bring. A vertical portrait gives you a vertical video, a wide one stays wide — no letterboxing, no forced crop.

Any voice, any language

Any voice, any language

The audio drives everything — pacing, pauses, the shape of every word. It can be recorded, dubbed or generated, in any language, and the face follows it.

Key Features

Two render modes, generated characters, voice tracks made in place

Turbo for drafts, Quality for the take

Turbo takes tracks up to 60 seconds and returns a short one in about a minute — enough to check that the timing and the read work. Quality caps at 20 seconds and takes considerably longer, but adds the head and hand movement that makes a take usable. Pick per shot, not per project.

Characters, not only people

Nothing here is tied to a real face. A character you generated, a painted portrait, a face that never existed — if it reads as a portrait, it speaks. The people you invent get a voice on the same terms as the people you photograph.

Voiceovers made in the same place

No track yet is not a blocker. Write the line in Voice Studio, generate the read, and bring it straight here — same workspace, same subscription, nothing exported and re-uploaded between two services.

How it works

Open Talking Avatar

Give your character a voice

One portrait and one track is the whole setup.

More tools to speed up your creative workflow.

Frequently asked questions