One portrait, one voice track
and your character speaks
Bring any speech track — a recording, a dub, a voice you generated — and the person in your photo delivers it. Lip-synced, in your photo's own framing.
Natural speech, your framing, ready to cut
Speech that reads as performance
Lip-sync alone looks like a puppet. The stronger mode drives head turns and hand movement from the same audio, so the take holds up in a cut instead of sitting there.
Your frame, not a template
The output inherits the proportions of the photo you bring. A vertical portrait gives you a vertical video, a wide one stays wide — no letterboxing, no forced crop.
Any voice, any language
The audio drives everything — pacing, pauses, the shape of every word. It can be recorded, dubbed or generated, in any language, and the face follows it.
Key Features
Two render modes, generated characters, voice tracks made in place
Turbo for drafts, Quality for the take
Turbo takes tracks up to 60 seconds and returns a short one in about a minute — enough to check that the timing and the read work. Quality caps at 20 seconds and takes considerably longer, but adds the head and hand movement that makes a take usable. Pick per shot, not per project.
Characters, not only people
Nothing here is tied to a real face. A character you generated, a painted portrait, a face that never existed — if it reads as a portrait, it speaks. The people you invent get a voice on the same terms as the people you photograph.
Voiceovers made in the same place
No track yet is not a blocker. Write the line in Voice Studio, generate the read, and bring it straight here — same workspace, same subscription, nothing exported and re-uploaded between two services.
How it works
Try other apps
More tools to speed up your creative workflow.