Skip to main content

Create a talking avatar video

Updated September 11, 2026Talking Avatar — coming soon

Talking Avatar is Workroom's AI app for talking-head video. You bring a portrait photo and a speech track, and the person in the photo delivers it. The result is lip-synced, at 720p, in the proportions of the photo you started from. Nothing is tied to a real face: a generated character or an illustrated portrait speaks on the same terms as a photograph.

Talking Avatar will be available in your Workroom account soon. Everything below describes how the app works — some of it may change by the time it ships.

Before you begin

You need two files.

A portrait photo. One person facing the camera with the face clearly visible, as JPG, PNG or WEBP. The short side has to be at least 720 pixels. A normal phone photo is well past that; only small images, like a frame pulled from low-resolution video, get rejected.

A speech track. Any recording of someone talking, as MP3 or WAV. It can be a voice memo, a dub, or a voice you generated — see how to generate a voiceover if you don't have one yet. The language doesn't matter.
The output keeps the proportions of your photo. Upload a vertical portrait and you get a vertical video, ready for social formats without cropping.

Open Talking Avatar

Open Apps from the sidebar and pick Talking Avatar. The form is three fields long — a photo, a track, and a mode — with the requirements written next to each one.

Talking Avatar workspace with the three-step form on the right
Talking Avatar workspace with the three-step form on the right

Add a portrait, a voiceover and a mode

Under Step 1: Portrait photo, drop in your image or click Select to pick one from your assets. Under Step 2: Voiceover, add the speech track the same way.

Step 3: Mode is optional and decides how the result moves:

  • Turbo takes tracks up to 60 seconds and returns a short one in about a minute. Use it to check that the timing and the read work.
  • Quality caps at 20 seconds and takes considerably longer, but drives head turns and hand movement from the audio instead of the mouth alone.

Leave the mode alone and the run goes through on Turbo.

Step 2 filled in with the mode list open on Turbo and Quality
Step 2 filled in with the mode list open on Turbo and Quality
Quality can only move what it can see. If you want hand movement in the result, upload a photo where the hands are in the frame — a tight head-and-shoulders crop gives it nothing to work with.
A track longer than the mode allows is refused before anything is generated. Trim it, or switch to Turbo for the full length.

Generate and take the video

Click Generate. The finished video plays in the workspace with its audio, and you can download it from the icon in the corner or rate the result with the thumbs.

Generated talking-avatar video playing in the workspace
Generated talking-avatar video playing in the workspace

Your photo and track stay in the form, so you can run the same pair again on the other mode and compare the two takes side by side.

Apps share a monthly pool of 5 free generations on the Free plan — once it's used, each run costs credits, with the exact cost shown on the Generate button.
Output is 720p. If you need it larger — for a big screen or a client delivery — run it through Topaz Video Upscaler afterwards.

Troubleshooting

The photo was rejected as too small

The short side is under 720 pixels. The error names the dimensions it found. Use a larger version of the same image rather than upscaling it by hand.

The voice track was refused

It runs longer than the selected mode allows — 60 seconds for Turbo, 20 for Quality. Trim the track or switch modes.

The head and hands don't move

Head and hand movement only comes from Quality, and only for what the photo shows. Check that the mode is set to Quality, and that the hands are inside the frame of the portrait you uploaded.

The lip-sync looks loose

Front-facing, evenly lit portraits sync best. A heavily turned head, a hand across the mouth, or a face in deep shadow all cost accuracy. A clean speech track with little background noise helps as much as the photo does.

I want a talking head from a video, not a photo

That's a different tool. Workroom's lip-sync takes an existing video and re-syncs the face in it to new audio, while Talking Avatar builds the video from one still image.