Talking Avatar is Workroom's AI app for talking-head video. You bring a portrait photo and a speech track, and the person in the photo delivers it. The result is lip-synced, at 720p, in the proportions of the photo you started from. Nothing is tied to a real face: a generated character or an illustrated portrait speaks on the same terms as a photograph.
Before you begin
You need two files.
A portrait photo. One person facing the camera with the face clearly visible, as JPG, PNG or WEBP. The short side has to be at least 720 pixels. A normal phone photo is well past that; only small images, like a frame pulled from low-resolution video, get rejected.
Open Talking Avatar
Open Apps from the sidebar and pick Talking Avatar. The form is three fields long — a photo, a track, and a mode — with the requirements written next to each one.

Add a portrait, a voiceover and a mode
Under Step 1: Portrait photo, drop in your image or click Select to pick one from your assets. Under Step 2: Voiceover, add the speech track the same way.
Step 3: Mode is optional and decides how the result moves:
- Turbo takes tracks up to 60 seconds and returns a short one in about a minute. Use it to check that the timing and the read work.
- Quality caps at 20 seconds and takes considerably longer, but drives head turns and hand movement from the audio instead of the mouth alone.
Leave the mode alone and the run goes through on Turbo.

Generate and take the video
Click Generate. The finished video plays in the workspace with its audio, and you can download it from the icon in the corner or rate the result with the thumbs.

Your photo and track stay in the form, so you can run the same pair again on the other mode and compare the two takes side by side.
Troubleshooting
The photo was rejected as too small
The short side is under 720 pixels. The error names the dimensions it found. Use a larger version of the same image rather than upscaling it by hand.
The voice track was refused
It runs longer than the selected mode allows — 60 seconds for Turbo, 20 for Quality. Trim the track or switch modes.
The head and hands don't move
Head and hand movement only comes from Quality, and only for what the photo shows. Check that the mode is set to Quality, and that the hands are inside the frame of the portrait you uploaded.
The lip-sync looks loose
Front-facing, evenly lit portraits sync best. A heavily turned head, a hand across the mouth, or a face in deep shadow all cost accuracy. A clean speech track with little background noise helps as much as the photo does.