Put new audio on a video or image
Upload a talking clip or a portrait, add new audio, a recording or text-to-speech, and the mouth is re-synced to the new words.
- Dubbing and localization
- Fixing or replacing a line
- Making a character sing or speak
New dialogue, dubs and restyles, with mouths that match
Give a clip new lines, dub it into another language, or restyle a talking video without losing sync. DomoAI analyzes the audio and reshapes the mouth to match every sound. It works on real people, anime characters and mascots.
The right tool depends on whether you’re adding new audio or keeping the original speech.
Upload a talking clip or a portrait, add new audio, a recording or text-to-speech, and the mouth is re-synced to the new words.
Turning a talking video into anime or cartoon? Switch on Lip Sync in Video to Video so the styled mouth still matches the original audio.
Send an image or a video plus audio to POST /v1/video/talking-avatar. Clips run 1–60 s, with aspect ratio and callback options.
| Visual input | Front-facing image or video with a visible face; works on real, anime and stylized characters |
|---|---|
| Audio input | Upload MP3, WAV or M4A up to 80 MB, record in the browser, or use text-to-speech (6 emotions, 6 tones) |
| Languages | Upload audio in any language |
| Length | Up to 60 s (Pro plan); API 1–60 s |
| Restyle lip sync | A toggle in Video to Video that keeps the original speech in sync |
| Output | Up to 1080p; 4K with the Video Upscaler |
| Speed | About 60 s for a 5 s clip; 60 s clips can take 10–15 min at peak times |
| API aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4 |
A steady shot with one face, front-on or 3/4, the mouth visible and well lit.
Record, use text-to-speech, or bring a voice file. Voice only, no music, silence trimmed.
Keep the new audio close to the clip’s length. Use the timing checker below.
Add the video or image plus the audio in Talking Avatar and generate.
The lips should close fully on these sounds. Review a short test first.
Add music and subtitles in your editor, and upscale if needed.
New audio that’s much longer or shorter than the clip looks rushed or frozen. Check the fit before you generate.
Re-voice your own clips in other languages for new audiences.
Change a price, a date or a call to action without reshooting.
Restyle a vlog or interview into anime while the speech stays in sync.

Make a character perform your track, using a clean vocal stem.
| Problem | Likely cause | Fix |
|---|---|---|
| Sync starts late | Silence at the start of the audio | Trim leading silence before upload |
| Lips drift or blur | Music or noise under the voice | Use an isolated, voice-only track |
| Mouth doesn’t close on p, b, m | Fast or mumbled speech | Re-record a little slower and clearer |
| Face changes over a long clip | Identity drift in long generations | Split into shorter segments |
| Odd teeth or jaw | Mouth covered or extreme angle | Pick a front-on shot with the mouth visible |
| Styled mouth off after restyle | Lip Sync turned off in Video to Video | Turn on Lip Sync and regenerate |
Lip sync changes what a person appears to say. Only re-voice yourself, characters you own, or people who have clearly agreed. Never use it to impersonate anyone, spread misinformation or put words in a real person’s mouth. DomoAI’s policies prohibit this, and accounts can be suspended.
Yes. Upload a talking video or a portrait with new audio, and DomoAI re-shapes the mouth to match the new speech. The API’s talking-avatar endpoint accepts an image or a video plus audio, for 1–60 seconds.
In Video to Video (Restyle), turn on the Lip Sync setting. It keeps the styled mouth in sync with the original speech or singing. It works best with a visible mouth, a front or 3/4 angle and good lighting.
MP3, WAV and M4A up to 80 MB. You can also record in the browser or use text-to-speech with 6 emotions and 6 voice tones. Audio can be in any language.
Yes. Create the translated voice track (recorded, text-to-speech or from a voice tool you have rights to), keep it close to the original clip length, then lip sync the clip to the new audio.
Up to 60 seconds (60-second clips on the Pro plan; the API accepts 1–60 s). Split longer videos into segments, which also reduces identity drift.
About 60 seconds for a 5-second clip. A 60-second clip can take 10–15 minutes at peak times.
Yes. DomoAI’s lip sync handles both realistic faces and stylized characters, such as anime, mascots and illustrations, as long as the mouth is clearly visible.
Only with their clear permission. Never use lip sync to impersonate someone or to make a real person appear to say something they didn’t. Label AI-dubbed content where platforms require it.
Re-voice, dub or restyle your clips with lip sync. Start with free credits.