To dub an existing video into another language, first prepare and review the translation, then supply either a finished target-language recording or the translated text with a suitable voice to AI Video Dubbing. FreeLipSync synchronizes that new speech with the filmed speaker. It does not translate the original words inside this workflow.
The example below uses an English presenter clip, a Spanish translation prepared in advance, and a finished Spanish voice track. That track drives the unedited Spanish lip-sync result. This is an Audio-mode example; Text mode is explained separately so the two paths are not confused.
When is AI video dubbing useful?
Share a recorded explanation with another language audience
A teacher or product specialist may already have an on-camera lesson that works visually, but an audience needs it in another language. Refilming the same demonstration with a second speaker changes the performance and takes another shoot. Dubbing lets you keep the filmed person, setting, gestures, and expressions while replacing the speech with a reviewed target-language version. This is useful for training introductions, product explainers, support videos, and regional announcements.
The translation should communicate the original meaning, not merely copy its word order. Have a fluent reviewer check names, dates, numbers, terminology, and tone. If a visible action refers to a particular object or moment, make sure the translated line still makes sense at that point in the shot.
Make a new-language edition without a new camera session
A team with an approved presenter video can prepare separate voice tracks for different audiences. Each language needs its own reviewed script and voice performance; the existing source footage can remain the visual starting point. The saved effort is in the camera session, location, lighting, and on-screen performance. The result still depends on the real movement and expression already present in the source, which is why a natural, versatile performance helps multiple versions feel believable.
This tutorial shows one English-to-Spanish edition, not an automatically translated set of languages.
Revise the voice performance as well as the language
Sometimes the target-language words are ready, but their pacing and emphasis matter. A finished recording gives a translator, voice actor, or producer direct control over pronunciation, pauses, and emotion before lip sync. If you do not have a final recording, Text mode can speak a reviewed translation using a language-matched preset or an authorized cloned voice. For a same-language line change, see the text-driven video rewrite tutorial. If you already have a finished recording and do not need translation guidance, see the uploaded-audio tutorial.
The exact English-to-Spanish example
The source is a fictional presenter created for this demonstration. It is approximately 7.96 seconds long and says in English:
Hello and welcome. In this short update, I will explain how our new workflow helps global teams share clear product messages.
The Spanish translation was prepared and reviewed before lip-sync generation:
Hola y bienvenidos. En esta breve actualización, explicaré cómo nuestro nuevo flujo de trabajo ayuda a los equipos globales a compartir mensajes de producto claros.
An existing Spanish preset voice was used to render this approximately 8.98-second recording. This is the finished audio supplied to the dubbing workflow, not an audio track extracted from the English source:
Finished Spanish dub audio
Here is the approximately nine-second unedited result. Compare the words with the Spanish script, and compare the person's movement and expression with the English source. No manual mouth retiming or replacement animation was added after generation.
Watch the unedited Spanish result on its dedicated page
The Spanish speech is about a second longer than the English source, and the result follows the Spanish recording. A target-language sentence does not have to be shortened simply to equal the source duration. FreeLipSync can match suitable parts of the filmed performance to the new speech. What still deserves review is whether the original gestures and expressions fit the translated meaning, especially in a longer or more action-specific scene.
How to make your own dub with finished audio
- Choose footage you may dub. Use a video you own or are licensed to adapt, including permission to portray an identifiable person saying new words. Select a performance whose actions and expressions make sense with the translated message. It need not be long or a rigid, straight-on mouth close-up.
- Translate and review the message. Write the target-language script outside the lip-sync step. A fluent reviewer should confirm meaning, product terms, names, numbers, and tone. Check any actions or on-screen labels that the speech references.
- Prepare the final voice track. Record a permitted speaker or render the reviewed script in a language-matched preset or authorized cloned voice. Listen to the whole track before uploading; this recording determines the words, pronunciation, and delivery used in Audio mode.
- Upload and generate. Open AI Video Dubbing, upload the source video, select Audio, add the finished dub track, and generate. The English-to-Spanish case above follows this route.
- Review before sharing. Listen for the approved words and pronunciation. Watch whether the source gestures remain plausible under the new speech. Check publication permission and any on-screen text or soundtrack that needs separate editing.
If you are updating only one section of a longer film, export that section as a clip, dub it, and return it to your edit. This workflow does not automatically find a sentence inside a completed timeline, translate text that appears on screen, add subtitles, or create a final music mix.
When to use translated text instead
Choose Text when the translation is approved but you have not recorded the final target-language performance. Paste the translated script and select a suitable preset voice or an authorized cloned voice. The tool generates the new speech and synchronizes it to the video. Choose Audio when a finished take already exists and its pacing, pronunciation, and expression are part of the approved dub.
The source-video principles are the same in either mode. The footage supplies the person's physical performance; the dubbing workflow changes visible speech rather than regenerating body movement and expressions. A short source may support a longer message when its movement is broadly reusable. A longer source can also be dubbed, but actions tied to a specific spoken moment should be checked against the translated line. Current script and audio limits depend on the selected plan, so check the pricing page for limits before preparing a long segment.
Frequently asked questions
Does FreeLipSync translate the English source into Spanish automatically?
No. The Spanish text in this example was prepared and reviewed before generation. AI Video Dubbing synchronizes the target-language speech with the video; it does not decide how to translate the original message.
Does this example use Text or Audio mode?
It uses Audio. The reviewed Spanish script was rendered with a Spanish preset voice, and that finished track was uploaded. Text mode is available when you want the tool to generate speech from a reviewed translation instead.
Must the target-language recording match the source video's length?
No. Here, the filmed English source is about eight seconds and the Spanish audio and result are about nine seconds. Check the output for gestures or expressions that may look out of place as the new speech continues.
Can I use an authorized voice actor's recording?
Yes. Upload the finished target-language take in Audio mode. Obtain permission to use both the voice performance and the video, especially when an identifiable person will appear to say new words.
Last updated:



