To dub a video with AI lip sync, prepare a translated script, upload one clear presenter video, then either select a voice for the translated text or upload the finished dub track. FreeLipSync synchronizes the new speech to the visible speaker; it does not translate the source inside this workflow. Translation accuracy and review remain your responsibility.
What you need
- Video: an authorized source with one clear, visible speaker.
- Translation: a reviewed target-language script or a finished dub recording.
- Tool: AI Video Dubbing.
The presenter in this tutorial is fictional and was generated for the demonstration. Obtain permission to dub real people, and never use a cloned voice without the speaker's authorization.
Exact English source video
Original English line:
Hello and welcome. In this short update, I will explain how our new workflow helps global teams share clear product messages.
Reviewed Spanish translation and audio
Hola y bienvenidos. En esta breve actualización, explicaré cómo nuestro nuevo flujo de trabajo ayuda a los equipos globales a compartir mensajes de producto claros.
Exact Spanish audio used for lip sync
We selected an existing Spanish female voice from the catalog. The tool can also use another language-matched preset, an authorized cloned voice, an uploaded finished track, or a browser recording.
Unedited Spanish result
The displayed result is the direct output. It has not been manually retimed or given replacement mouth animation in post-production.
Open the dedicated video watch page
Step-by-step workflow
- Prepare the source — Use a steady clip with one speaker, a clear mouth, and minimal cuts.
- Translate before generation — Preserve meaning, names, numbers, tone, and approximate spoken length.
- Choose Text or Audio — Text uses a preset or authorized cloned voice; Audio uses a finished upload or browser recording.
- Generate a short proof — Test the most difficult sentence before processing a long piece.
- Review the final dub — Check pronunciation, timing, meaning, identity permissions, and disclosure.
Text dubbing versus audio dubbing
Use Text when the translation is still being revised or you need to choose among existing voices for the target language. Use Audio when a human actor or another production tool has already delivered the final performance. A finished track gives you direct control of pacing, emotion, and pronunciation.
Free accepts up to 133 text code points or 20 seconds of audio. Starter Generate Pro accepts up to 800 code points or 3 minutes. Pro supports up to 16,000 code points or 60 minutes. These are input ceilings, not duration billing units.
Translation and timing tips
- Translate for spoken meaning, not word-for-word length.
- Read the target script aloud and compare its duration with the source.
- Check product names, abbreviations, dates, and numbers with a fluent reviewer.
- Add punctuation where the target speaker should breathe.
- If one sentence is much longer in translation, rewrite it concisely or prepare a paced audio performance.
Source-video limits
One continuously visible speaker is the strongest case. Rapid cuts, profile turns, hands over the face, large camera motion, multiple speakers, or long periods without a visible face reduce reliability. The workflow changes visible speech timing; it does not translate on-screen text, create subtitles, remove background music, or mix a finished soundtrack.
Pricing is per high-resolution result, not duration credits
Free generation deducts no Pro Video, and downloads are watermark-free subject to the free sign-in and resolution rules. Starter includes 20 Pro Videos per 30-day entitlement window; one high-resolution result uses one Pro Video rather than a credit charge for every second or minute. Active Pro subscription entitlement is unlimited.
The 2026 AI video pricing report shows an effective Starter comparison of about $0.25 per output minute when all 20 included results average one minute. That number is planning math, not a rate charged against the timeline.
Troubleshooting
- The dub runs longer than the shot: shorten the translation or record a slightly faster natural performance.
- Names sound wrong: use phonetic spelling in Text mode or record the finished line in Audio mode.
- The wrong face appears active: use a source with only one visible speaker.
- Music competes with speech: prepare a speech-first dub track, then mix music after lip sync.
Questions people ask
Does FreeLipSync translate English into Spanish?
Not in this workflow. The displayed Spanish translation was prepared and reviewed before generation.
Can I use a voice actor's recording?
Yes, with permission. Add the finished performance in Audio mode.
Does one result consume more because it is longer?
No. Starter Generate Pro uses one Pro Video per completed high-resolution result, not a per-second or per-minute balance.
Make your own dub
Start with a short authorized presenter clip and one reviewed target-language sentence. Once meaning, voice, and timing pass together, move to a longer segment.



