If a singing photo looks out of sync, fix the inputs before regenerating: use a front-facing image with an uncovered mouth, trim silence before the first vocal, choose a short section with clear lead vocals, and avoid heavy reverb or competing voices. FreeLipSync does not provide a manual mouth-timing track after generation, so clean source timing is the reliable control.
What you need
- Source image: The fictional singer faces the camera and the microphone is deliberately placed to the side. This removes two common image problems—profile angle and an object crossing the lips—while keeping the performance context.
- Input audio used in this example: The 18-second original song begins clearly and has one lead vocal. If your own clip has a long instrumental intro, crowd vocals, overlapping harmonies, or a beat that masks consonants, trim to a simpler section for the diagnostic pass.
- Tool: Make Photo Sing
Diagnose one variable at a time. First trim the audio, then test the image, then try a different vocal section. Changing everything together makes it impossible to know which correction worked.
Source image and input
Input audio used in this example
Exact English transcript or lyrics used by the shared demo media:
This is my moment, watch me rise, Spotlight shining in my eyes.
I found my voice, I'm breaking free, This is the start of the real me — Here I go, this is my time!
Raw result
The raw result is a reference for a clean input pair, not a promise that every song and portrait will behave identically. Watch the beginning first: a consistent offset usually points to leading silence; unstable movement often points to a difficult image or crowded vocal mix.
Open the dedicated video watch page
Step-by-step workflow
- Prepare one clear source image — Use one front-facing image with a visible mouth, even lighting, and no hand, microphone, or heavy shadow covering the lips.
- Open the matching FreeLipSync tool — Open the linked tool for this workflow. Confirm that its input mode matches your source: text, uploaded speech, or song audio.
- Add the script or audio input — Add a short clean input. Keep the first test simple so it is easy to tell whether the image and audio are a good match.
- Generate the lip-sync result — Generate one raw result without changing several variables at once. A controlled first pass makes troubleshooting much faster.
- Review the full result before publishing — Watch with sound from beginning to end. Check mouth visibility, timing, pronunciation, and whether the clip represents the source honestly.
Why this setup works
Lip-sync generation relies on visible facial landmarks and audible vocal timing. Cleaning those two signals is more effective than repeatedly submitting the same problematic pair or trying to hide a mismatch with a later crop.
Quality checklist
- Remove silence before the first sung syllable and export the corrected clip before regenerating.
- Choose one clear lead vocal; avoid duet overlaps, chorus stacks, and extreme echo for the first pass.
- Use a front-facing portrait with a large, uncovered mouth and no microphone directly across the lips.
Common mistakes to avoid
- Diagnose one variable at a time. First trim the audio, then test the image, then try a different vocal section. Changing everything together makes it impossible to know which correction worked.
- Do not resubmit the same problematic image and audio repeatedly. Correct one input, export it, and run a controlled comparison.
- Do not publish a face, character, recording, or song unless you have the necessary permission or license.
Questions people ask
Why does the mouth start moving late?
Leading silence or a soft, unclear vocal entrance is a common cause. Trim the file so the first intended sound begins cleanly, then regenerate.
Can I manually move the lip-sync track after generation?
Not in the current Make Photo Sing workflow. Correct the audio timing or source image and generate a new result.
Why does the mouth move weakly even when timing is correct?
A small, side-facing, shadowed, or obstructed mouth can reduce visible movement. Test a clearer front-facing crop with the same audio.
Make your own version
Trim one clean 10–20 second vocal section, pair it with an unobstructed portrait, and regenerate in Make Photo Sing.


