Tool
AI lip sync
Sync a face to any audio track. Works from a still portrait or existing footage, with Sync, Veed, Creatify, LatentSync, Infinite Talk and OVI.
Lip sync takes two inputs — a face and an audio track — and returns footage where the mouth matches the words. The face can be a still portrait, in which case the model animates it from scratch, or existing video, where it re-times the mouth against your new audio. Those are different models, and the studio splits them so you are not offered an image model for a video job.
Pairing this with the Voiceover studio is the common path: generate the narration, then drive a face with it. Because both produce artifacts in your gallery with stable storage paths, the audio from one becomes the input to the other without a download-and-re-upload round trip.
A sync job costs 15 credits and, like every generation here, runs durably server-side — the work is checkpointed, so a closed tab or a dropped connection does not lose it.
Open this studio
Lip Sync
Drive a face with an audio track
FAQ
AI lip sync — questions
Can I lip sync a still photo?
Yes. Image-driven models animate a still portrait from the audio alone. Video-driven models re-time the mouth in footage you already have.
Which languages work?
The sync models are driven by the audio waveform rather than a transcript, so they are broadly language-independent. Pair with a voice from any of the 39 language buckets in the Voiceover studio.
What does lip sync cost?
15 credits per job, scaling with the length of the audio track.
Keep reading
Related guides
Start generating today
100 free credits on signup — roughly fifty images, or twenty-five voiceovers, or three premium video clips.
No card required.
