On this section
Guide
Choosing and using a lipsync model
Fabric 1.0 for an image that talks, Lipsync 2.0 for new audio on existing footage — what each takes, and how to get good results.
VEED has two lipsync models, and the one you want depends on what you start from. A still image: Fabric 1.0 makes it talk. Existing footage: Lipsync 2.0 re-syncs it to new audio.
Which one
| Fabric 1.0 | Lipsync 2.0 | |
|---|---|---|
| You send | An image and an audio track | A video and an audio track |
| You get | A talking video | The same video, re-synced to the new audio |
| Use it for | Avatars, mascots, personalised videos | Dubbing, translation, replacing a voice track |
| Price | $0.08–$0.15 per second of the generated video | $0.07 per second of the generated video |
Both are asynchronous: submit, then poll until the job settles. The quickstart walks that loop.
Fabric 1.0: an image that talks
audio_urlurlrequiredURL of the audio track to lip-sync to.
image_urlurlrequiredURL of the source image to animate.
resolutionenumrequiredOutput video resolution.
720p480p# Submit a job
curl -X POST "https://api.veed.io/v1/fabric-1.0" \
-H "Authorization: Bearer $VEED_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio_url": "https://static-assets.veed.io/api-examples/fabric-input.mp3",
"image_url": "https://static-assets.veed.io/api-examples/fabric-input.png",
"resolution": "720p"
}'
# Poll until COMPLETED (response includes data.job_id)
curl "https://api.veed.io/v1/fabric-1.0/{job_id}" \
-H "Authorization: Bearer $VEED_API_KEY"resolution is one of 720p, 480p,
and the price follows it. The model takes speech, not text: to make the image
say a script, generate the audio with a text-to-speech tool and send its URL.
Lipsync 2.0: new audio for existing footage
audio_urlurlrequiredURL of the new audio track to lip-sync the video to.
video_urlurlrequiredURL of the source video to dub.
# Submit a job
curl -X POST "https://api.veed.io/v1/lipsync-2.0" \
-H "Authorization: Bearer $VEED_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio_url": "https://static-assets.veed.io/api-examples/lipsyncv2-input.mp3",
"video_url": "https://static-assets.veed.io/api-examples/lipsyncv2-input.mp4"
}'
# Poll until COMPLETED (response includes data.job_id)
curl "https://api.veed.io/v1/lipsync-2.0/{job_id}" \
-H "Authorization: Bearer $VEED_API_KEY"To translate a video, send the translated speech with the original footage. One speaker per video: more than one is not supported.
Getting good results
These hold for both models.
- Send clean speech. One voice, with no music or background noise. If the recording is noisy, run it through Clean Audio first.
- Show the face clearly. Front-facing, well lit, and not covered by hands, a microphone or hair.
- Use a sharp, high-resolution source.
- For Lipsync 2.0, avoid fast movement. A speaker who turns away or moves quickly gives worse results.
- Try a short clip first, before you send a long job.
If a job ends FAILED, its error.code says why, and each poll endpoint lists
what every code means:
/v1/fabric-1.0/{job_id}
/v1/lipsync-2.0/{job_id}