API
On this section

Guide

Choosing and using a lipsync model

Fabric 1.0 for an image that talks, Lipsync 2.0 for new audio on existing footage — what each takes, and how to get good results.

VEED has two lipsync models, and the one you want depends on what you start from. A still image: Fabric 1.0 makes it talk. Existing footage: Lipsync 2.0 re-syncs it to new audio.

Which one

Fabric 1.0Lipsync 2.0
You sendAn image and an audio trackA video and an audio track
You getA talking videoThe same video, re-synced to the new audio
Use it forAvatars, mascots, personalised videosDubbing, translation, replacing a voice track
Price$0.08–$0.15 per second of the generated video$0.07 per second of the generated video

Both are asynchronous: submit, then poll until the job settles. The quickstart walks that loop.

Fabric 1.0: an image that talks

audio_urlurlrequired

URL of the audio track to lip-sync to.

Must be an http(s) URL.
image_urlurlrequired

URL of the source image to animate.

Must be an http(s) URL.
resolutionenumrequired

Output video resolution.

Accepts720p480p
terminal
# Submit a job
curl -X POST "https://api.veed.io/v1/fabric-1.0" \
  -H "Authorization: Bearer $VEED_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "audio_url": "https://static-assets.veed.io/api-examples/fabric-input.mp3",
  "image_url": "https://static-assets.veed.io/api-examples/fabric-input.png",
  "resolution": "720p"
}'

# Poll until COMPLETED (response includes data.job_id)
curl "https://api.veed.io/v1/fabric-1.0/{job_id}" \
  -H "Authorization: Bearer $VEED_API_KEY"

resolution is one of 720p, 480p, and the price follows it. The model takes speech, not text: to make the image say a script, generate the audio with a text-to-speech tool and send its URL.

Lipsync 2.0: new audio for existing footage

audio_urlurlrequired

URL of the new audio track to lip-sync the video to.

Must be an http(s) URL.
video_urlurlrequired

URL of the source video to dub.

Must be an http(s) URL.
terminal
# Submit a job
curl -X POST "https://api.veed.io/v1/lipsync-2.0" \
  -H "Authorization: Bearer $VEED_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "audio_url": "https://static-assets.veed.io/api-examples/lipsyncv2-input.mp3",
  "video_url": "https://static-assets.veed.io/api-examples/lipsyncv2-input.mp4"
}'

# Poll until COMPLETED (response includes data.job_id)
curl "https://api.veed.io/v1/lipsync-2.0/{job_id}" \
  -H "Authorization: Bearer $VEED_API_KEY"

To translate a video, send the translated speech with the original footage. One speaker per video: more than one is not supported.

Getting good results

These hold for both models.

  • Send clean speech. One voice, with no music or background noise. If the recording is noisy, run it through Clean Audio first.
  • Show the face clearly. Front-facing, well lit, and not covered by hands, a microphone or hair.
  • Use a sharp, high-resolution source.
  • For Lipsync 2.0, avoid fast movement. A speaker who turns away or moves quickly gives worse results.
  • Try a short clip first, before you send a long job.

If a job ends FAILED, its error.code says why, and each poll endpoint lists what every code means:

GET
/v1/fabric-1.0/{job_id}
GET
/v1/lipsync-2.0/{job_id}