Trusted by 1M+ creators

Lip Sync API

Swap the audio on any video and the lips match every time. No reshoot, no retraining

Lip Sync API

4.6

2,100+ reviews

Visa logo
Venture Foods logo
Merck logo
Target logo
Pentax logo
Procter & Gamble logo
Meta company logo with infinity symbol and text 'Meta' in gray.
Amazon logo with company wordmark and the brand's signature curved smile.
Google logo
BBC logo featured in VEED's customer showcases.
NBCUniversal
UBS logo with stylized snowflake symbol and text.
Netflix logo in gray lettering.
OpusClip logo
Elevenlabs logo

AI Lip Sync API: Change a hook, fix a line, or localize any video

Reshooting a video to change one hook, fix a flubbed line, or record a translation is slow, expensive, and impossible to automate. VEED's AI lip sync API replaces that whole cycle. Send it a source video and a new audio track, and it returns a fully lip-synced result that matches the original take. VEED's original lip sync model also remains available on fal.ai for existing integrations.

Built on VEED's Lip Sync 2.0 model, the API carries emotion, style, and timing across from your new audio with no per-subject training or fine-tuning. Marketing teams can use it to A/B test ad hooks, product teams can use it to patch mistakes after sign-off, and localization teams can ship every market version from a single shoot. It runs alongside the rest of VEED's AI video APIs.

How to lip sync a video with the AI lip sync 2.0 API:

Step 01

Send your video and audio

Make a request with your source video and the new audio track. The API accepts standard video and audio files, subjects up to 4K, clips up to 10 minutes, and files up to 5GB. No training data or reference footage needed.

Step 02

Let the model sync and transfer

The model detects the face and re-renders the mouth and lower face to match your new audio. Emotion and speaking style carry across from the recording, so the delivery reads as the speaker's own. Results come back in a fraction of the time a retake would cost you.

Step 03

Get your lip-synced video back

The API returns a finished video at the same resolution you sent in, with the mouth synced to the new audio. Pair it with translated voiceovers from VEED's AI video translator to ship localized versions at scale.

Learn More

Also see Fabric 1.0 – VEED’s image-to-video model:

Multiple frames of a woman talking next to a lip sync API form with fields for video and audio URL inputs

Change the audio, keep the performance

One take becomes endless versions. Swap in a new hook to A/B test a paid social ad, drop in a corrected line, or replace the voiceover with a translation, and the lips match every time. No reshoot, no re-record, no manual animation. The original visual take stays intact while the audio does the work, whether the output is going to TikTok, Reels, Shorts, or a client's YouTube channel.

A woman presenting a sneaker in a video with a cursor hovering over a green download button for the generated file

Dubbing that looks like the real take

The model transfers emotion and speaking style straight from your audio, so the result reads as the speaker's own performance instead of a generic overlay. It also holds up where most lip sync models break: a hand or object crossing the face, low light, fast camera movement, and footage cut from multiple angles. Extreme angles and close-up mouth detail are where the gap against competing models shows most, which is what makes the output brand-safe enough to publish under your name.

A man speaking next to a dropdown menu listing multiple programming languages for developer ready API integration

Built for pipelines and scale

Call it over REST or run it as a Fal deployment, and process footage up to 4K and 10 minutes at $0.07 per second, the same price at every resolution. Zero-shot means no per-face setup, so you can batch thousands of clips across different speakers, including non-human and animated subjects. Combine the lip sync 2.0 API with VEED's Fabric 1.0 API for talking-video generation or the Subtitles API to caption every localized version. VEED's original lip sync model also remains available on fal.ai for existing integrations.

FAQ

  • An AI lip sync API is a video-to-video endpoint that re-renders a speaker's mouth to match a new audio track. You send a source video and a separate audio file, and it returns a finished video with the lips synced to the new recording. VEED's lip sync 2.0 API also transfers emotion and speaking style from the audio, so the delivery matches the performance rather than looking pasted on.

  • VEED's lip sync 2.0 API costs $0.07 per second of video processed, and the price is identical at every resolution. A one-minute clip costs $4.20 whether it is 720p or 4K, so a 4K deliverable costs no more than a vertical social cut. There is no per-subject training fee and no fine-tuning step. You pay only for the seconds you process.

  • No. The API is zero-shot, so it works on a new face the first time with no training data, reference clips, or setup step. Send a video and an audio track and the model handles the sync directly. That makes it practical for automated pipelines where you process hundreds of different speakers and could never train a model for each one.

  • The API takes a video file plus a separate audio file in standard formats. It supports resolutions up to 4K, clips up to 10 minutes long, and files up to 5GB, and the per-second price does not change with resolution. It returns a video at the resolution you sent in, with the mouth synced to the new audio and the speaker's emotion and timing preserved.

  • The model delivers its highest fidelity on a forward-facing speaker, and it also handles side-angle subjects up to a medium close-up. Non-human and animated subjects are supported, so character-led and mascot content works as well as live-action talent. It currently expects one active speaker per scene. For dubbing, pair it with translated audio to build market-ready versions without reshooting.

  • Yes, and it is the most common reason teams pick it up. The API accepts audio in any language, so you can pair one original shoot with translated voiceovers and ship a version per market. It works with any audio source, including tracks generated by VEED's text-to-speech tool, which makes fully automated localization pipelines straightforward to build.

Loved by creators.

Loved by the Fortune 500

The first four videos i created with VEED got over 40,000 impressions on LinkedIn.

Travis Tyler

Travis Tyler

Senior Content Producer,
PandaDoc

I found VEED at the right time: if you're making UGC ads and not using Eye Contact Correction and Noise Reduction on VEED you're missing out on ROI.

Sebastian Schurgers

Sebastian Schurgers

Head of Growth Marketing,
Gronda

With VEED I didn't need 15 tutorials - I jumped straight in and started editing.

Gwenne Wilcox

Gwenne Wilcox

Founder,
Brand Brainery

You can go beyond choppping things. VEED actually makes my videos look great.

Michael Glover

Michael Glover

Demand Gen Manager,
ConvertFlow

Trusted by 1M+ creators

When it comes to amazing videos, all you need is VEED

Lip Sync API

More than a lip sync API

The Lip Sync 2.0 API is one endpoint in VEED's AI video creation platform. Generate talking video from a single image with Fabric 1.0, spin up presenters with AI avatars, then refine what you get back: trim it, caption it, translate it, and drop it into your brand kit so every version looks like your brand made it, not AI. From the first generation through dubbing, branding, and exporting market-ready cuts at scale, VEED turns one video into every version you need, in minutes.