VideoGenReview

What each AI video model's release notes say changed

Sound in AI video models: dated release notes

23 entries on sound, voice or lip-sync from 7 of the 17 lines, dated 23 December 2024 to 19 August 2026. The table lists each with its line and version; the chart counts them per line. As of 2026-09-25.

1Reading the notes

Sound entries come from seven lines. Lip-sync is the older thread, starting with Kling's in December 2024 and growing to 60 seconds and several faces on screen in 2025. The earliest note here of sound generated with the picture is Veo 3's preview in July 2025, then the Wan 2.5 preview in September, then a sound switch on Kling 2.6 and joint audio and video output on Vidu in late 2025. Voice choice, voice binding and voice cloning follow from December 2025.

Entries on sound per lineCounted from the table. Lines without an entry of this kind are left out.Entries on sound per lineKling AI77PixVerse44Wan44LTX33Grok Imagine22Vidu22Google Veo11
Fig. 1 Counted from the table. Lines without an entry of this kind are left out.
Entries on sound, voice or lip-sync per line, read 2026-09-25.
LineEntries
Google Veo1
Grok Imagine2
Kling AI7
LTX3
PixVerse4
Vidu2
Wan4

Inclusion rule. Lines with at least one entry of this kind. Order. Alphabetical by line.

2Every entry

Every entry on sound, voice or lip-sync, oldest first.
DateLineVersionWhat changed
23 December 2024Kling AINot named in the noteLip-sync for videos made with the 1.0 and 1.5 models
30 June 2025Kling AINot named in the noteLip-sync videos can run 60 seconds instead of 10
14 July 2025PixVerseNot named in the noteLip-sync and extend arrive
17 July 2025Google VeoVeo 3Veo 3 preview arrives in the Gemini API and generates sound with the picture
28 July 2025ViduNot named in the noteLip-sync from text or audio
1 August 2025Kling AINot named in the noteVideo-to-audio adds sound to any Kling video, and text-to-audio arrives
5 August 2025PixVerseNot named in the noteSound effects and Fusion (reference to video) arrive
11 September 2025PixVerseNot named in the noteSound effects and text-to-speech lip-sync inside generation calls
15 September 2025Kling AINot named in the noteLip-sync handles several people on screen and a set start time
23 September 2025WanWan 2.5Wan 2.5 preview generates synchronised audio and 10 second clips
13 November 2025ViduNot named in the noteAudio and video in one output for reference-to-video and image-to-video
3 December 2025WanWan 2.6Wan 2.6 image-to-video handles dialogue between several speakers
9 December 2025LTXNot named in the noteAudio-to-video drives a clip from an audio track
15 December 2025Kling AIKling 2.6Kling 2.6 launches with a switch for generating sound
16 December 2025Kling AIKling 2.6Kling 2.6 can speak with a chosen voice
16 December 2025WanWan 2.6Wan 2.6 reference-to-video keeps a person's look and voice, with several characters at once
23 March 2026Kling AINot named in the noteElements made from several images can be tied to a voice
3 April 2026WanWan 2.7Wan 2.7 reference-to-video mixes up to five image or video references and clones a voice timbre
22 July 2026PixVerseNot named in the noteVoice cloning for image avatar and lip-sync
31 July 2026Grok ImagineGrok Imagine 1.5Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080p
13 August 2026Grok ImagineGrok Imagine 1.5Grok Imagine 1.5 reaches Runway's MCP, 1 to 15 seconds with audio
18 August 2026LTXLTX-2.5Audio-to-video gets the same resolution tiers as the other modes
19 August 2026LTXNot named in the noteAudio-to-video takes a frame rate, a last frame and camera motion

Inclusion rule. All entries tagged with this kind of change, including platform arrivals and undated rows. Order. By date, oldest first; undated rows last.

3One line at a time

4Other kinds of change