Sound in AI video models: dated release notes
23 entries on sound, voice or lip-sync from 7 of the 17 lines, dated 23 December 2024 to 19 August 2026. The table lists each with its line and version; the chart counts them per line. As of 2026-09-25.
1Reading the notes
Sound entries come from seven lines. Lip-sync is the older thread, starting with Kling's in December 2024 and growing to 60 seconds and several faces on screen in 2025. The earliest note here of sound generated with the picture is Veo 3's preview in July 2025, then the Wan 2.5 preview in September, then a sound switch on Kling 2.6 and joint audio and video output on Vidu in late 2025. Voice choice, voice binding and voice cloning follow from December 2025.
| Line | Entries |
|---|---|
| Google Veo | 1 |
| Grok Imagine | 2 |
| Kling AI | 7 |
| LTX | 3 |
| PixVerse | 4 |
| Vidu | 2 |
| Wan | 4 |
Inclusion rule. Lines with at least one entry of this kind. Order. Alphabetical by line.
2Every entry
| Date | Line | Version | What changed |
|---|---|---|---|
| 23 December 2024 | Kling AI | Not named in the note | Lip-sync for videos made with the 1.0 and 1.5 models |
| 30 June 2025 | Kling AI | Not named in the note | Lip-sync videos can run 60 seconds instead of 10 |
| 14 July 2025 | PixVerse | Not named in the note | Lip-sync and extend arrive |
| 17 July 2025 | Google Veo | Veo 3 | Veo 3 preview arrives in the Gemini API and generates sound with the picture |
| 28 July 2025 | Vidu | Not named in the note | Lip-sync from text or audio |
| 1 August 2025 | Kling AI | Not named in the note | Video-to-audio adds sound to any Kling video, and text-to-audio arrives |
| 5 August 2025 | PixVerse | Not named in the note | Sound effects and Fusion (reference to video) arrive |
| 11 September 2025 | PixVerse | Not named in the note | Sound effects and text-to-speech lip-sync inside generation calls |
| 15 September 2025 | Kling AI | Not named in the note | Lip-sync handles several people on screen and a set start time |
| 23 September 2025 | Wan | Wan 2.5 | Wan 2.5 preview generates synchronised audio and 10 second clips |
| 13 November 2025 | Vidu | Not named in the note | Audio and video in one output for reference-to-video and image-to-video |
| 3 December 2025 | Wan | Wan 2.6 | Wan 2.6 image-to-video handles dialogue between several speakers |
| 9 December 2025 | LTX | Not named in the note | Audio-to-video drives a clip from an audio track |
| 15 December 2025 | Kling AI | Kling 2.6 | Kling 2.6 launches with a switch for generating sound |
| 16 December 2025 | Kling AI | Kling 2.6 | Kling 2.6 can speak with a chosen voice |
| 16 December 2025 | Wan | Wan 2.6 | Wan 2.6 reference-to-video keeps a person's look and voice, with several characters at once |
| 23 March 2026 | Kling AI | Not named in the note | Elements made from several images can be tied to a voice |
| 3 April 2026 | Wan | Wan 2.7 | Wan 2.7 reference-to-video mixes up to five image or video references and clones a voice timbre |
| 22 July 2026 | PixVerse | Not named in the note | Voice cloning for image avatar and lip-sync |
| 31 July 2026 | Grok Imagine | Grok Imagine 1.5 | Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080p |
| 13 August 2026 | Grok Imagine | Grok Imagine 1.5 | Grok Imagine 1.5 reaches Runway's MCP, 1 to 15 seconds with audio |
| 18 August 2026 | LTX | LTX-2.5 | Audio-to-video gets the same resolution tiers as the other modes |
| 19 August 2026 | LTX | Not named in the note | Audio-to-video takes a frame rate, a last frame and camera motion |
Inclusion rule. All entries tagged with this kind of change, including platform arrivals and undated rows. Order. By date, oldest first; undated rows last.