VideoGenReview

What each AI video model's release notes say changed

Which AI video models let you set or clone a voice?

Kling 2.6 took a chosen voice in December 2025 and let elements carry a voice from March 2026. Wan 2.6 reference-to-video kept a person's voice in December 2025 and Wan 2.7 clones a timbre; PixVerse clones voices for avatar and lip-sync, and Grok Imagine 1.5 offers preset voices. Vidu's lip-sync takes text or audio. As of 2026-09-25.

The entries by dateOne dot per dated entry.The entries by date2025-07-28Vidu, Not named in the note2025-12-03Wan, Wan 2.62025-12-16Kling AI, Kling 2.62025-12-16Wan, Wan 2.62026-03-23Kling AI, Not named in the note2026-04-03Wan, Wan 2.72026-07-22PixVerse, Not named in the note2026-07-31Grok Imagine, Grok Imagine 1.5
Fig. 1 One dot per dated entry.
Entries that answer this question, oldest first, read 2026-09-25.
DateLineVersionWhat changed
28 July 2025ViduNot named in the noteLip-sync from text or audio
3 December 2025WanWan 2.6Wan 2.6 image-to-video handles dialogue between several speakers
16 December 2025Kling AIKling 2.6Kling 2.6 can speak with a chosen voice
16 December 2025WanWan 2.6Wan 2.6 reference-to-video keeps a person's look and voice, with several characters at once
23 March 2026Kling AINot named in the noteElements made from several images can be tied to a voice
3 April 2026WanWan 2.7Wan 2.7 reference-to-video mixes up to five image or video references and clones a voice timbre
22 July 2026PixVerseNot named in the noteVoice cloning for image avatar and lip-sync
31 July 2026Grok ImagineGrok Imagine 1.5Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080p

Inclusion rule. Entries whose note bears directly on the question. Order. By date, oldest first; undated rows last.

Once a model speaks, the next question is whose voice it uses. The notes give three answers: pick one from a set, as with Kling's voice list and Grok's preset voices; carry it with a reference, as with Wan 2.6's reference-to-video; or clone it from audio, as with Wan 2.7 and PixVerse in 2026.

Voice notes cluster from December 2025. Wan 2.6's image-to-video note in the same month adds dialogue between several speakers, which needs a way to keep the voices apart.

1In the makers' words

Vidu, 28 July 2025

Input text/audio to precisely match lip movements in video

Vidu, API update notice

Wan, 3 December 2025

supports stable multi‑speaker dialogue with more natural and realistic vocal timbres

Alibaba Cloud Model Studio, newly released models

Kling AI, 16 December 2025

By using prompt and voice_list, a specified voice can be used to generate videos

Kling AI, API updates
  • 12/16/2025
    Kling 2.6 can speak with a chosen voiceKling AI, API updates / since 2025-12-16 / checked 2026-09-25

Wan, 16 December 2025

supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances.

Alibaba Cloud Model Studio, newly released models

Kling AI, 23 March 2026

Multi-image elements also support binding timbres

Kling AI, API updates
  • 03/23/2026
    Elements made from several images can be tied to a voiceKling AI, API updates / since 2026-03-23 / checked 2026-09-25

Wan, 3 April 2026

Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning.

Alibaba Cloud Model Studio, newly released models

PixVerse, 22 July 2026

Voice Cloning Is Now Available for Image Avatar and Lip Sync Generation!

PixVerse, API changelogs

Grok Imagine, 31 July 2026

grok-imagine-video-1.5 now supports text-to-video, image-to-video, and reference-to-video (including optional preset voices), with native 1080p for T2V and I2V.

xAI, API release notes
  • July 31
    Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080pxAI, API release notes / since 2026-07-31 / checked 2026-09-25

Sound, voice and lip-sync · When did AI video models add vertical video? · Which AI video models copy a performance from a video?

2Notes read