Which AI video models generate sound with the picture?
In the notes read: Veo 3 from its preview in July 2025, the Wan 2.5 preview in September 2025 and later Wan versions, Vidu's joint audio and video output from November 2025, Kling 2.6's sound switch in December 2025 and Grok Imagine 1.5 with preset voices in July 2026. Kling's earlier video-to-audio adds sound afterwards. As of 2026-09-25.
| Date | Line | Version | What changed |
|---|---|---|---|
| 17 July 2025 | Google Veo | Veo 3 | Veo 3 preview arrives in the Gemini API and generates sound with the picture |
| 1 August 2025 | Kling AI | Not named in the note | Video-to-audio adds sound to any Kling video, and text-to-audio arrives |
| 23 September 2025 | Wan | Wan 2.5 | Wan 2.5 preview generates synchronised audio and 10 second clips |
| 13 November 2025 | Vidu | Not named in the note | Audio and video in one output for reference-to-video and image-to-video |
| 3 December 2025 | Wan | Wan 2.6 | Wan 2.6 image-to-video handles dialogue between several speakers |
| 15 December 2025 | Kling AI | Kling 2.6 | Kling 2.6 launches with a switch for generating sound |
| 31 July 2026 | Grok Imagine | Grok Imagine 1.5 | Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080p |
| 11 August 2026 | Seedance | Seedance 2.5 | Seedance 2.5 reaches Runway through its MCP, up to 30 seconds |
| 13 August 2026 | Grok Imagine | Grok Imagine 1.5 | Grok Imagine 1.5 reaches Runway's MCP, 1 to 15 seconds with audio |
| 24 August 2026 | Wan | Wan 3.0 | Wan 3.0 arrives in Runway at 480p, 720p or 1080p |
Inclusion rule. Entries whose note bears directly on the question. Order. By date, oldest first; undated rows last.
The difference that matters is whether sound is generated in the same pass as the picture or laid onto a finished clip. Kling's August 2025 note is the second kind: video-to-audio works on any Kling video, including one made earlier. Its December 2025 note on Kling 2.6 is the first kind, a switch on the generation itself.
Platform posts repeat the claim for the models they host: Runway describes Seedance 2.5, Grok Imagine 1.5 and Wan 3.0 as generating with native audio or with audio. Those are the platforms' words about someone else's model, so they sit here as arrivals rather than as the maker's own note.
Sound entries that only take audio in, such as LTX's audio-to-video, are not on this page; they are on the audio kind-of-change page.
1In the makers' words
Google Veo, 17 July 2025
Launched veo-3.0-generate-preview , the latest update to Veo introducing video with audio generation.
Google, Gemini API changelog
- July 17Veo 3 preview arrives in the Gemini API and generates sound with the picture2025
Kling AI, 1 August 2025
Supports adding audio to all videos generated by Kling models
Kling AI, API updates
- 08/01/2025Video-to-audio adds sound to any Kling video, and text-to-audio arrives
Wan, 23 September 2025
newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation
Alibaba Cloud Model Studio, newly released models
- wan2.5-t2v-previewWan 2.5 preview generates synchronised audio and 10 second clips
Vidu, 13 November 2025
Audio&video direct output is now available for Vidu Reference to Video and Image to Video.
Vidu, API update notice
- November 13Audio and video in one output for reference-to-video and image-to-video2025
Wan, 3 December 2025
supports stable multi‑speaker dialogue with more natural and realistic vocal timbres
Alibaba Cloud Model Studio, newly released models
- wan2.6-i2v-usWan 2.6 image-to-video handles dialogue between several speakers
Kling AI, 15 December 2025
Control whether to include audio when generating videos through the sound parameter.
Kling AI, API updates
- 12/15/2025Kling 2.6 launches with a switch for generating sound
Grok Imagine, 31 July 2026
grok-imagine-video-1.5 now supports text-to-video, image-to-video, and reference-to-video (including optional preset voices), with native 1080p for T2V and I2V.
xAI, API release notes
- July 31Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080p
Seedance, 11 August 2026
Generate up to 30 seconds of video with native audio
Runway, changelog
- Seedance 2.5Seedance 2.5 reaches Runway through its MCP, up to 30 seconds
Grok Imagine, 13 August 2026
Generate 1-15 second clips with native audio, up to 1080p resolution.
Runway, changelog
- Grok Imagine 1.5 VideoGrok Imagine 1.5 reaches Runway's MCP, 1 to 15 seconds with audio
Wan, 24 August 2026
Wan 3.0 is now available in Runway in tool mode and workflows.
Runway, changelog
- Wan 3.0Wan 3.0 arrives in Runway at 480p, 720p or 1080p
Sound, voice and lip-sync · When did AI video clips reach 30 seconds? · When did AI video models get 4K output?
2Notes read
- Google, Gemini API changelog, read 2026-09-25
- Kling AI, API updates, read 2026-09-25
- Alibaba Cloud Model Studio, newly released models, read 2026-09-25
- Vidu, API update notice, read 2026-09-25
- xAI, API release notes, read 2026-09-25
- Runway, changelog, read 2026-09-25