VideoGenReview

What each AI video model's release notes say changed

Reference-to-video in AI video release notes

Generating a video from one or more reference images or clips that fix how a person, object or place looks, with a prompt for what happens. As of 2026-09-25.

It differs from image-to-video, where the image becomes the first frame. In reference-to-video the references say who or what appears, and the model decides the framing. Vidu's reference video generation in August 2025 took one to seven images; Wan 2.6's reference-to-video in December 2025 kept both look and voice.

Later notes widen what counts as a reference: HappyHorse 1.0 takes up to nine images, PixVerse V6 takes video references, and Grok Imagine 1.5's reference-to-video adds optional preset voices.

Where the term appearsOne dot per dated entry.Where the term appears2025-08-26Vidu, Not named in the note2025-12-16Wan, Wan 2.62026-04-26HappyHorse, HappyHorse 1.02026-07-26PixVerse, PixVerse V62026-07-31Grok Imagine, Grok Imagine 1.52026-09-14Vidu, Vidu Q3
Fig. 1 One dot per dated entry.
Entries that show the term, oldest first, read 2026-09-25.
DateLineVersionWhat changed
26 August 2025ViduNot named in the noteReference video generation from one to seven images
16 December 2025WanWan 2.6Wan 2.6 reference-to-video keeps a person's look and voice, with several characters at once
26 April 2026HappyHorseHappyHorse 1.0HappyHorse 1.0 reference-to-video takes up to nine reference images
26 July 2026PixVersePixVerse V6V6 reference-to-video takes video references in an omni mode
31 July 2026Grok ImagineGrok Imagine 1.5Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080p
14 September 2026ViduVidu Q3Vidu Q3 reference-to-video models for drama and for ads are listed on Model Studio

Inclusion rule. Entries whose note uses or shows the term. Order. By date, oldest first; undated rows last.

1In the notes

Added reference video generation - Supports uploading 1–7 images

Vidu, API update notice

supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances.

Alibaba Cloud Model Studio, newly released models

Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance.

Alibaba Cloud Model Studio, newly released models

V6 Reference-to-Video now supports `reference_mode: "omni"` and `video_references`.

PixVerse, API changelogs

grok-imagine-video-1.5 now supports text-to-video, image-to-video, and reference-to-video (including optional preset voices), with native 1080p for T2V and I2V.

xAI, API release notes

A Vidu reference-to-video model designed for drama and AI comic series production.

Alibaba Cloud Model Studio, newly released models

2Other terms

3Notes read