VideoGenReview

What each AI video model's release notes say changed

How many reference images do AI video models take?

The notes that give a number: Veo 3.1 up to three images, Vidu one to seven, Wan 2.7 up to five images or videos mixed, HappyHorse 1.0 up to nine for generation and five for its edit model. Runway, Sora, Kling, PixVerse and Grok describe references without a count in the notes read. As of 2026-09-25.

The entries by dateOne dot per dated entry.The entries by date2025-04-30Runway, Gen-42025-08-26Vidu, Not named in the note2025-10-15Google Veo, Veo 3.12025-12-16Wan, Wan 2.62026-02-25Kling AI, Kling 3.02026-03-12Sora, Sora 22026-04-03Wan, Wan 2.72026-04-26HappyHorse, HappyHorse 1.02026-04-26HappyHorse, HappyHorse 1.02026-07-26PixVerse, PixVerse V62026-07-31Grok Imagine, Grok Imagine 1.5
Fig. 1 One dot per dated entry.
Entries that answer this question, oldest first, read 2026-09-25.
DateLineVersionWhat changed
30 April 2025RunwayGen-4Gen-4 References keeps characters and locations consistent
26 August 2025ViduNot named in the noteReference video generation from one to seven images
15 October 2025Google VeoVeo 3.1Veo 3.1 preview adds extension, up to three reference images and first-and-last-frame input
16 December 2025WanWan 2.6Wan 2.6 reference-to-video keeps a person's look and voice, with several characters at once
25 February 2026Kling AIKling 3.0Kling 3.0 and 3.0 Omni reach the API; elements can be built from video
12 March 2026SoraSora 2Character references, clips up to 20 seconds, 1080p on Sora 2 Pro, extensions and batch jobs
3 April 2026WanWan 2.7Wan 2.7 reference-to-video mixes up to five image or video references and clones a voice timbre
26 April 2026HappyHorseHappyHorse 1.0HappyHorse 1.0 reference-to-video takes up to nine reference images
26 April 2026HappyHorseHappyHorse 1.0HappyHorse 1.0 video edit changes parts of a clip from instructions and up to five images
26 July 2026PixVersePixVerse V6V6 reference-to-video takes video references in an omni mode
31 July 2026Grok ImagineGrok Imagine 1.5Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080p
Undated version rowSeedanceSeedance 2.0Seedance 2.0: 4 to 15 seconds, up to 4K, multimodal references, editing and extension

Inclusion rule. Entries whose note bears directly on the question. Order. By date, oldest first; undated rows last.

A reference image tells the model what a person, object or place looks like so it stays the same across shots. The notes describe it under several names: references at Runway and Veo, elements at Kling, character references at Sora, Fusion at PixVerse. The idea is the same; the limits differ, and most notes do not state them.

Counts rose quickly where they are given, from three in Veo 3.1's October 2025 note to nine in HappyHorse 1.0's in April 2026. Two 2026 notes go further than pictures: Wan 2.7 mixes images with video references, and PixVerse V6 takes video references in an omni mode. Several notes tie a voice to the reference too.

1In the makers' words

Runway, 30 April 2025

Generate consistent characters, locations and more.

Runway, changelog
  • Gen-4 References
    Gen-4 References keeps characters and locations consistentRunway, changelog / since 2025-04-30 / checked 2026-09-25

Vidu, 26 August 2025

Added reference video generation - Supports uploading 1–7 images

Vidu, API update notice
  • August 26
    Reference video generation from one to seven images2025Vidu, API update notice / since 2025-08-26 / checked 2026-09-25

Google Veo, 15 October 2025

Released Veo 3.1 and 3.1 Fast models in public preview, with new features including: Extending Veo-created videos. Referencing up to three images to generate a video. Providing first and last frame images to generate videos from.

Google, Gemini API changelog
  • October 15
    Veo 3.1 preview adds extension, up to three reference images and first-and-last-frame input2025Google, Gemini API changelog / since 2025-10-15 / checked 2026-09-25

Wan, 16 December 2025

supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances.

Alibaba Cloud Model Studio, newly released models

Kling AI, 25 February 2026

Supports the creation of elements through video

Kling AI, API updates
  • 02/25/2026
    Kling 3.0 and 3.0 Omni reach the API; elements can be built from videoKling AI, API updates / since 2026-02-25 / checked 2026-09-25

Sora, 12 March 2026

Expanded the Sora API with reusable character references, longer generations up to 20 seconds, 1080p output for sora-2-pro , video extensions, and Batch API support for POST /v1/videos .

OpenAI, API changelog
  • March 12
    Character references, clips up to 20 seconds, 1080p on Sora 2 Pro, extensions and batch jobs2026OpenAI, API changelog / since 2026-03-12 / checked 2026-09-25

Wan, 3 April 2026

Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning.

Alibaba Cloud Model Studio, newly released models

HappyHorse, 26 April 2026

Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance.

Alibaba Cloud Model Studio, newly released models

HappyHorse, 26 April 2026

It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness.

Alibaba Cloud Model Studio, newly released models

PixVerse, 26 July 2026

V6 Reference-to-Video now supports `reference_mode: "omni"` and `video_references`.

PixVerse, API changelogs
  • 2026/07/26
    V6 reference-to-video takes video references in an omni modePixVerse, API changelogs / since 2026-07-26 / checked 2026-09-25

Grok Imagine, 31 July 2026

grok-imagine-video-1.5 now supports text-to-video, image-to-video, and reference-to-video (including optional preset voices), with native 1080p for T2V and I2V.

xAI, API release notes
  • July 31
    Grok Imagine video 1.5 adds reference-to-video with preset voices and native 1080pxAI, API release notes / since 2026-07-31 / checked 2026-09-25

Seedance, Undated version row

4K (10-bit color depth)

BytePlus ModelArk, model list
  • dreamina-seedance-2-0-260128
    Seedance 2.0: 4 to 15 seconds, up to 4K, multimodal references, editing and extensionBytePlus ModelArk, model list / recorded 2026-09-25

References and consistency · Which AI video models take a first and a last frame? · Which AI video models can extend a clip?

2Notes read