VideoGenReview

What each AI video model's release notes say changed

Audio-to-video in AI video release notes

Audio-to-video generates a clip driven by an audio track, so speech or music shapes what happens on screen. As of 2026-09-25.

LTX opened an audio-to-video endpoint in December 2025, with its own prompt enhancement. In August 2026 it brought audio-to-video level with its other modes: the same resolution tiers, and settings for frame rate, a last frame and camera motion.

It is the reverse of generating sound with the picture. Here the sound comes first and the video follows it.

For dialogue scenes and music videos, that order can be the useful one: the performance is fixed in the audio track, and the picture is made to fit it, instead of hoping a generated voice lands on the right beat. LTX's August 2026 notes also added a last frame to audio-to-video, so the clip can be steered toward a set ending while the sound drives what happens before it.

Other notes treat audio as one input among several rather than a mode of its own. MiniMax describes H3 as reading text, image, video and audio together, and Runway's note on the model it calls Hailuo 3.0 lists reference audio next to reference images and videos.

Where the term appearsOne dot per dated entry.Where the term appears2025-12-09LTX, Not named in the note2026-07-31MiniMax Hailuo, MiniMax H32026-08-05MiniMax Hailuo, Hailuo 3.02026-08-18LTX, LTX-2.52026-08-19LTX, Not named in the note
Fig. 1 One dot per dated entry.
Entries that show the term, oldest first, read 2026-09-25.
DateLineVersionWhat changed
9 December 2025LTXNot named in the noteAudio-to-video drives a clip from an audio track
31 July 2026MiniMax HailuoMiniMax H3MiniMax H3 reads text, image, video and audio together as creative context
5 August 2026MiniMax HailuoHailuo 3.0Runway's developer API adds a model it names MiniMax Hailuo 3.0
18 August 2026LTXLTX-2.5Audio-to-video gets the same resolution tiers as the other modes
19 August 2026LTXNot named in the noteAudio-to-video takes a frame rate, a last frame and camera motion

Inclusion rule. Entries whose note uses or shows the term. Order. By date, oldest first; undated rows last.

1In the notes

New audio-to-video endpoint for generating videos driven by audio input, with dedicated prompt enhancement.

LTX, API changelog

It understands creative intent across multimodal context — text, image, video, and audio — and delivers more natural, coherent generation and expression.

MiniMax, model release notes

MiniMax Hailuo 3.0 is now live on Runway Dev.

Runway, changelog

now supports the same resolution tiers as text-to-video and image-to-video

LTX, API changelog

Audio-to-video now supports fps (default 24), last_frame_uri (requires image_uri), and camera_motion

LTX, API changelog
  • Audio-to-Video API
    Audio-to-video drives a clip from an audio trackLTX, API changelog / since 2025-12-09 / checked 2026-09-25
  • MiniMax H3
    MiniMax H3 reads text, image, video and audio together as creative contextMiniMax, model release notes / since 2026-07-31 / checked 2026-09-25
  • MiniMax Hailuo 3.0
    Runway's developer API adds a model it names MiniMax Hailuo 3.0Runway, changelog / since 2026-08-05 / checked 2026-09-25
  • Audio-to-video resolution parity
    Audio-to-video gets the same resolution tiers as the other modesLTX, API changelog / since 2026-08-18 / checked 2026-09-25
  • Audio-to-video generation params
    Audio-to-video takes a frame rate, a last frame and camera motionLTX, API changelog / since 2026-08-19 / checked 2026-09-25

2Other terms

3Notes read