Audio-to-video in AI video release notes
Audio-to-video generates a clip driven by an audio track, so speech or music shapes what happens on screen. As of 2026-09-25.
LTX opened an audio-to-video endpoint in December 2025, with its own prompt enhancement. In August 2026 it brought audio-to-video level with its other modes: the same resolution tiers, and settings for frame rate, a last frame and camera motion.
It is the reverse of generating sound with the picture. Here the sound comes first and the video follows it.
For dialogue scenes and music videos, that order can be the useful one: the performance is fixed in the audio track, and the picture is made to fit it, instead of hoping a generated voice lands on the right beat. LTX's August 2026 notes also added a last frame to audio-to-video, so the clip can be steered toward a set ending while the sound drives what happens before it.
Other notes treat audio as one input among several rather than a mode of its own. MiniMax describes H3 as reading text, image, video and audio together, and Runway's note on the model it calls Hailuo 3.0 lists reference audio next to reference images and videos.
| Date | Line | Version | What changed |
|---|---|---|---|
| 9 December 2025 | LTX | Not named in the note | Audio-to-video drives a clip from an audio track |
| 31 July 2026 | MiniMax Hailuo | MiniMax H3 | MiniMax H3 reads text, image, video and audio together as creative context |
| 5 August 2026 | MiniMax Hailuo | Hailuo 3.0 | Runway's developer API adds a model it names MiniMax Hailuo 3.0 |
| 18 August 2026 | LTX | LTX-2.5 | Audio-to-video gets the same resolution tiers as the other modes |
| 19 August 2026 | LTX | Not named in the note | Audio-to-video takes a frame rate, a last frame and camera motion |
Inclusion rule. Entries whose note uses or shows the term. Order. By date, oldest first; undated rows last.
1In the notes
New audio-to-video endpoint for generating videos driven by audio input, with dedicated prompt enhancement.
LTX, API changelog
It understands creative intent across multimodal context — text, image, video, and audio — and delivers more natural, coherent generation and expression.
MiniMax, model release notes
MiniMax Hailuo 3.0 is now live on Runway Dev.
Runway, changelog
now supports the same resolution tiers as text-to-video and image-to-video
LTX, API changelog
Audio-to-video now supports fps (default 24), last_frame_uri (requires image_uri), and camera_motion
LTX, API changelog
- Audio-to-Video APIAudio-to-video drives a clip from an audio track
- MiniMax H3MiniMax H3 reads text, image, video and audio together as creative context
- MiniMax Hailuo 3.0Runway's developer API adds a model it names MiniMax Hailuo 3.0
- Audio-to-video resolution parityAudio-to-video gets the same resolution tiers as the other modes
- Audio-to-video generation paramsAudio-to-video takes a frame rate, a last frame and camera motion
2Other terms
3Notes read
- LTX, API changelog, note of 9 December 2025, read 2026-09-25
- MiniMax, model release notes, read 2026-09-25
- Runway, changelog, read 2026-09-25
- LTX, API changelog, note of 18 August 2026, read 2026-09-25
- LTX, API changelog, note of 19 August 2026, read 2026-09-25