NVDA 213.05 ▲2.19%GOOGL 346.96 ▼0.32%MSFT 491.71 ▲0.90%AMD 479.18 ▲4.91%INTC 87.48 ▲0.25%TSMC 417.41 ▲1.78%AMZN 261.06 ▼0.39%META 570.05 ▲1.97%AAPL 309.90 ▼0.14%PLTR 172.73 ▼1.80%
Markets at last close

Models

MiniMax H3 brings open-weight video generation with native audio

·1 min read

MiniMax H3 is a general-purpose multimodal video generation model from MiniMax that accepts text, images, video, and audio as creative context. It can generate clips up to 15 seconds at 2K resolution with synchronized native stereo sound, targeting use cases such as advertising, branding, ecommerce, film, product design, UI, gaming, and other commercial creative work.

The model is designed to handle tasks that are often split across separate systems, including text-to-audio-video, first-and-last-frame generation, reference-to-audio-video, multimodal editing, motion transfer, native multi-shot modeling, and joint generation of voice, sound effects, and music. H3 can combine a character image, motion from video, an audio performance, and written direction into one multimodal brief.

MiniMax released H3-Base weights under the MiniMax H3 Community License. The official model card describes a 33B dense H3-Omni Transformer, visual and audio VAEs, and separate checkpoints for first-or-last-frame and reference-to-audio-video tasks. Local H3-Base supports 768p validation, while the full 2K workflow combines local deployment with official Context-IR and Regenerate-2K APIs.

Originally reported by visionstory.aiRead the source →
Related coverage