Skip to content

MiniMax H3 brings 33B audio-video generation to Hugging Face

Original: MiniMax H3 lands on Hugging Face as a 33B audio-video model View original →

Read in other languages: 한국어日本語
AI Aug 4, 2026 By Insights AI (Twitter) 1 min read 1 views Source
MiniMax H3 brings 33B audio-video generation to Hugging Face

A wider open path for audio-video generation

MiniMax H3 became a fresh Hugging Face signal on August 3, when the multimodalart account posted that the model is available for text-to-video, image-to-video, and reference-to-video workflows with audio. The tweet also pointed developers to model weights and an app on Spaces.

MiniMax H3 just dropped on Hugging Face — text-to-video, image-to-video, reference-to-video — all with audio.

The Hugging Face model card describes MiniMax H3 as a general-purpose omni-modal generative system. It supports multimodal contexts made of text, images, video, and audio, and can generate video at 24 FPS for 4 to 15 seconds. The card lists 33B parameters in the repository metadata, default output with a 768-pixel shorter side, 2K generation through H3-Regenerate-2K, and native 32 kHz stereo audio.

multimodalart is a Hugging Face staff-linked account that regularly surfaces Spaces, Diffusers demos, and multimodal releases. That context matters here because the tweet is not only a model claim; it points to usable distribution paths. The linked collection includes the MiniMaxAI/MiniMax-H3 model, a reference-oriented Space, and a Qwen3-VL conditioner service.

The release is worth watching because audio-video generation has often been split across separate video and sound models. A single Hugging Face workflow for text, image, video, and audio references could change the build-versus-buy decision for short-form video tools. The open questions are licensing, hardware requirements, and whether the 2K and reference modes remain reliable under real creator workloads. Source tweet

Share: Long

Related Articles