MiniMax H3 brings 33B audio-video generation to Hugging Face
Original: MiniMax H3 lands on Hugging Face as a 33B audio-video model View original →
A wider open path for audio-video generation
MiniMax H3 became a fresh Hugging Face signal on August 3, when the multimodalart account posted that the model is available for text-to-video, image-to-video, and reference-to-video workflows with audio. The tweet also pointed developers to model weights and an app on Spaces.
MiniMax H3 just dropped on Hugging Face — text-to-video, image-to-video, reference-to-video — all with audio.
The Hugging Face model card describes MiniMax H3 as a general-purpose omni-modal generative system. It supports multimodal contexts made of text, images, video, and audio, and can generate video at 24 FPS for 4 to 15 seconds. The card lists 33B parameters in the repository metadata, default output with a 768-pixel shorter side, 2K generation through H3-Regenerate-2K, and native 32 kHz stereo audio.
multimodalart is a Hugging Face staff-linked account that regularly surfaces Spaces, Diffusers demos, and multimodal releases. That context matters here because the tweet is not only a model claim; it points to usable distribution paths. The linked collection includes the MiniMaxAI/MiniMax-H3 model, a reference-oriented Space, and a Qwen3-VL conditioner service.
The release is worth watching because audio-video generation has often been split across separate video and sound models. A single Hugging Face workflow for text, image, video, and audio references could change the build-versus-buy decision for short-form video tools. The open questions are licensing, hardware requirements, and whether the 2K and reference modes remain reliable under real creator workloads. Source tweet
Related Articles
Meta has rolled out Muse Image across Meta AI, meta.ai, Instagram Stories in the U.S., and WhatsApp in limited countries. The notable shift is an image model that uses search, code execution, self-refinement, and a watermarking system inside consumer social products.
NVIDIA Research’s MOTIVE targets a specific video-model bottleneck: which fine-tuning clips actually improve motion. The ICML 2026 honored paper reports a 74.1% human preference result against the base model.
A security incident tied to model evaluation drew unusually intense HN debate. The real issue is not only the breach, but how far cyber benchmarks can safely push models against realistic infrastructure.