Today, AI company MiniMax officially launched the general-purpose multimodal generation model MiniMax H3. This model achieves unified understanding and generation of text, images, videos, and audio for the first time, breaking the long-standing fragmentation between tasks and modalities in the video generation field, marking a crucial step towards multimodal generation models moving from "specialized experts" to "general assistants."
According to the official introduction, H3 supports native dual-channel audio-video output, with a maximum generation capacity of 15 seconds of 2K resolution content. During early invitation tests, H3 demonstrated commercial-level stability in scenarios such as instruction following, brand information presentation, and V2V action transfer, and is already applicable to multiple commercial fields such as advertising, e-commerce, gaming, and UI design.

In terms of technology, H3 introduces the Contextual Omni Representation (Contextual Multimodal Representation) technology, making natural language a universal bridge connecting various modalities. Users need only describe complex relationships through language—for example, "refer to the Hitchcock-style camera movement from a certain video and have a person in a specified image sing the melody of an audio clip"—the model can automatically complete the entire chain of understanding and generation. Combined with its self-developed H3-VAE and In-context Regeneration technologies, H3's per-second generation cost at 2K resolution is less than one-third of mainstream models, and less than half at 768P resolution, providing the industry's currently known most cost-effective solution.
Notably, MiniMax announced that it will open-source the H3 model weights in the coming days, under the premise of complying with relevant laws and regulations. This is also the first time a large-scale domestic video generation model has opened its weights to the community, aiming to promote the development of open-source ecosystems and accelerate the adaptation of domestic chips. From the initial design stage, H3 considered compatibility with multiple domestic chips, and its open-sourcing will further reduce the technical barriers for domestic companies in multimodal AI applications.
MiniMax stated that H3 still has room for improvement in terms of image fineness and model scale. In the future, it will promote the integration of capabilities between the H series and M series models, and continue to expand the model's upper limits through scaling.
Industry analysts believe that the release and open-sourcing of H3 may change the long-term dominance of closed-source models in the video generation field, accelerating the popularization of multimodal AI applications.


