After continuously investing in large language model technology, the AI company MiniMax has recently officially launched its new multimodal generation model MiniMax H3 and pledged to fully open-source the model weights within a few days.

As a generation model that supports multimodal context input, MiniMax H3 can receive any combination of text, images, videos, and audio as input, and finally output high-definition videos with native stereo audio. In practical applications, users only need to provide reference videos, images, or audio materials to the model, combined with natural language instructions, to seamlessly complete complex audio-visual content creation.

image.png

On the technical architecture side, MiniMax H3 has completely broken away from the traditional approach of splitting expert models by task in multimodal models. The development team integrated various tasks such as text-to-image, text-to-video, image-to-video, and audio processing into a single pre-training framework, introducing a more concise H3-Omni Transformer architecture, which increased end-to-end training throughput by nearly 30%. At the same time, through the H3-VAE architecture brought about by rewriting the Tokenizer, it achieved extremely high compression rates, effectively reducing inference costs at high resolutions.

Regarding pricing and ecosystem, MiniMax H3 costs less than one-third of mainstream similar models per second at 2K resolution, and during the design phase, full consideration was given to compatibility with domestic chips. Currently, the model shows broad application potential in scenarios such as movie openers, game UI animations, dynamic posters, and e-commerce ads. Its accuracy in text rendering and brand information presentation has also been significantly improved.