MiniMax, the Shanghai-based AI lab behind the Hailuo video models, has launched H3 — a general-purpose multimodal generation model it says breaks down the walls between media types.

H3 understands unified context across text, images, video and audio, and can generate videos up to 15 seconds long at 2K resolution with native stereo sound, meaning the audio is generated as part of the same model rather than stitched on afterwards. MiniMax highlights strength in instruction following, text and brand presentation in video, and video-to-video motion transfer.

Pricing is the other headline: at 2K, the model costs less than one-third of mainstream per-second prices, and at 768p it is less than half the price of mainstream 720p offerings. The company says it plans to open-source the weights in the coming days to support the open-source community and accelerate hardware compatibility — a step in line with MiniMax's strategy of releasing open-weight models.

For creators and developers, H3's combination of open weights, native audio and aggressive pricing is a direct challenge to the paid-video-model status quo. The model is already available through the company's API and third-party platforms, and observers will be watching whether the open-source release follows as promised — and how quickly rivals answer.