Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Spain attacks ‘selfish’ response of some EU countries to Ceuta migrant crossings

    August 1, 2026

    The 5 laptop features worth spending extra on (and 3 that are mostly hype)

    August 1, 2026

    Ruby on Rails Patches Critical Vulnerability

    August 1, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Spain attacks ‘selfish’ response of some EU countries to Ceuta migrant crossings
    • The 5 laptop features worth spending extra on (and 3 that are mostly hype)
    • Ruby on Rails Patches Critical Vulnerability
    • Russia Expands Crypto Mining Ban to Moscow
    • Corey Ruiz shooting: what happened, what’s next and the reporting to know
    • India’s Modi says he forgives students who abused him in Cockroach protests | Narendra Modi News
    • Ceuta crisis: 22 EU leaders gang up against Spain’s Sánchez over his migration policy
    • On the ground in the English seaside town rightwing influencers say is being ‘invaded’ | Immigration and asylum
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Saturday, August 1
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 1, 2026 Artificial Intelligence No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    MiniMax releases MiniMax H3, a general-purpose multimodal generation model. MiniMax H3 is not a text-to-video model with add-ons. MiniMax describes it as a general-purpose multimodal generation model that reads text, images, video, and audio as one unified context and returns video with native stereo sound. The mains specs include: 2K output, 4–15 seconds, integer durations only.

    Previous video stacks split into text-to-video, image-to-video, first-and-last-frame, subject reference, motion reference, and video editing, each often a separate expert model. MiniMax H3 folds those into one pretraining paradigm where reference and editing relationships are expressed in natural language. MiniMax’s example prompt makes the point: reference the camera movement from Video 1, have the character in Image 2 sing, match the vocals to Audio 3.

    Is it deployable?

    Today: yes, through the API and no, on your own hardware. MiniMax launched H3 on July 31, 2026 with the model live in the platform API under the model ID MiniMax-H3 and in the consumer Hailuo AI app.

    Industries: MiniMax positions MiniMax H3 for advertising, branding, e-commerce, product design, UI/UX, and gaming along with film pre-visualization and retail catalog media.

    Applications: Ad variant generation, product and listing videos, animated posters, film title sequences, website hero loops, character-consistent game cinematics, and video-to-video motion transfer.

    The API surface

    The video generation guide documents three entry modes: text-to-video, first/last-frame image-to-video, and reference generation. Behind one endpoint and an asynchronous three-step flow: create a task, poll task_id, download content.url.

    Input limits worth designing around:

    • Reference images: up to 9. Reference videos: up to 3 clips, 2–15 s each, ≤15 s total. Reference audio: up to 3 clips, and audio cannot be sent without an accompanying image or video.
    • Mixed input caps at 12 files total. Prompt length ≤7,000 characters; request body ≤64 MB, with URL input recommended for large assets.
    • File sizes: video ≤50 MB, image ≤30 MB, audio ≤15 MB, per asset.
    • Formats: H.264/H.265 video, JPG/PNG/WEBP/HEIC/HEIF images, WAV/MP3 audio.

    Four technical pieces doing the work

    Contextual Omni Representation: MiniMax rebuilt captioning so it describes the relationship between context and target video, not just the target. Most source material requires roughly 100K tokens of inference, distilled to about 4K tokens on average. Language is the bridge that turns a fixed task set into an open, descriptive one.

    H3-VAE: A full tokenizer overhaul. Its high compression ratio delivers a stated 4× gain in effective sequence length, cutting training and inference cost and it is the enabling technology for native 2K.

    H3-Omni Transformer: MiniMax explicitly set aside the Hailuo-02 architecture here. Multimodal context tripled sequence-length variance, so the training architecture separates understanding and generation workloads and tunes hardware utilization for each. Reported result: end-to-end training throughput up nearly 30%.

    In-Context Regeneration: Instead of a bolt-on super-resolution module, the base model regenerates its own low-resolution output in-context, re-reading the original multimodal context. That is what recovers small text and fine detail that conventional upscalers guess at — directly relevant to brand and product rendering.

    Price and standing

    MiniMax’s own claim: at 2K, H3’s per-second price is less than a third of mainstream models; at 768p, less than half the price of mainstream 720p. The company amplified both the launch and the pricing framing on X (1, 2). Third-party trackers and launch coverage put the 2K pay-as-you-go rate at $0.13 per second, about $1.95 for a 15-second clip, but MiniMax’s pay-as-you-go page still listed only Hailuo 2.3 tiers at the time of writing, so treat that figure as reported, not primary.

    On placement: SCMP reports, citing Artificial Analysis, that H3 leads in video editing while trailing Google’s Gemini Omni Flash in text-to-video and sitting behind both Seedance 2.0 and Gemini Omni Flash in image-to-video.

    Key Takeaways

    • H3 unifies text, image, video, and audio into one generation model — 2K, 4–15s, native stereo.
    • Open weights are promised “in the coming days,” not shipped; the API is the only path today.
    • H3-VAE’s 4× effective sequence-length gain is what makes native 2K economically viable.
    • In-context regeneration replaces super-resolution, preserving small text and brand marks.
    • Artificial Analysis ranks H3 first in video editing, behind rivals in text-to-video and image-to-video.

    Sentimental Analysis


    Check out the Technical details. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

    15Second Audio Clips Generates MiniMax model native OmniModal Releases Stereo video
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

    Three Recent Chrome Releases Fix 1,442 Flaws, More Than Prior 23 Updates Combined

    Does video show thousands of migrants entering Spanish territory of Cueta? What we know

    DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains

    LingBot-Map Tutorial: GPU-Aware Inference and Point Cloud Export

    OpenAI aligns safety practices with EU AI Act’s GPAI Code

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Spain attacks ‘selfish’ response of some EU countries to Ceuta migrant crossings

    August 1, 2026

    The 5 laptop features worth spending extra on (and 3 that are mostly hype)

    August 1, 2026

    Ruby on Rails Patches Critical Vulnerability

    August 1, 2026

    Russia Expands Crypto Mining Ban to Moscow

    August 1, 2026
    Latest Posts

    New to Linux? This 10-day checklist will help you settle in nice and easy

    July 22, 2026

    Tories ask HMRC to investigate whether Nigel Farage owes tax on £5m gift | Nigel Farage

    July 22, 2026

    Greece derails EU’s Russia sanctions plan – POLITICO

    July 22, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Spain attacks ‘selfish’ response of some EU countries to Ceuta migrant crossings

    August 1, 2026

    The 5 laptop features worth spending extra on (and 3 that are mostly hype)

    August 1, 2026

    Ruby on Rails Patches Critical Vulnerability

    August 1, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.