The Viral AI Video Nobody Can Legally Render Locally | Edition 320
Edition 320 — MiniMax H3's viral clip really costs $19–$44, not $9.60 — plus a robot that learns a new task from one demo, no retraining.

On August 19, a clip of Seinfeld and Trump doing a bit on a New York street went viral on X. The caption: “Minimax h3 is just way too good. Rendered this locally on my machine.”
That claim is the hook, not the story. The story is that almost nobody comparing AI video tools is pricing the thing readers actually want to make.
The wrong unit
Every AI video comparison benchmarks a single clip: price per second, at whatever the shortest, cheapest setting is. Real work is a 60-120 second piece — a product demo, an explainer, a short. At that length, the bill is decided by three things a per-second number never shows:
1. Max clip length. MiniMax H3 caps at 15 seconds. A 2-minute piece is 8 separate generations, 8 prompts, 8 continuity problems to paper over.
2. Retry rate. A failed clip bills the same as a good one — there's no partial refund. Runway confirms this outright: a bad generation and its regenerate attempt are billed as two separate charges. At 4 takes per shot instead of 2 (typical for anything needing character consistency), a “cheap” model can lose to an expensive one.
3. Native audio in the same pass. Models that don't generate audio natively need a second model plus an editor timeline. That cost never shows up on a pricing page.
The unit that matters is cost per usable minute, not price per second.
The math nobody runs, run once
Take MiniMax H3 at its 768P rate: $0.08/second. A 120-second video, priced like a single clip, looks like $9.60.
But H3 caps generations at 15 seconds, so a 2-minute video is 8 separate generations:
2 takes per shot (optimistic): ≈ $19
4 takes per shot (typical for anything needing a consistent character across shots): ≈ $38
Regenerate the 8 keepers up to 2K (at $0.05/second, with reference inputs rebilled): +≈$6 → ≈$44 all-in
Same model, same video, a 4.6x spread depending on which number you were shown. Confirmed against MiniMax's own pricing page (platform.minimax.io/docs/guides/pricing-paygo) as of August 21, 2026.
| Model | Max clip | Native audio? | Reference control | Price/sec | Cost, 2-min video (2×/4×) |
|---|---|---|---|---|---|
| MiniMax H3 (Hailuo 3.0) | 15s | Yes | Up to 5 free ref. images | 768P $0.08/s · 2K $0.13/s | $19 / $38 (+$44 w/ 2K regen) |
| Runway Gen-4.5 | 5–10s | No — separate pass | Ref. images; retries billed same as originals | ≈$0.12/s | $29 / $58 |
| Kling 3.0 | Up to 15s | Yes, Pro mode | Up to 7 ref. images or a ref. video | 6–12 credits/s ≈ $0.07–$0.16/s by plan | ≈$38 / ≈$77 (1080p + audio) |
| Google Veo 3.1 | 8s | Yes — every tier | Up to 3 ref. images | Lite $0.05 · Fast $0.10 · Std $0.40 (720p) | $96 / $192 (Std) · $24 / $48 (Fast) |
| ByteDance Seedance 2.5 | 30s | Yes | Up to 50 refs (images/video/audio) | Billed per-token, not per-second | $24–54 / $48–108 |
Retry assumption, stated plainly: every cost figure above assumes you keep the first acceptable take at 2× the shots, or need a 4th attempt on average at 4×. Character-consistency work — the same face or product across every generation — runs closer to the 4× column in practice.
*Where these numbers come from: MiniMax and Runway were confirmed direct against their own pricing pages (platform.minimax.io, runway.com/pricing). Kling, Veo, and Seedance are now confirmed too, verified August 23, 2026 — Veo against Google's own Gemini API pricing page, Kling against Kling's own VIDEO 3.0 credit guide, Seedance against Volcano Engine's model-pricing page. One caveat survives on Kling: it bills in credits whose dollar value moves with your plan tier, from $1.33 down to $0.62 per 100 credits, so its per-second figure is a range rather than a single price. Worth noting on Veo: audio is included at every tier, not only Standard, which is what makes its Fast tier the cheapest audio-included option in this table. Seedance specifically bills per-token, not per-second, and ByteDance has never published a USD rate; the dollar figures above are this piece's own conversion at the day's spot rate, not a ByteDance number.
“Rendered locally” isn't a path most readers can follow
MiniMax shipped H3's open weights on Hugging Face on August 2, 2026 — but the license excludes the US, EU, UK, and South Korea from local deployment entirely, for both running the model and using its outputs. Section V.4 of the license restricts the outputs, not just the act of running the weights. MiniMax's own Q&A on the Hugging Face discussion thread ties that exclusion directly to “ongoing copyright-related legal proceedings specifically concerning generative video AI.”
Two more gaps between the demo and what you'd actually get running it yourself:
The widely-shared “12GB VRAM” figure is a community quantization-plus-offloading trick — MiniMax's own reference implementation (BF16, SGLang) uses four GPUs. And the open release, H3-Base, tops out at 768p; the prompt-preprocessing and 2K-upscale steps that make the hosted product look as good as the viral clip — H3-Context-IR and H3-Regenerate-2K — stay API-only. Local output isn't what the demo showed.
Worth one neutral line: MiniMax is a defendant in a copyright suit brought by Disney, Warner Bros. Discovery, and NBCUniversal (C.D. Cal., filed September 16, 2025). The court denied MiniMax's motion to dismiss on May 26, 2026, finding the studios plausibly alleged direct and secondary infringement including “near perfect likenesses” of characters. The case is in discovery, unresolved.
Reality check: every number above is one company's or one tracker's price sheet on one day in August 2026. None of it is a benchmark of output quality — a cheap take that needs a 5th regenerate isn't actually cheap, and none of these numbers move if you need a 6th.
A robot that learns a chore from watching it once
Everyone comparing AI video models is still arguing about which one looks best in a demo. A few days after that viral clip, a much quieter release argued that the more interesting AI progress right now isn't happening in a chat window or a render queue — it's showing up in a robot arm.
On August 19, Generalist AI announced GEN-1.5, its embodied foundation model. Show it a task once — a 3-12 second video clip — and it attempts the task immediately. No gradient updates, no fine-tuning. The company calls the demonstration a “physical prompt”: load it into the model's 30-second context window, and a robot arm just tries it, emitting action commands 100 times a second from video, sensor, and language input.
The numbers, as reported by Generalist and confirmed across independent coverage:
• 59% average success, one-shot, straight from the pretrained model, across 10 tasks (±10 percentage points)
• 83% average success (±9 percentage points) after 10 gradient-update steps on about 5 minutes of data (~50 demonstrations) — and those 10 steps changed the model's weights by less than 0.15%
• The capability wasn't explicitly trained for — it emerged from roughly 8 months of pretraining on large-scale physical interaction data
• The tasks were deliberately small: twisting a jar lid, retrieving money from a purse, stacking cups, sweeping trash, opening a book, unzipping a pencil pouch, removing a vacuum pad
The part that matters more than the headline number: the one-shot mode is in-context learning — the model's weights never change. The few-shot mode, the one that uses 1-10 gradient steps, is the one that actually updates the model. Most coverage conflates the two. The real milestone isn't that a robot got a task right 59% of the time off one demo — it's that eight months of broad physical pretraining made a few seconds of new experience useful at all.
A few things it does that a fixed script can't: prompts compose — two separate demonstrations shown in context chain into one continuous skill, with the model filling in intermediate motions (repositioning, regrasping, error recovery) that appeared in neither original demo. It improvises with tools it's never seen — shown a brush sweeping blocks into a bowl, then handed a banana, it used the banana as a brush; handed a dustpan instead, it switched from sweeping to scooping. And skills transfer from simulation, and even from a human hand — prompts built entirely in simulation produce working behavior on a real robot, and in some cases a human demonstrating with their own hands, seen only through the robot's cameras, is enough to teach it.
Everyone covering this is reaching for the same line: the GPT-2-to-GPT-3 moment for robotics. That's Generalist's own comparison — CEO Pete Florence frames it that way directly — not an independent assessment, so take it as an attributed claim, not a fact.
One claim circulating is flatly wrong and worth naming so you don't repeat it: at least one aggregator credits GEN-1.5 to MIT's Improbable AI Lab and Stanford, and calls it open source. Neither is true. It's Generalist AI, founded in 2024 by ex-Google DeepMind researchers Pete Florence and Andy Zeng and ex-Boston Dynamics roboticist Andrew Barry, and there's no open-source release. The company raised $400M in June 2026 at a $2B valuation, led by Radical Ventures with Nvidia, Bezos Expeditions, and Union Square Ventures participating.
Reality check: ten short, deliberately simple tasks, all numbers self-reported by the company that built the model, no independent replication yet. A demo generalizing across 10 curated tasks is not the same as dependable performance in an actual warehouse or home.
The thread connecting both stories
Generative video is getting cheap enough to run that the sticker price stopped being the interesting number — the interesting number is what a finished piece actually costs once you account for how the tool really behaves. Embodied AI just crossed a different threshold: the sticker price to teach a robot something new just dropped from a training run to a single video clip. Different domains, same shape of story — the frontier is moving from “can it do the demo” to “what does it actually cost you to use this for real.”