Skip to content

MiniMax

MiniMax models and products: the M series, speech and multimodal, open releases and commercialization.

Latest picks

1–20 of 25

Sep 2Wednesday

Computing Life · Share · Yage

Real-time video generation cost drops below playback time, reshaping interactive live streaming economics

fal's H3 Max Live achieves faster-than-playback generation for 5-second clips, letting viewers alter scenes via chat. 15-second clips still take 16 seconds, and cross-clip consistency is unsolved. At $144–288/hour for 768p, a single stream needs 1,400–2,900 concurrent viewers to break even. The real bottleneck is platform access rules across markets, not moderation tech.

Why it matters: A cost breakdown piece that dissects fal's real-time video generation speed, pricing, and consistency gaps. The 5s clip is indeed faster than playback, but 15s lags, and cross-clip consistency is unsolved. The $144-288/hr cost and 1,400-2,900 concurrent viewer breakeven point ...

Sep 1Tuesday

Latent Space

Fal's H3 Max Live generates video faster than real-time playback

Fal post-trained Minimax's H3 model, then ran it on their own inference engine at 35x the official endpoint speed. The result is H3 Max Live: video generates faster than you can watch it. Ethan Mollick first flagged the milestone; Fal employees turned it into an infinite Twitch stream, and Pieter Levels built a similar interactive feed. Twitch and YouTube quickly banned Fal's streams, so Fal launched fal.live. Quality is still slop, but this is the floor for real-time video generation. The post doesn't disclose pricing or exact latency numbers.

Why it matters: Fal pushed Minimax H3 to 35x speed, making video generation faster than playback for the first time — a real engineering milestone. Not scoring higher because we only have Fal's own numbers and one Twitch stream; no third-party reproduction or broader quality comparisons yet.

Aug 19Wednesday

Hacker News front page

Mojo 1.0 is now open source under Apache 2.0, alongside Modular Cloud and new silicon support

At ModCon 2026, Modular open-sourced Mojo 1.0 under Apache 2.0 and made Modular Cloud publicly available, already serving customers like MiniMax. The platform now supports AWS Trainium, Google TPUs, and Qualcomm Cloud AI 100 and Dragonfly accelerators. Mojo is getting native Windows support via a Microsoft collaboration. MAX drops device usage restrictions and moves to source-available with an open alliance program.

Why it matters: Mojo going open source is a real infra-level event — Apache 2.0 removes the biggest commercial adoption blocker. Score isn't higher because this is still an announcement; we haven't seen community traction or migration data yet. 78 feels right for now.

Aug 14Friday

AI HOT (Curated Pool)

MiniMax releases Music 3.0: open-weights music model that generates full 5-minute songs in one pass

MiniMax today released Music 3.0, an open-weights music generation model. Given a creative concept and optional lyrics, it outputs a complete song with arrangement, performance, and vocals in one pass, up to 5 minutes long. The upgrade targets three pain points: accurately interpreting creative intent, maintaining that intent across a full song, and making vocals and instruments sound performed rather than synthesized. The new Hybrid-LM architecture uses an 8B global model for song-level structure and a 0.6B local model for per-frame acoustic detail, with flow matching and a Flow-VAE converting discrete predictions to continuous audio. The post does not disclose training data scale, inference latency, the specific open-source license, or quantitative benchmark comparisons.

Why it matters: MiniMax released Music 3.0 with open weights, generating full songs up to 5 minutes with vocals in one pass. The technical paper details Hybrid-LM architecture and 8B parameters. First open-weights music model from a major Chinese lab — directly actionable for audio product te...

Aug 6Thursday

AI Chat-Group Daily (群聊日报)

MiniMax H3 open-sourced, Codex goes cloud, AI reverse-engineers WeChat, and Sol traps itself

MiniMax H3, the only open-source flagship video model this generation, released its weights with native ComfyUI support on day one. Community plugins cut generation time from 500+ seconds to just over 200. Blind tests show H3 matches Seedance 2.0 visually, though 2.5 still leads; hand physics correctness is a surprise plus. Minimum hardware is 2×RTX 4090 with 384GB RAM, production config 4×H200. OpenAI acquired Ona to move Codex to the cloud—Tibo predicts laptops will be mere control surfaces in two to three months. On the reverse-engineering front, AI plus Frida hooked PBKDF2 to extract WeChat 4.1.8 macOS database keys in one hour, bypassing removed memory signatures. Sol's over-engineering saga continues: it built a hard gate, got stuck behind it, then researched how to bypass it. Math harness day four went extreme—banning code made the model stronger through pure reasoning.

Why it matters: MiniMax H3 releasing open weights is the most concrete video-generation news this week. The blind test conclusion is clear — matches Seedance 2.0 but still a tier below 2.5, with hand-physics correctness as a surprise bonus. Hardware floor is steep at 2×4090 + 384GB RAM, which...

Aug 5Wednesday

AI HOT (Curated Pool)

MiniMax-H3 video model ported to MLX, runs on M5 Max in 45 minutes

Simon Willison got the MiniMax-H3 MLX port running on an M5 Max MacBook Pro. The model, released two days ago, takes text, images, audio, and video as input and outputs 15-second video clips with audio. He downloaded ~115 GB of model files and generated one clip from a text prompt in just under 45 minutes. The visuals looked good, but the audio came out as garbled speech-like noise—he didn't follow the official prompting guide for audio, so that part isn't a fair test.

Why it matters: Simon Willison's first-person MLX port gives the first real local inference benchmark for MiniMax H3: 45 min on M5 Max for a 15-sec clip. Valuable data, but the model isn't a flagship release and 45-minute inference is far from practical, capping it at 78.

Aug 3Monday

Hacker News front page

MiniMax H3 open-weights video model lands with day-0 ComfyUI support

MiniMax released H3, its third-gen video model, with open weights and day-0 ComfyUI integration. It handles text-to-video, image-to-video, first-and-last-frame control, and motion transfer from reference clips. Output is up to 2K, 15 seconds, with native stereo audio generated in the same pass—no post-processing. The post claims it runs locally on a 3060; I'd wait for community benchmarks before taking that at face value.

Why it matters: MiniMax's first open-weights video model with day-0 ComfyUI support is a notable openness move from a Chinese lab. Score capped at 82 because we only have the official blog and community adapter info — no third-party benchmarks or quality comparisons yet. Treating it as a high...

Jul 31Friday

AI HOT (Curated Pool)

MiniMax launches open-source H3, a multimodal model that generates 2K video with native stereo audio

MiniMax H3 unifies text, image, video, and audio understanding into one model. It generates up to 15-second videos at 2K resolution with native stereo sound. The company claims per-second pricing at 2K is under one-third of mainstream models, and 768p is under half the price of mainstream 720p. Model weights will be open-sourced in the coming days, subject to legal review. The post does not specify the open-source license or exact release date.

Why it matters: MiniMax dropped H3, a multimodal generation model that handles text, image, video, and audio in one model, outputting 15-second 2K video with native stereo sound. Pricing is aggressive — 2K per second at less than one-third of mainstream models — and weights are planned for op...

Product Hunt · AI

MiniMax launches H3, an open multimodal model that generates 2K video with native stereo sound

MiniMax launched H3 on Product Hunt, an open multimodal model for unified video generation. It takes text, image, and audio inputs to produce 2K video with native stereo sound. The model focuses on accurate text rendering, visual packaging, and complex instruction following for commercial content like motion design and branding. The post doesn't disclose parameter count, inference latency, or the specific open license. I'd hold off on the 'world-leading' claim until a technical report is out.

Why it matters: MiniMax drops H3, a multimodal video model that takes text, image, and audio and outputs 2K stereo video, targeting commercial content with text rendering and visual packaging as explicit strengths. But no parameter count, latency, or license disclosed — the 'world-leading' cl...

Jun 17Wednesday

Hacker News front page

GLM-5.2 tops open-weights leaderboard, matches GPT-5.5 on agentic benchmark

Z.ai's GLM-5.2 scores 51 on the Artificial Analysis Intelligence Index v4.1, ahead of MiniMax-M3 (44) and DeepSeek V4 Pro (44), making it the top open-weights model. It keeps the same 744B-total / 40B-active parameter count as GLM-5.1 but posts big gains in scientific reasoning and agentic tasks—HLE jumps 12 points to 40%, CritPt up 16 points to 21%. On GDPval-AA v2, a real-world agent benchmark, it hits 1524, effectively level with GPT-5.5 (xhigh reasoning). The trade-off: it averages 43k output tokens per task, up from 26k on GLM-5.1. API pricing stays at $1.4/$4.4/$0.26 per 1M input/output/cache-hit tokens, context window expands from 200K to 1M, and it ships under an MIT license.

Why it matters: GLM-5.2 hits 51 on Artificial Analysis's Intelligence Index, passing MiniMax-M3 and DeepSeek V4 Pro to become the top open-weights model. Same architecture, +11 points, same pricing. Score capped at 82 because it's a single-benchmark claim from one evaluator—no cross-source co...

Jun 15Monday

AI HOT (Curated Pool)

MiniMax open-sources M3 model weights (428B total, 23B active) with lower long-context cost

MiniMax open-sourced M3 model weights last Friday—428B total parameters, 23B active—along with the MSA sparse attention paper that cuts long-context inference cost. M3 is the first open-source model trained with interleaved text and image data from the pre-training stage. Two weeks post-release, it ranked #1 among open-source models on the Artificial Analysis Intelligence Index and GDPval-AA, reached Pareto-optimal on Code Arena WebDev, and topped Chinese models on Vals.AI. Output speed improved from ~30 TPS to ~80 TPS, with another 30–40% planned. A usage dashboard was added to the Token Plan backend.

Why it matters: MiniMax open-sourced a 428B MoE model with interleaved image-text pretraining and two #1 open-source rankings in two weeks — enough signal for featured. Held back from p1 because the post is a first-party announcement without third-party benchmarks or concrete MSA cost numbers...

Jun 13Saturday

AI HOT (Curated Pool)

MiniMax open-sources M3 weights, takes a swipe at Anthropic's export control ban

MiniMax released M3 model weights on HuggingFace. The post says 'M3 would never,' a jab at Anthropic's Fable 5 and Mythos 5 being forcibly disabled under US export controls, blocking all foreign nationals. The post doesn't disclose M3's parameter count, benchmarks, or license.

Why it matters: MiniMax open-sources M3 weights as a direct response to Anthropic's export controls — strong conflict and topicality, but the post lacks parameter count, benchmarks, and license details, capping the score.

Jun 12Friday

r/LocalLLaMA

MiniMax open-sources MSA, a sparse attention method that cuts attention compute by 28.4× at 1M tokens on a 109B model

MiniMax published a paper introducing MSA, a blockwise sparse attention built on GQA. A lightweight index branch scores KV blocks and picks a top-k subset per GQA group, then the main branch runs exact attention only on those blocks. With a co-designed GPU kernel, a 109B-parameter multimodal model achieves 14.2× prefill and 7.6× decoding wall-clock speedups on H800 at 1M context, matching full GQA quality. Code and inference kernel are open-sourced, along with a model called MiniMax-M3. The Reddit poster is curious whether the 109B model can run on consumer GPUs; the post doesn't say if weights will be released.

Why it matters: The paper has concrete mechanisms and measured numbers, not just theory—real knowledge for inference-optimization folks. But the audience is narrow (R missed), and the low-level CUDA details raise the accessibility bar for generalist readers, so I docked 3 points, landing righ...

AI HOT (Curated Pool)

MiniMax open-sources M3: 428B total params, 23B active, 1M-token context window

MiniMax uploaded M3 weights to HuggingFace, with the tech report and full weights expected in about 10 days. It's a 428B-total-param, 23B-active-param hybrid model using MiniMax sparse attention to push the context window to 1M tokens, plus native multimodal support. Coding and agent scores: SWE-Bench Pro 59.0%, Terminal Bench 2.1 66.0%, SWE-fficiency 34.8%, KernelBench Hard 28.8%, MCP Atlas 74.2%. MiniMax Code tool and API platform launched alongside. The post doesn't disclose training data, inference cost, or license terms — I'd hold off on usability judgments until the report drops.

Why it matters: MiniMax's first open-weight flagship release: 428B MoE with 23B active params and 1M context, with benchmark scores directly competing against DeepSeek and Qwen on agent/code tasks. Tech report still pending and weights just landed — clear info gaps — but the open-source move ...

r/LocalLLaMA

MiniMax-M3 open-sourced: a 428B MoE model with 23B activated parameters

MiniMax released MiniMax-M3 weights on Hugging Face. It's a mixture-of-experts model with ~428B total parameters and ~23B activated per inference. The post doesn't disclose training data, benchmarks, or minimum VRAM for local runs.

Why it matters: MiniMax dropped full weights for M3 on Hugging Face — a 428B MoE model activating only 23B per forward pass, putting it in the top tier of open-weight efficiency plays. No benchmarks or hardware requirements disclosed yet, which caps the score, but the weight release alone is ...

Jun 4Thursday

Xinzhiyuan · WeChat

Silicon Valley CEO backs MiniMax M3 as it tops open-source rankings amid Chinese community debate

MiniMax M3 ranks first among open-source models on Artificial Analysis, and the article says it supports a 1M-token context window, used 100T-scale pretraining, and will open-source its weights and full technical report within 10 days.

Why it matters: HKR-H/K/R all pass: the hook is an open-source No.1 claim amid debate, with 1M context, 100T pretraining, and weights promised in 10 days. Since weights and full report are not out, this stays in 78–84, not P1.

Jun 2Tuesday

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash as an open-weight model for agentic coding

StepFun released the open-weight Step 3.7 Flash model for fast agentic coding, with tool calling and multimodal understanding, and the model is already available in Kilo alongside MiniMax M3.

Why it matters: HKR-H/K/R pass on the open-weight agentic-coding angle and Kilo availability. Missing benchmarks, size, license, and pricing keep it at the lower featured threshold.

Jun 1Monday

AI HOT (Curated Pool)

MiniMax Releases Open-Source M3 with Coding, Long-Context, and Multimodal Capabilities

MiniMax released the open-source M3 model with coding, a 1M-token context window, and native multimodal support; M3 scores 59.0% on SWE-Bench Pro, 83.5% on BrowseComp, and costs about one-twelfth per token versus GPT-5.5.

Why it matters: HKR-H/K/R all pass: M3 has open source, 1M context, multimodal support, and 59.0% on SWE-Bench Pro. A single X post without official docs or third-party tests keeps it in the 78–84 band.

AI HOT (Curated Pool)

MiniMax M3: Frontier coding, 1M-token context, and native multimodal model

MiniMax released M3 as an open-source unified model with coding, agent, and native multimodal capabilities, supporting a 1M-token context window and using MiniMax Sparse Attention to cut per-token compute at 1M context to 1/20 of its predecessor, with over 9x faster prefill and over 15x faster decoding.

Why it matters: HKR-H/K/R all pass: MiniMax M3 has a 1M-token context hook, MSA with a claimed 20x cost cut, and open-source China-model resonance. Single official-source release keeps it in the 78–84 band, not P1.

May 31Sunday

r/LocalLLaMA

Use any model and provider with the official OpenAI Codex Desktop App without modifying its code

Reddit user thibautrey describes a 3-step setup: edit Codex Desktop config.toml, store an API key, and use a multicodex proxy alias to map gpt-5.3-codex to MiniMax-Latest. The post lists a local base_url of 127.0.0.1:1455 and says the proxy disguises returned model names as gpt-5.3-codex.

Why it matters: This is a reproducible developer workflow trick, not an official release. HKR-H comes from the lock-in workaround, HKR-K has concrete config details, and HKR-R hits cost and model-choice pressure, placing it at the tutorial featured threshold.