Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

21–40 of 514

Sep 10Thursday

AI HOT (Curated Pool)

DeepSeek-V4.1-Flash lands on SiliconFlow, a 552B MoE with 1M context window

SiliconFlow launched DeepSeek-V4.1-Flash on Day 0. It's a 552B MoE model with ~8B active params during prefill and ~16B during decode, native vision, and a 1M-token context window. KV cache footprint is about 1/4 of V4 Flash, which helps with deployment cost. MIT license keeps commercial use straightforward.

Why it matters: Same-day availability of DeepSeek V4.1-Flash on SiliconFlow, with KV cache reduced to 1/4 of V4 Flash — a clear deployment cost signal. Score held at 78 because this is a platform availability announcement; no benchmarks or real-world performance data yet.

AI HOT (Curated Pool)

DeepSeek V4.1-Flash cuts KV cache memory for AI agents to a quarter of its predecessor

DeepSeek released V4.1-Flash, a 552B-parameter model built to slash memory costs for AI agents. Its KV cache in fast GPU memory is about a quarter the size of V4-Flash, and the offloaded portion shrinks to roughly an eighth. The model splits into an encoder and decoder: only 8B parameters activate per token during input processing, versus 16B during text generation, nearly halving input compute. It supports 1M-token contexts and stores the main KV cache in FP4. On the DeepSWE v1.1 coding benchmark it scores 74.2%, narrowly beating Anthropic Opus 5 and OpenAI GPT-5.6 Sol, but it still trails on complex scientific tasks and image analysis. Weights are on Hugging Face under the MIT license. The post does not disclose inference latency or specific hardware requirements.

Why it matters: DeepSeek drops V4.1-Flash targeting agent memory costs — KV cache down to 1/4 of predecessor. Concrete architecture numbers, not vapor. Held at featured rather than p1 because only one source so far (no cross-source cluster yet) and the post doesn't disclose real latency/throu...

AI HOT (Curated Pool)

DeepSeek V4.1-Flash drops with native vision and a big price cut

DeepSeek released V4.1-Flash with a new Causal Encoder-Decoder architecture and native vision understanding — no separate vision model needed. It's a 552B MoE, activating 8B params for input and 16B for output. The post doesn't disclose the exact price cut or benchmark numbers, so I'd wait for third-party evals before getting excited.

Why it matters: DeepSeek ships a new model with a causal encoder-decoder architecture that replaces bolt-on vision components. 552B total params, only 8B/16B active during inference. The post claims a price cut but gives no specific numbers or benchmarks, so the score stays below 85 until thi...

AI HOT (Curated Pool)

DeepSeek Releases V4.1-Flash: New Causal Encoder-Decoder Architecture with Native Vision

DeepSeek V4.1-Flash is the smallest model in the new architecture family: a 552B MoE with 8B active params for input and 16B for output. It uses a Causal Encoder-Decoder design with native vision. KV cache drops to 1/4 of HBM and 1/8 of SSD storage vs the previous generation, and API pricing is lower. The post doesn't disclose exact pricing or vision benchmarks.

Why it matters: DeepSeek ships a new architecture — not a V4 refresh but a Causal Encoder-Decoder with native vision and dramatically reduced KV cache. The 552B total / 8B+16B active MoE config directly impacts deployment economics. Domestic Chinese flagship model release triggers the positiv...

Sep 9Wednesday

AI HOT (Curated Pool)

OpenAI launches ChatGPT Images 2.5 with two models and sharper editing

OpenAI released ChatGPT Images 2.5 with two models: Flare for speed (up to 50% lower latency) and Sunburst for tighter editing control at longer generation times. API pricing is $8/1M input tokens and $30/1M output tokens. New 'xhigh' and 'max' quality tiers push a 1024x1024 max image to roughly $0.21. No batch pricing is offered yet, and OpenAI didn't share an average per-image cost. ChatGPT users can't manually pick the model; the post's tests show Work mode keeps edits more stable than Chat mode. Both models top the image arena rankings, and watermarking is added in partnership with Google DeepMind.

Why it matters: Substantive image-gen upgrade from OpenAI with clear dual-model split and concrete performance/pricing numbers. But it's an iteration, not a paradigm shift, and Plus users can't access Sunburst yet — that caps the score.

Sep 8Tuesday

OpenAI News

OpenAI launches ChatGPT Images 2.5 with faster generation and sharper editing

OpenAI released Images 2.5, a new image model that cuts generation latency by up to 50% and improves lighting, textures, and multi-turn editing consistency. Over 3 billion images are already created weekly across ChatGPT and the API. A new Sketch feature lets users draw directly in ChatGPT as a reference. API availability is confirmed, but the post does not disclose pricing details.

Why it matters: OpenAI officially released Images 2.5 with 50% lower latency, quality improvements, a new Sketch feature, and 3B images/week volume. It's a substantive update to a core ChatGPT capability, hitting all three HKR axes. Not scored higher because this is an iterative upgrade rathe...

Sep 6Sunday

QbitAI · WeChat

GPT-6 Astra directs ByteDance Seedance 2.5, handling script-to-edit pipelines

Users chained GPT-6 Astra with ByteDance Seedance 2.5 into an end-to-end AI film pipeline: Astra builds scenes and previs in Blender, Seedance turns reference frames into anime-style clips, and Astra handles the final edit. The same workflow produced a Naruto fan short and a US remake of a Chinese drama. Fable 5.1 and Gemini 3.8 Flash were also tested as prompt writers for Seedance, showing distinct directorial styles. Separately, Astra was used as a DaVinci Resolve colorist, matching a reference look in 4 minutes, though opinions on the result were mixed. The post does not disclose Seedance 2.5 technical specs or pricing.

Why it matters: A hands-on experiment chaining OpenAI and ByteDance's latest models into an automated filmmaking pipeline, with concrete steps and outputs. But it's a personal workflow share, not a product update or official partnership, so it lands right at the featured threshold.

Hacker News front page

GPT-6 Astra on robot arms: 95% on block-in-bowl, still stuck on puzzle insertion

Robocurve gave GPT-6 Astra control of YAM arms on two tasks, head-to-head with Claude Fable 5.1. On block-into-bowl, Astra scored 19/20 (95%) vs Fable 5.1's 8/20, averaging 2.5 min and $0.94 per run—less than half the time and cost of Fable 5.1's 6.8 min and $2.12. On the puzzle-insertion task, Astra managed 2/20, same as Fable 5.1; both stall at the final alignment step, at $1.36 per run. Clear win on pick-and-place, no progress on fine insertion.

Why it matters: Named first-person experiment with numbers and a direct model comparison — hits all three HKR axes. The puzzle-task stall for both models adds credibility. Not p1 because it's a third-party eval, not an official release, and only two tasks tested.

Sep 2Wednesday

Hacker News front page

World Labs introduces Atlas, a world model that natively understands 3D space

Atlas is a multimodal autoregressive diffusion transformer pretrained from scratch to handle text, images, video, and 3D. It stitches reference images into a coherent 3D scene with pixel-perfect camera control, generating up to 1 minute of 1440p video. On sparse-view 3D reconstruction (2–3 images), Atlas beats specialized reconstruction models; more input images reduce guesswork. Early access is open for request, but the post doesn't disclose pricing or a public launch date.

Why it matters: World Labs drops Atlas, a from-scratch omni world model that unifies text, images, video, and 3D into a shared spatial context. The sparse-view reconstruction claim — beating specialized models with only 2-3 images — is concrete and testable. Not a 95 because it's a blog post ...

TechCrunch · AI

Google Pics is an AI-first design tool that takes prompts instead of manual editing

Google launched Pics, an AI design tool that generates posters, social posts, and illustrations from text prompts. It's part of Workspace for business users and Google AI Pro/Ultra subscribers, running on the Nano Banana image model. Unlike Canva or Adobe Express, there's no marketplace for creator templates—everything is AI-generated from scratch. The post doesn't disclose pricing or exact rollout dates, only 'over the coming weeks.'

Why it matters: Google's Pics is a prompt-to-design tool powered by its Nano Banana model, targeting Canva's space with an enterprise-first rollout. Score capped at 72 because the post lacks details on template ecosystems and collaboration — it reads more like a feature demo than a full produ...

AI HOT (Curated Pool)

Gemini gets agentic video understanding that can watch and act on screen

Google DeepMind added agentic video understanding to Gemini: it can watch a video of a UI and then perform the same clicks, typing, and scrolling itself. Instead of just describing what it sees, Gemini executes multi-step tasks like filling web forms or completing an order in a mobile app. The feature is now available for testing in the Gemini app and Google AI Studio. The post doesn't disclose latency or success rates—real-world UI agent reliability is still a big open question.

Why it matters: Google DeepMind added agentic video understanding to Gemini — it learns UI workflows from screen recordings and executes multi-step tasks, now available in the Gemini app and AI Studio. Hits all three HKR axes, but the post doesn't disclose latency or success rate, the two num...

Aug 31Monday

AI Chat-Group Daily (群聊日报)

Astra frontend one-shot leak, coding growth economics, and Claude safety downgrade that deleted 700GB

OpenAI is gray-testing Astra, a model that one-shots full frontend webpages from scratch—testers declared 'frontend is solved.' Anthropic is rushing Fable 5.1, and both sides are already trading SVG stability comparisons. Meanwhile, Claude Code's safety mechanism downgraded a dangerous file-cleanup task to the weaker Opus 4.8, which correctly identified the home directory as off-limits, then deleted 700GB of it anyway. A coding growth analysis shows non-engineer Codex usage growing 108x in legal, 41x in sales, with broad coding tasks driving 60–70% of OpenAI ARR. Hy4 preview scaled up urgently after a usage spike, but real-world prefill hits ~20K tokens and long sessions take 24.7s. Dual GB10 running DeepSeek V4 Flash hit 200.3 tok/s aggregate throughput at 6 concurrency. Fireworks delayed GLM-5.3-Flash by two days after discovering EvalScope prompts caused 2–3x overthinking. The group also discussed orthogonal design for cheaper code review and a prescription for vibe coding addiction: no agent one hour before bed.

Why it matters: The Astra leak vs Fable 5.1 head-to-head is the most watchable narrative this week — four concrete technical directions give it substance, and the 'frontend is solved' claim hits a nerve. But the source is a chat-group digest relaying a WeChat article and tweet screenshots, wi...

Aug 30Sunday

Computing Life · Share · Yage

The value of multimodal models isn't understanding images—it's deciding to look

Meta, Z.ai, and DeepSeek each released multimodal models in August with strikingly similar demos: the model observes a video or screenshot, calls tools to generate a webpage, slides, or a mini-game, then inspects its own output. This shifts vision from a passive input channel to an action the model initiates. The article likens it to the 2023 shift from static RAG to agentic RAG, but notes the loop direction is reversed—here the model self-verifies after producing. Evaluation moves beyond image Q&A: Meta's WildArtifactBench uses pairwise comparisons and Elo scores to assess full artifact creation. Training also changes; both GLM and Meta train models in generate-inspect-revise loops, logging interaction trajectories as training data. For builders, the key question is no longer static image accuracy but whether the model can complete an observe-generate-inspect closed loop.

Why it matters: Three labs independently demo the same multimodal pattern—shifting from passive image understanding to an active observe-produce-verify loop—with a convincing analogy to the 2023 agentic RAG paradigm shift. Points off because this is a commentary synthesis rather than a primar...

Aug 26Wednesday

AI HOT (Curated Pool)

Alibaba Qwen releases Qwen3.8-Flash, a 125B MoE model activating only 6B per token, as an early preview of the Qwen4 architecture

Alibaba Qwen open-sourced Qwen3.8-Flash with full weights. It's a multimodal MoE model with 125B total parameters, activating only 6B per token. Training cost is 1/9 of Qwen3.7-Plus while outperforming it across the board. Production API pricing is $0.16/1M input tokens and $0.47/1M output tokens, with 262K native context expandable to 1M. The model also serves as an early preview of the Qwen4 architecture.

Why it matters: Alibaba Qwen open-sources Qwen3.8-Flash, a 125B MoE model activating only 6B per inference, with 1/9 the training cost of its predecessor and claimed performance gains, plus a Qwen4 architecture preview. Domestic flagship release with concrete numbers — HKR all hit. Not 90+ be...

TechCrunch · AI

Z.ai confirms it built Ox Alpha, the anonymous model topping leaderboards

Z.ai confirmed it is the lab behind Ox Alpha, the open-weight model that appeared anonymously on OpenRouter and immediately topped rankings. The company calls it the newest GLM iteration, built for coding, sustained agentic work, and multimodal reasoning. Weights drop Wednesday for developers to build on. Earlier GLM-5.3 already matched Anthropic's Fable 5 on some benchmarks. Ox Alpha adds more pressure on frontier pricing from OpenAI and Anthropic.

Why it matters: Revealing the identity of a chart-topping anonymous model is inherently newsworthy; Z.ai also commits to open-sourcing weights on Wednesday and clearly positions the model for code, agents, and multimodal reasoning. The score is held back because the article provides no benchm...

AI HOT (Curated Pool)

Alibaba Qwen releases Qwen3.8-Flash: a 125B multimodal MoE activating only 6B per token, trained at 1/9 the cost of Qwen3.7-Plus

Qwen3.8-Flash is an early preview of the Qwen4 architecture: 125B total params, only 6B active per token. Native context is 262K, extendable to 1M. Training cost is just 1/9 of Qwen3.7-Plus, with better coding and office-task performance. Weights are open. The post doesn't disclose specific benchmark scores or license details.

Why it matters: Alibaba Qwen drops Qwen3.8-Flash as an early Qwen4 architecture preview: 125B total params, 6B active, trained at 1/9 the cost of Qwen3.7-Plus. Weights are open. The efficiency numbers are concrete, but the post doesn't disclose specific benchmarks or the open-source license, ...

Aug 25Tuesday

Hacker News front page

Qwen 3.8-Flash-Next open-release tomorrow: 125B total, 6B active MoE model

Qwen teased Qwen3.8-Flash-Next on ModelScope, a multimodal MoE model built on the next-gen Qwen4 architecture with 125B total and ~6B active parameters. The early release is meant to preview Qwen4's design for the community. It drops 2026-08-26 15:00 UTC, with an FP8 variant alongside. The post doesn't disclose benchmarks, inference speed, or specific multimodal capabilities—I'll hold judgment until the model card lands.

Why it matters: Qwen is previewing the Qwen4 architecture with a 125B-total / 6B-active MoE design — real new information with high attention in the Chinese open-source community. The deduction is because it's not open-sourced until tomorrow, and no benchmarks or inference speed data are avai...

OpenAI News

OpenAI bans Russian accounts behind a covert influence campaign posing as an Israel-based think tank

OpenAI banned a cluster of Russia-based ChatGPT accounts used to promote the International Burke Institute (IBI), a fake think tank claiming to be in Israel. The site copied academic work, used machine translation, and published a sovereignty index favoring Russia. Operators prompted the model in Russian to generate English social media posts while hiding linguistic clues. OpenAI calls this the most elaborate Russia-linked IO they've disrupted since the Ukraine war began, though it reached relatively small audiences.

Why it matters: OpenAI's first-party disclosure of a Russian covert influence campaign using ChatGPT, with concrete operational details. Held at 78 because it's a routine security takedown rather than a product capability leap, and the audience fit is narrower.

Aug 22Saturday

Hacker News front page

Hollywood creatives are training AI to do their own jobs

The Guardian reports on Hollywood concept artists, voice actors, and writers hired by AI firms to train models at $25–$150/hour. They know they are teaching AI to replicate their own craft—some call it 'digging the grave of my profession.' Runway and Sora are named as key employers, but the article doesn't disclose contract scale or training data volume. Treat this as an industry mood piece rather than a quantified displacement forecast.

Why it matters: A conflict-rich mood piece on Hollywood creatives training their own AI replacements. The headline and first-person quotes carry strong tension, but the body lacks hard numbers on contract scale or training data volume—more feature than hard news. H and R hit, K absent, lands ...

Aug 18Tuesday

Hacker News front page

Muse Glimmer fits an agent on-device with a memory hierarchy disguised as a 30B Transformer

Meta's Muse Glimmer is a ~30B multimodal model built to run agentic tasks offline on consumer hardware. The BF16 checkpoint is 55 GiB; Meta ships ~4-bit quantized versions that bring the language model under 20 GB. Architecturally, only every fourth of the 52 layers uses full-context attention—the other 39 use a 2,048-token sliding window. Global layers drop RoPE and retrieve by content. The KV cache stores just two key/value heads while 32 query heads provide diverse retrieval behaviors. A large ViT handles perception once, compresses neighboring patches 4:1, and feeds them as tokens. The result is a memory hierarchy: local layers build ordered representations, global layers search across the full sequence, and the tiny KV cache means quantization savings translate directly into longer context or larger batches.

Why it matters: A solid architecture deep-dive with real numbers on quantization cost, attention hierarchy, and QK norm. But it's a third-party analysis, not a Meta launch, and the pure-architecture focus raises the bar for readers outside on-device deployment — so it lands right at the featu...