Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

41–60 of 514

Aug 17Monday

AI HOT (Curated Pool)

Qwen 3.8 27B is excellent, but defaults to wildly overthinking things

Simon Willison tested Alibaba's Qwen 3.8 27B and found the default xhigh reasoning effort causes absurd overthinking. A simple circle prompt triggered minutes of animated SVG generation; a pelican-on-a-bike SVG burned 22,276 reasoning tokens over 21 minutes. Turning reasoning off cut the same task to just over two minutes. He recommends starting with low or no reasoning. The model also nailed bounding-box detection on a pelican photo with near-perfect accuracy.

Why it matters: Simon Willison's hands-on test of Qwen 3.8 27B reveals severe overthinking from default reasoning settings, with concrete token and time comparisons. A data-backed first-person experiment directly useful for local deployment users. Not above 80 because the core finding is a co...

Aug 16Sunday

TechCrunch · AI

Woman claims stepfather used Grok to turn her childhood photo into 7,000+ explicit images

A woman, Jane Doe 4, joined a class-action lawsuit against xAI, alleging her stepfather used Grok to create over 7,000 explicit images from a photo taken when she was 11. Her stepfather died by suicide two days after a law enforcement raid uncovered the images. Three Tennessee teenagers had previously sued xAI, claiming Grok lacked basic safeguards to prevent generating explicit imagery of real people, including minors. Earlier this year, X was flooded with millions of Grok-generated sexualized images. TechCrunch has reached out to xAI; the post does not include a response.

Why it matters: This isn't a product update or a paper — it's a new filing in a class-action lawsuit that pins Grok's safety gaps to a horrifyingly specific case. 7,000+ images, an age-11 source photo, a suicide — every detail forces the question of where xAI's content moderation line actuall...

Aug 15Saturday

AI HOT (Curated Pool)

MOSS-VL: An open VLM family that treats real-time interaction as a first-class capability

Fudan's MOSS-VL makes real-time interaction—perceiving while speaking—a first-class capability. Gated cross-attention keeps visual tokens outside the decoded sequence, giving it a 2.8× to 5.1× time-to-first-token advantage over same-backbone Qwen3-VL-8B. MOSS-VL-Realtime tops three of four streaming benchmarks, hitting 66.0 vs. 37.5 on OmniMMI Proactive Alerting. The offline variant leads temporal-reasoning video sets at comparable scale. All five checkpoints, the training curriculum, and inference code are open.

Why it matters: Fudan open-sourced a VLM family that treats real-time interaction as a first-class capability. Gated cross-attention cuts time-to-first-token by 2.8–5.1× vs. Qwen3-VL-8B on the same base, with the gap widening as frames increase. H and K are solid hits, but R is weak—the open-...

Computing Life · Share · Yage

Give DeepSeek V4 a 3B vision front-end — it works today

DeepSeek V4's API rejects image input, but Liquid AI's new LFM2.5-VL-3B can serve as a perception layer. The author tested it zero-shot on garage camera data — 94.3% accuracy for door open/close, 1.5s locally. Perception runs on-device, images never leave, only structured text hits the cloud for reasoning. The latency, privacy, and cost advantages of this split architecture hold even if DeepSeek adds native vision later. One gap: extracting a reliable confidence signal from natural-language output — the post doesn't detail a calibration method.

Why it matters: The author built a perception frontend for DeepSeek V4 using Liquid's LFM2.5-VL-3B, tested it on a real garage camera with 94.3% accuracy and 1.5s latency — concrete data, clear architecture. Also traces DeepSeek's vision research history (VL2, the retracted Thinking with Visu...

Aug 14Friday

AI HOT (Curated Pool)

Qwen releases Qwen3.8 series: a 27B dense multimodal model and open weights for a 2.4T-A95B Max variant

Qwen delivered on its open-source promise with the Qwen3.8 series. Qwen3.8-27B is a natively multimodal dense model that beats Qwen3.7-Plus at only 27B parameters, supports 262K context natively and up to 1M tokens via YaRN, under Apache 2.0. Open weights for the Max-tier Qwen3.8-2.4T-A95B are also available. The post doesn't cover training data, inference cost, or release timeline details.

Why it matters: Alibaba Qwen drops Qwen3.8 series: a 27B dense multimodal model that beats Qwen3.7-Plus on benchmarks, with native 262K context and Apache 2.0 license, plus a 2.4T MoE Max variant. This is a same-day must-cover for a major Chinese open-source release. Not pushing past 90 yet b...

Aug 13Thursday

AI Chat-Group Daily (群聊日报)

Closed-source reasoning chains extracted at scale; Coze CLI hijacks AI tools

The big one today: researchers extracted hidden reasoning chains from Anthropic, OpenAI, and Google models at scale. The trick is absurdly simple—take Opus 4.8's encrypted CoT and feed it to Haiku 4.5, which decodes it verbatim. All three API families were broken, and decoding 10K trajectories costs about $720. A separate paper shows you can even reverse-engineer reasoning from public outputs alone using a 1.5B-param model. Separately, Coze CLI was caught silently scanning local Codex and Claude Code directories and injecting its own skills into workflows. On the engineering side, the group discussed how prompt debt now rivals traditional code debt—old rules pile up, evals lag behind model iterations, and nobody dares delete anything.

Why it matters: Strong cross-source cluster signal (chat digest + original paper + study notes). First systematic validation that encrypted CoT from three major vendors is cross-model decodable, with concrete $720/10k cost. All three HKR axes hit, but the source is a secondary digest rather t...

Aug 12Wednesday

Google DeepMind

Google DeepMind releases SL2T sign language-to-text model, first in Pixel 11 Gboard and Live Transcribe

Google DeepMind released SL2T, a multilingual sign language-to-text model, bringing sign language AI into consumer products for the first time. On Pixel 11, Gboard and Live Transcribe support American Sign Language (ASL) to English dictation, with more devices and languages to follow.

Why it matters: It gives SL2T's training scale, benchmark results and privacy design, so readers can judge the real limits of sign language translation in consumer products.

AI HOT (Curated Pool)

Google Gemini app hits 1B monthly users, matching ChatGPT's June milestone

Sundar Pichai announced on X that the standalone Gemini app surpassed 1 billion monthly active users—Google's 14th product to hit that mark. The figure excludes AI Mode in Search and other channels. ChatGPT reached 1B MAU in June; Gemini is now keeping pace. Usage stats: 63% of users have tried the voice feature, the app generates over 150 million images daily, and iOS has more than 100 million active users. The post doesn't disclose paid user share or revenue.

Why it matters: Gemini's standalone app hitting 1B MAU, neck-and-neck with ChatGPT, is a major consumer AI milestone. Score stays at 78 rather than higher because this is a growth metric, not a capability breakthrough, and the 63% vision usage stat lacks detail on what 'using vision' actually...

Aug 11Tuesday

New York Times Chinese

Meta releases open-weight Muse Glimmer, a free version of its paid Muse Spark model

Meta released Muse Glimmer on Monday, an open-weight AI model nearly identical to its paid, closed-source Muse Spark launched in July—capable of generating code, text, and images. Mark Zuckerberg also published a 14-page essay arguing superintelligence should not be concentrated in a few companies, and announced a $1 billion fund for communities hosting its data centers. Muse Glimmer is open-weight, not fully open-source; the underlying code isn't fully public. Meta also teased a more powerful model codenamed Watermelon but didn't disclose whether it will be open or closed.

Why it matters: Meta open-weights a near-clone of its paid closed model Muse Spark, paired with a 14-page Zuck essay arguing superintelligence shouldn't be locked in a few companies and a $1B community pledge. It's a product launch, a positioning statement, and a funding move rolled into one ...

Aug 10Monday

AI HOT (Curated Pool)

SGLang adds Day-0 inference support for Meta's local agent model Muse Glimmer

Meta released Muse Glimmer, a 30B multimodal model built for local agentic workflows. SGLang ships Day-0 support with dedicated optimizations: on a single RTX 5090 with NVFP4 quantization and DFlash speculative decoding, per-user decode hits 236 tok/s and total throughput reaches 1,452 tok/s. The model uses a hybrid of sliding-window and full-sequence attention with a 128k+ context window. Apple Silicon is supported via the MLX backend, though speculative decoding isn't available there yet.

Why it matters: Meta shipping a new model is an industry event, but this post centers on SGLang's inference optimization, not the model itself. Concrete perf numbers (236 tok/s, 1452 tok/s total throughput) give it enough knowledge density to clear the featured bar, though the narrow audience...

Hacker News front page

Meta open-sources Muse Glimmer, a 30B agentic model that runs locally on a single GPU

Meta released Muse Glimmer weights under Apache 2.0. It's a 30B model built for always-on local agent workflows, small enough to run on a Mac or PC with a single consumer GPU. 4-bit quantization shrinks it below 20 GB, leaving room for KV cache and the vision encoder within a 24 GB or 32 GB envelope. Training used logit distillation from a larger Muse Spark teacher, followed by mid-training on long-context agent data and post-training with SFT, on-policy distillation, and RL. Meta's benchmarks show it outperforming Gemma4-31B and Qwen3.6-27B on agentic, coding, multimodal, and safety evals. The post doesn't disclose specific latency numbers, only that inference optimizations were applied to keep it responsive.

Why it matters: Meta drops a 30B local agent model under Apache 2.0, quantized under 20 GB for consumer GPUs. Clear positioning — not a general chatbot but purpose-built for always-on agent workflows. Score held back from higher bands because we only have the launch blog; third-party benchmar...

Aug 6Thursday

AI HOT (Curated Pool)

Alibaba Cloud launches Qwen-Image-3.0 with high-res image generation starting at $0.03

Qwen-Image-3.0 is pitched as production-ready: 4.5K-token prompts, 100%+ text accuracy with no broken logos, and native support for 12 languages. High-res generation starts at $0.03. The post only provides a headline and links—no model architecture, inference speed, or benchmarks are disclosed, so I'd hold off on the accuracy claim until third-party tests appear.

Why it matters: Qwen's first dedicated image gen model, priced at $0.03 with a 4,500-token prompt ceiling and 12-language support — real differentiators. Score held back because the source is a single tweet with no architecture details, inference speed, or third-party benchmarks; the '>100% t...

Product Hunt · AI

Mistral launches Shieldstral: a 3B open-weight multimodal guardrail that defines safety policies in natural language at inference

Mistral AI launched Shieldstral on Product Hunt this week—a 3B open-weight multimodal guardrail. You define safety policies in natural language at inference time, and it evaluates text, images, or both with a single token output. It runs locally on one 16GB GPU. The Product Hunt page shows the basics and screenshots, but doesn't disclose latency, accuracy benchmarks, or comparisons with other guardrails like Llama Guard. Treat it as a lightweight, self-hostable compliance component for now; real-world performance needs community benchmarks.

Why it matters: Mistral dropped a 3B multimodal safety guardrail that runs locally and takes plain-language rules—practical for AI app builders. Score held back because the Product Hunt page gives no latency or accuracy numbers, so production readiness is unclear.

AI HOT (Curated Pool)

Simon Willison one-shots a full 3D Raccoon Heist game with Claude Fable 5

Simon Willison fed a 2022 tweet and two concept images to Claude Fable 5 and let it build a playable browser 3D game with zero further input. The model chose Three.js, called OpenAI's gpt-image-2 for textures, and added mechanics like a patrol dog with scent tracking. The whole project was done on mobile, deployed via GitHub Pages. The gameplay is basic, but the zero-intervention workflow is the real story.

Why it matters: Simon Willison's first-person experiment is a quality signal on its own. One old tweet plus two concept images, and Claude Fable 5 autonomously handled tech stack, texture generation, and deployment — the information density is high. Not scoring higher because the gameplay is ...

Aug 5Wednesday

Hacker News front page

Qwen Image 3.0 Pro targets production use with 4.5k-token layouts, 10px text, and realistic detail

Qwen Image 3.0 Pro handles up to 4,500 input tokens and generates dense layouts—newspapers, storyboards, menus—in one pass. It reliably renders text down to 10px across 12 languages and 20+ fonts, and reproduces micro-expressions, pores, and hair strands at near-photographic quality. Output pricing is $0.04 per 1K image and $0.075 per 2K image, but the rate limit is just 1 request per minute, so high-throughput use cases are off the table for now.

Why it matters: Qwen ships a new image model with concrete specs — 4.5k token input, nested-image layouts, 10px text rendering — not marketing fluff. $0.04 per 1K images is competitive, but the 1 request/minute rate limit bottlenecks batch use, capping the score.

Hacker News front page

Mistral releases Shieldstral: a 3B open-weights model for multimodal moderation

Mistral introduced Shieldstral, a 3B-parameter open-weights model built to moderate both text and image content. It can run locally or on edge devices, checking user inputs and model outputs for policy violations. The post doesn't disclose benchmark scores, latency figures, or pricing—only that it's positioned as a safety filter. I'd wait for third-party evals, but the 3B size is genuinely lightweight for self-hosted moderation.

Why it matters: Mistral dropped a 3B open-weights multimodal moderation model — light enough for local deployment, useful for teams running their own safety stack. But no benchmarks or latency numbers in the post, so real-world performance is still an open question, capping the score here.

Aug 4Tuesday

AI HOT (Curated Pool)

SenseTime open-sources SenseNova U1: unified reasoning and image generation in one model

SenseTime open-sourced SenseNova U1, a model that handles reasoning and image generation in a single pipeline. It can turn a prompt into a structured slide deck or generate step-by-step illustrated content, like a six-step dragon drawing tutorial. Available on HuggingFace, GitHub, and SenseNova Studio. The post doesn't disclose parameter count, training data, or benchmarks.

Why it matters: SenseTime open-sourced SenseNova U1, unifying reasoning and image generation in one model with concrete demos, not just a headline. Missing param count, training data, and benchmarks means we can't assess real capability ceiling, so score stays below 85. But releasing weights ...

AI Chat-Group Daily (群聊日报)

Qwen 3.8 Max matches Fable 5 at 2.4T params, open weights next week

Qwen 3.8 Max launched with Terminal Bench 2.1 score 86.6 and PaperBench 93.0, beating Fable 5's 88.8. A 500-yuan token plan burned out in one day; the model lands between Luna and Terra, with price as the main draw. Open weights for both Qwen 3.8 Max and Qwen 3.8-27B drop next week. DS V4 Flash hit 8T tokens consumed in a single day, topping weekly charts—the group sees tokens becoming a commodity. On tools: an M5Stick voice dongle turns a keychain into an agent remote, LoopX keeps agent state across 200+ hours, and reverse-skill injects reverse-engineering toolchain knowledge into coding agents. The wildest methodology story: an agent autonomously downloaded a local Qwen mid-translation task and auto-installed Whisper when its API key ran out of funds.

Why it matters: Qwen 3.8 Max official release, 2.4T params matching Fable 5 with open weights coming next week — a major domestic flagship model update. The chat digest provides concrete benchmarks and real-world impressions, high information density. Deduction because the source is a group c...

Latent Space

Alibaba Qwen drops Qwen3.8-Max and 27B, open weights coming next week

Alibaba Qwen announced Qwen3.8-Max, a 2.4T-parameter model, and Qwen3.8-27B, both promised as open weights. Max claims 10+ days of autonomous coding, a 125-hour self-directed research loop beating the original paper by 2.71 points, and a 4.16x return in a 365-day e-commerce sim. API pricing is $2/M input, $6/M output. I'd hold the champagne: the post doesn't include standard academic benchmarks, and the exact open-weight date and license aren't specified.

Why it matters: Alibaba Qwen drops a 2.4T Qwen3.8-Max targeting long-horizon coding and agent tasks, with concrete benchmarks. Domestic flagship release triggers the positive bump. Not 95 because we only have the official blog and Latent Space's secondhand coverage — no independent repro or c...

AI HOT (Curated Pool)

EU AI Act transparency rules kick in, with fines up to €15M for non-compliance

As of today, tech companies in the EU must label AI-generated content and disclose when users interact with a machine. Fines reach €15 million or 3% of global annual turnover. The rules cover chatbots, deepfakes, and AI-generated audio, video, and text. The EU also published official label icons companies can use instead of designing their own.

Why it matters: EU AI Act transparency rules effective today, €15M fine cap — direct compliance event for AI companies and platforms operating in Europe. HKR all hit, but this is enforcement of existing rules rather than new regulation, so capped at 78, featured threshold.