Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

1–20 of 514

Yesterday · Sep 29Tuesday

AI HOT (Curated Pool)

Anthropic launches Claude Sonnet 5.5, now the free-tier default on claude.ai

Claude Sonnet 5.5 beats Sonnet 5 on every benchmark, runs 30%+ faster, and costs up to 30% less for most work. The big move: it's now the free-tier default on claude.ai, which Simon Willison tested and got a solid WebGL 3D pelican on a bicycle. The 'max' thinking effort still hits the same bug as Opus 5.5—128K tokens of thought with no output, costing $1.28. 'xhigh' delivered a decent SVG in 41 seconds for 5.74 cents. Anthropic says Haiku 5.5 is coming in weeks; Simon hopes it's price-competitive with GPT-6 Luna.

Why it matters: Putting the latest Sonnet on the free tier is a real product strategy shift, not a routine model update. Simon's hands-on test delivers concrete numbers ($1.28 burned, 5.74 cents for the working render, 41-second latency), and the max-mode bug matching Opus 5.5 is a useful sig...

AI HOT (Curated Pool)

Anthropic releases Claude Sonnet 5.5, over 30% faster than Sonnet 5

Anthropic launched Claude Sonnet 5.5, claiming over 30% speed gains and clearer writing for fast-turnaround tasks like bug fixes, docs, and slide decks. Opus 5.5 targets complex judgment work, and Haiku 5.5 is coming in a few weeks. The post doesn't disclose pricing or latency numbers.

Why it matters: Anthropic model line refresh with a concrete 30% speed claim for Sonnet 5.5 and clear product-line differentiation. Held below 85 because the post doesn't disclose pricing, latency benchmarks, or the baseline for the 30% figure.

Sep 28Monday

AI HOT (Curated Pool)

Human contractors are reviewing Microsoft Copilot user prompts and uploaded images

404 Media obtained internal documents showing Microsoft hires at least hundreds of contractors to review Copilot users' prompts and uploaded images. They are not filtering for safety—they judge output quality, such as whether AI-enlarged breasts are big enough. Reviewers are flooded with sexual requests: shortening skirts, foot fetish images of children's cartoon characters, pro-anorexia content. Uploaded faces are never blurred, and many prompts are dubiously consensual. One contractor said it's hard to take the work seriously when the focus is which model generated the right bust size.

Why it matters: 404 Media obtained internal documents with solid evidence. The story exposes how Copilot's content review actually works, hitting all three HKR axes. Score capped below 85 because it's a single investigative piece, not a product launch or model release — industry shake-up is l...

Sep 25Friday

Google DeepMind

Google DeepMind releases Gemini 3.8 Live with Live Avatar

Google DeepMind released Gemini 3.8 Live with Live Avatar, adding near-real-time video generation to its native real-time conversation model. The result is a dynamic visual avatar with lip sync, natural expressions and smooth turn-taking.

Why it matters: The post details Live Avatar's real-time video conversation, async tool calls and 97-language support, a useful read on enterprise multimodal interaction.

Sep 24Thursday

The Verge · AI

Meta unveils Muse Charm, a standalone AI gadget that looks like a strapless smartwatch

Meta teased the Muse Charm at the end of Connect, a dedicated hardware device for its Muse AI agent. It resembles a chunky strapless smartwatch with a lanyard. A fingerprint sensor on the top right activates voice input; the front has at least three mic holes and a small camera. Zuckerberg noted you don't need to unlock a phone to use it. The post doesn't disclose pricing, battery life, or a release date.

Why it matters: Meta teased a standalone Muse AI gadget at the end of Connect — a thick watch-face on a lanyard with fingerprint wake, voice, and a camera. Only looks and interaction logic are disclosed; no price, battery, or launch date, so substance is thin and the score sits right at the f...

Computing Life · Share · Yage

Qwen-Image-2.1: A version rollback that packs text rendering, editing, and native RGBA into one open-weight model

Qwen released Qwen-Image-2.1 on Sep 20, a 7B open-weight image model that unifies text-to-image, local editing, and native RGBA output in a single pipeline. The version number rolled back from 3.0 to 2.1 reflects a 2026 split: 3.0 is a closed-source commercial API, while 2.1 continues the open research branch. A built-in RGBA VAE outputs PNGs with transparency, skipping external matting. The interface supports up to 10 reference images and three mask types at native 2K. The license shifted from Apache 2.0 to a research-only agreement; commercial use requires a separate license. Community tests show ~25s per megapixel image on RTX 5070/5080 at 25 steps, ~15.6GB VRAM with Q8 quantization. Text rendering remains a strength, but multi-subject consistency shows facial generalization on well-known public figures—official demos don't guarantee universal performance.

Why it matters: Qwen open-sourced a model that combines image generation, editing, and native transparency output into one pipeline — a clear engineering increment, not a reskin. The backward version jump is inherently clickable, and it resonates with both designers and developers. Not scorin...

The Verge · AI

Meta is bringing its Muse AI agent to smart glasses with voice activation

Two weeks after launching Muse, Meta says it's working on bringing the agent to its smart glasses, including the new ones shown at Connect. You'll activate it by saying its name and can ask it to guide workouts, log meals, or help shop for products you're looking at. The glasses are also getting an FDA-cleared hearing enhancement feature for adults with mild to moderate hearing loss. The post doesn't specify a launch date or which models will get Muse.

Why it matters: Putting Muse on glasses is a key step in Meta's push to move AI assistants from phones to wearables, with three concrete use cases. But the post doesn't give a launch date or supported models, so the score sits right at the featured threshold.

Sep 23Wednesday

AI HOT (Curated Pool)

Xiaomi releases open-source MiMo-V2.6 Pro and Flash multimodal models; Pro matches Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks

Xiaomi open-sourced two multimodal models: MiMo-V2.6 Pro and Flash. Pro scored 46 on the Artificial Analysis Intelligence Index—the highest among open-source models—and matches Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks. The post doesn't disclose parameter counts, training cost, inference latency, or the exact open-source license, so I'd hold off on production assumptions for now.

Why it matters: Xiaomi open-sourced MiMo-V2.6 Pro, matching Claude Opus 5 and GPT-5.6 Sol on agent benchmarks and hitting the highest open-source score on the Intelligence Index. Domestic flagship model release gets full weight per policy. Missing parameter count is a gap, but the signal is s...

AI HOT (Curated Pool)

Qwen-Image-2.1 Tops Arena's Open-Source Leaderboards for Image Editing and Text-to-Image

Alibaba's Qwen-Image-2.1 ranks first among open-source models on Arena's Image Edit and Text-to-Image leaderboards. It scored 1367 in Image Edit Arena, placing 16th overall, just 3 points behind GPT-Image-1.5-high-fidelity at #15. The post doesn't disclose parameter count, architecture, or release timeline.

Why it matters: Qwen-Image-2.1 hitting #1 open-source on Arena's image editing leaderboard, just 3 points behind GPT-Image-1.5, is a concrete cross-model signal. Score stays at 78 rather than higher because the post doesn't disclose parameter count, architecture, or release timeline — the inf...

Sep 22Tuesday

AI HOT (Curated Pool)

Qwen-Image-2.1 released as open weights, tops Image Edit Arena among open-source models

Qwen-Image-2.1 is out with open weights. It scored 1367 on the Arena Image Edit Arena, ranking #1 among open-source models and #16 overall — just 3 points behind GPT-Image-1.5-high-fidelity at #15. It also landed #1 open-source on the Text-to-Image Arena. The post doesn't disclose parameter count, architecture details, or the exact open license.

Why it matters: Qwen-Image-2.1 open weights dropped, hitting #1 open-source on Arena's image editing leaderboard at #16 overall, just 3 points behind GPT-Image-1.5. Score held back because the post doesn't disclose parameter count, architecture, or license — we're grading on the leaderboard n...

Latent Space

Xiaomi MiMo-V2.6-Pro tops open weights leaderboard, trained for $3M

Xiaomi released the MiMo-V2.6 series. The Pro version ranks #1 among open weights models on the Artificial Analysis Intelligence Index with a score of 46, at a training cost of $3M. A Flash variant targets efficiency, and an UltraSpeed variant offers 20x faster output. The technical report details RL scaling across three axes: larger batches and throughput, richer multi-task environments, and more grader compute. Code and training recipes are open-sourced, but the 7k+ task datasets are not yet released. Former DeepSeek engineer Fuli Luo, now at Xiaomi, previously live-streamed the training runs.

Why it matters: Xiaomi's MiMo-V2.6-Pro hit #1 on the Artificial Analysis open-weights leaderboard with a $3M training budget — price-performance right at the frontier. Flash and UltraSpeed variants cover efficiency and speed use cases, and the tech report details an async RL architecture. Not...

AI HOT (Curated Pool)

Xiaomi MiMo-V2.6-Pro hits ~10th on Code Arena WebDev, ~3rd among open-weight models

Xiaomi released two omni-modal models: MiMo-V2.6-Pro and Flash. The Pro version scored 1628 on Code Arena WebDev, up 153 points from MiMo-V2.5-Pro's 1475, landing around 10th overall and ~3rd among open-weight models under MIT license. The post doesn't disclose Flash's benchmark numbers or parameter counts.

Why it matters: Xiaomi's multimodal model hits ~10th on Code Arena WebDev and ~3rd among MIT open-weight models, with a 153-point gain for Pro. Flags a domestic flagship release with concrete benchmark data. Flash variant lacks params and scores, capping it below 80.

Sep 21Monday

AI HOT (Curated Pool)

Qwen-Image-2.1: A 7B Single-Checkpoint Model for Both Image Generation and Editing

Qwen-Image-2.1 is a 7B native image generation and editing model. It uses a single checkpoint for both tasks and supports up to 10 reference images. The model includes a built-in prompt-enhancement LLM, integrates with diffusers and ComfyUI, and offers a no-install browser demo on Hugging Face Spaces. The post doesn't disclose training data, inference latency, or benchmark comparisons.

Why it matters: Qwen drops an image model with a 7B single-checkpoint design for both generation and editing, plus a built-in prompt optimizer — a fresh combo. Score stays at 78 rather than 85+ because the post doesn't disclose training data, inference speed, or real image-quality comparisons...

Sep 20Sunday

AI HOT (Curated Pool)

Qwen-Image-2.1 now works with ComfyUI, open weights available

Alibaba Qwen released open weights for Qwen-Image-2.1 with native ComfyUI support. A single 7B checkpoint handles both image generation and editing, outputs up to 2K natively, accepts up to 10 reference images per instruction, and supports RGBA with alpha channel. The post doesn't spell out license terms or hardware requirements.

Why it matters: Alibaba Qwen drops a 7B unified generation/editing model with native ComfyUI support, 2K output, and RGBA transparency — a direct win for the local image-gen community. Held below 84 because hardware requirements and license terms aren't disclosed, so real-world adoption is st...

Hacker News front page

StepFun launches Step 5 Preview, a 600B MoE flagship model targeting coding and finance

StepFun introduces Step 5 Preview, a 600B-parameter MoE model with 27B active per token, a 1M-token context window, and vision support. It scores 67.7 on DeepSWE v1.1, ahead of Kimi K3 and GLM-5.3 but behind GPT-6 Astra and Claude Opus 5. On the in-house StepCodeBench it hits 49.0, again leading domestic models and trailing the two US labs. On FrontierFinance it reaches 66.4, second only to Claude Opus 5. Artificial Analysis gives it an intelligence index of 44; StepFun claims substantially lower cost per task at comparable intelligence. The post does not disclose API pricing, release timeline, or training details.

Why it matters: StepFun's Step 5 Preview is a 600B MoE model that edges out Kimi K3 and GLM-5.3 on coding benchmarks but still trails GPT-6 Astra and Claude Opus 5 by 6-7 points. Scored 78 because it's a substantive domestic model push in agentic coding with real numbers, but not industry-sha...

Sep 18Friday

AI HOT (Curated Pool)

OpenRouter tested 20 image gen models: cheapest at $0.006, priciest at $0.134

OpenRouter sent the same prompt to 20 image models and read the actual billed cost. GPT Image 2 was cheapest at $0.006 per 1024×1024 PNG; Gemini 3 Pro Image was priciest at $0.134—a 22x spread. Pricing units differ across providers (tokens, megapixels, per image), so side-by-side list prices mislead; generate once and check usage.cost. Five of six models rendered text correctly, including the cheapest. Recraft V4.1 Vector outputs editable SVG at $0.08. The post also details formats, resolution caps, and seed support per model.

Why it matters: OpenRouter ran one prompt through 20 image models and posted the actual bills — the kind of real cost data pricing pages never show. All three HKR axes hit: the headline pulls you in, the billing breakdown is genuinely new info, and it nails a daily pain point for builders. No...

Sep 15Tuesday

Computing Life · Share · Yage

Runway demos code-free UI rendering; Google lets agents write their own manuals; GitHub Enterprise goes air-gapped

All three are early-stage. Runway Solaris generates interactive UIs frame-by-frame with no frontend code—only curated demos and a waitlist so far, no public testing, pricing, or API. Google WikiSkill distills agent failure logs into reusable skill manuals, lifting Gemini-3.5-Flash accuracy from 49.5% to 68.1%, but skills from a small model can hurt a larger one; no official code repo. GitHub GHES 3.22 lets enterprises self-host Copilot CLI inside air-gapped networks with admin-managed model endpoints, though many features are disabled and it's labeled a technical preview.

Why it matters: Three items bundled, with Solaris as the main hook. Runway's frame-by-frame interface rendering is genuinely novel, but there's only a curated demo and waitlist — no public access, no third-party testing, and the cost comparison dodges standard web rendering. That keeps it bel...

Sep 14Monday

Hacker News front page

30 SVG prompts benchmark 2025–2026 LLMs on pelican-bicycle-style drawing tests

Tom Gally built a site with Claude Fable 5.1 that extends Simon Willison's pelican-riding-a-bicycle test into 30 SVG drawing prompts. The 2026 run covers six models—GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen3.8 Max, and Fugu Ultra v2—while the 2025 run includes ten models like Claude Sonnet 4.5 and GPT-5.1. Each image shows generation time and cost: DeepSeek V4 Pro finished in 1 min 36 s at $0.10, Qwen3.8 Max took over 12 minutes, and Fugu Ultra v2 cost $1.02. The post presents raw SVG outputs without subjective ratings, so you compare the drawings directly.

Why it matters: Simon Willison's pelican test is a community staple, and this expands it to 30 prompts across 6 models with timing and cost — dense, useful signal. The deliberate lack of subjective scoring means readers have to flip through images themselves, which costs it a bit of immediate...

Sep 12Saturday

Latent Space

DeepSeek V4.1-Flash: a 763B encoder-decoder MoE with 8B prefill, 16B decode, and native vision

DeepSeek dropped V4.1-Flash on Sep 10. Despite the 4.1 label, Sebastian Raschka called it a V5-level rewrite. It's a 763B total-parameter MoE with a causal encoder-decoder split: 8B active for prefill, 16B for decode, yielding 1–2% sparsity and up to 8× smaller KV cache vs V4 Flash. Native vision is built in, and V4 Pro has been quietly retired. The post doesn't include benchmark tables but argues current evals miss the point—the real advance is context efficiency for long-running agents.

Why it matters: DeepSeek drops V4.1-Flash with a 763B causal encoder-decoder MoE, 8B/16B active params, 1%-2% sparsity, and vision. Sebastian Raschka says it should've been V5. This is a major domestic flagship architecture update with a cross-source cluster forming. HKR all hit. Not 90+ yet ...

AI HOT (Curated Pool)

DeepSeek V4.1-Flash open-sourced: CED architecture cuts prefill cost for coding agents

DeepSeek released open weights for V4.1-Flash, a 552B MoE model with a Causal Encoder-Decoder architecture tuned for coding agents. It splits compute asymmetrically: 8B active params during prefill, 16B during decode, plus improved KV cache efficiency. On Terminal Bench 2.1 it hits 90.6; on Automation-Bench it scores 54.8—better than V4-Pro but still failing roughly half of complex workflows, so keep a human in the loop. It is also DeepSeek's first non-experimental model with native image input. Chartography reaches 78.9, but ZeroBench logical reasoning over images is only 49. DeepSeek has already retired V4-Flash traffic and will reroute V4-Pro traffic to V4.1-Flash starting September 14.

Why it matters: DeepSeek open-sourced V4.1-Flash, a 552B MoE that splits prefill and decode via CED architecture, directly targeting coding agent latency. Terminal Bench 2.1 scores are concrete, and Baseten's analysis adds deployment perspective. Not 85+ because this is a third-party writeup ...