Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

421–440 of 514

Apr 17Friday

MIT Technology Review · AI

How robots learn: A brief, contemporary history

Companies and investors put $6.1 billion into humanoid robots in 2025, 4x 2024, and MIT Technology Review attributes the surge to a shift in how robots learn. The piece highlights two mechanisms: around 2015, simulation plus reward signals enabled millions of trial-and-error runs; after ChatGPT in 2022, robotics models took images, sensors, and joint states to predict dozens of motor commands per second. The key change is data-driven learning over hand-written rules; the provided text is truncated, so later examples are not fully disclosed.

Why it matters: HKR-H/K/R all pass: the $6.1B and 4x funding jump provide the hook, and the piece maps the shift from sim+RL to multimodal action models. It stays in the lower featured band because this is commentary rather than a new release, and the excerpt is truncated on company-level detail

X · @dotey

Seedance 2.0 API is now available on Volcano Engine and BytePlus

Volcano Engine has released the Seedance 2.0 API for enterprises, individual developers, and overseas users via BytePlus; China pricing is RMB 46 per million tokens, or about RMB 1 per second for pure video generation. The post says it supports text, image, audio, and video inputs, plus face verification, portrait authorization, and 10,000+ preset avatars for workflow automation; overseas pricing is not disclosed here. The part to watch is orchestration: the post cites up to 10x efficiency gains, but does not disclose a common benchmark or model specs.

Why it matters: HKR-H/K/R all pass: the overseas rollout is a real hook, and the post includes usable pricing and modality details for practitioners. It stays at 74 because this is an API availability update, not a major model launch, and the post does not disclose model params, benchmark method

X · @op7418

Seedance 2.0 API is now fully open

Volcano Engine has opened the Seedance 2.0 API to domestic users, while BytePlus serves overseas access; the API currently accepts 4 input modalities: text, image, audio, and video. The post also confirms face registration, portrait authorization, and preset virtual avatars, but does not disclose pricing, rate limits, model variants, or regional availability. The real watchpoint is whether video-agent workflows can be wired through Skills and MCP, not the ecosystem rhetoric.

Why it matters: This is a real product update from ByteDance’s stack: HKR-H on full API availability, HKR-K on 4-modal input and consent mechanics, and HKR-R on builder demand for deployable video APIs. I keep it at 75 because pricing, rate limits, regional rollout details, and quality evidence

Latent Space

[AINews] Anthropic Claude Opus 4.7 - one step better than 4.6 in every dimension

Anthropic launched Claude Opus 4.7 at the same $5/$25 per million input/output tokens; the post says 4.7-low through 4.7-high each outperform the matching higher 4.6 tiers. Reported changes include a new xhigh reasoning tier, Claude Code defaulting to xhigh, an 11-point gain on SWE-Bench Pro, and image input up to 2,576 px on the long edge (~3.75 MP). Do not overread the tokenizer change: the same input can use up to 35% more tokens, but the post says total token use still falls by up to 50% from prior equivalents.

Why it matters: Anthropic's flagship-model release fits the policy's 85–94 band. HKR-H/K/R all pass because the post gives concrete pricing, benchmark, image-limit, and token-accounting changes that hit Claude users' core coding and cost concerns.

Hacker News front page

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

Simon Willison ran a 20.9GB quantized Qwen3.6-35B-A3B on a MacBook Pro M5 and judged its SVG pelican output better than Claude Opus 4.7. He used LM Studio with an Unsloth Q4_K_S GGUF, then repeated the test with “a flamingo riding a unicycle” and again scored Qwen higher. This is not a general capability result; the author says this joke benchmark no longer tracks overall model usefulness in this comparison.

Why it matters: A named first-person experiment with reproducible setup gives this strong HKR-H/K/R: the headline has a sharp contrast, the post includes a 20.9GB GGUF on an M5 MacBook Pro via LM Studio, and it hits the open-local-vs-closed-frontier debate. It stays in featured, not higher, لأن/

X · @dotey

browser-use open-sources video-use, a Claude Code skill that turns raw camera footage into edited videos

browser-use released video-use, a Claude Code skill that turns raw footage into a final.mp4 automatically. It converts footage into ElevenLabs word-level timestamp transcripts, shrinking one asset to about 12KB; the post says feeding frames directly would cost about 45 million tokens. The key detail is the structured editing pipeline: the model mostly reads text, uses timeline images only at uncertain cuts, and runs up to 3 self-check repair passes after rendering.

Why it matters: Strong HKR-H/K/R: the result is instantly clickable, and the post includes a concrete text-first editing architecture with 12KB vs about 45M-token economics. Kept below higher bands because this is a builder-facing Claude Code skill, not a platform-level release.

Apr 16Thursday

X · @dotey

Anthropic officially releases Claude Opus 4.7 at unchanged pricing

Anthropic released Claude Opus 4.7 at unchanged pricing: $5 per million input tokens and $25 per million output tokens; the API name is claude-opus-4-7, now live across Claude, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. The post gives two concrete changes: vision input now supports up to 2576 pixels on the long edge, and the new tokenizer can raise token usage to 1.0-1.35x for the same text. Watch migration cost, not list price; higher reasoning settings and multi-turn runs can increase output length and bills.

Why it matters: An Anthropic substantive model release belongs in the 85+ band, and this is not just a rename: the 2576px vision limit and 1.0–1.35x tokenization shift affect migration tests and billing immediately. HKR-H/K/R all pass, so it clears p1.

X · @op7418

Anthropic releases Claude Opus 4.7 with the following main updates

Anthropic has rolled out Claude Opus 4.7 across all Claude products and the API, with pricing unchanged from Opus 4.6. The post lists better long-horizon task handling, more precise instruction following, self-verification before reporting, vision support up to 2,576-pixel long-edge images, plus Claude Code Ultra Review, an xhigh thinking level, and auto-approval for Max users.

Why it matters: This is a substantive Anthropic model release across Claude and the API, with testable details: unchanged pricing, a 2,576px vision limit, self-checking outputs, and Claude Code workflow changes. HKR-H/K/R all pass; it fits the same-day must-write band, so p1.

Hacker News front page

Introducing Claude Opus 4.7

Anthropic released Claude Opus 4.7 on Apr. 16 at the same price as Opus 4.6: $5 per million input tokens and $25 per million output tokens. The post says it improves on Opus 4.6 in advanced software engineering, long-running tasks, and higher-resolution vision, and ships across Claude, the API, Amazon Bedrock, Vertex AI, and Microsoft Foundry. The key detail is the first deployment of Anthropic’s cyber request blocking on a less capable model; the post cites benchmark gains but does not fully disclose every score in text.

Why it matters: Anthropic shipping Claude Opus 4.7 is a same-day write: GA, unchanged $5/$25 pricing, and rollout across Claude, API, Bedrock, Vertex AI, and Foundry give it direct workflow impact. HKR-H/K/R all pass, but the post does not publish full benchmark scores.

Hacker News front page

Qwen3.6-35B-A3B: Agentic coding power, now open to all

Qwen released Qwen3.6-35B-A3B as open weights, with 35B total parameters and 3B active parameters. The post reports 73.4 on SWE-bench Verified, 51.5 on Terminal-Bench 2.0, and 92.0 on RefCOCO. The key point is agentic coding and multimodal performance at a 3B active-parameter budget, with weights, Qwen Studio, and API access available.

r/LocalLLaMA

Qwen3.6-35B-A3B released

Qwen released Qwen3.6-35B-A3B as open source under Apache 2.0; it is a sparse MoE with 35B total parameters and 3B active. The post also claims agentic coding, strong multimodal perception and reasoning, plus thinking and non-thinking modes; the post does not disclose benchmarks, context length, or latency.

Why it matters: HKR-H/K/R all pass: a new open Qwen model is timely, and the post confirms 35B total, 3B active, and Apache 2.0. The score stays at 82 because this is still a launch post; benchmarks, context window, latency, and multimodal details are not disclosed here.

Hacker News front page

Darkbloom – Private inference on idle Macs

Eigen Labs launched Darkbloom, linking 100M+ Apple Silicon Macs into a decentralized inference network. It offers an OpenAI-compatible API, claims end-to-end encryption plus hardware attestation, and lists prices up to 70% below OpenRouter comps. The real point is the trust model: hardware keys, hardened runtime, and signed outputs are disclosed, but enterprise audit scope still needs the paper.

Why it matters: HKR-H/K/R all pass: the idle-Mac inference angle is novel, and the post includes concrete scale, API, encryption, and price claims. I keep it at 80 because this is still a self-published research preview; audit scope, network reliability, and attack boundaries are not yet third-p

TechCrunch · AI

Google rolls out a native Gemini app for Mac

Google launched a native Gemini app for Mac on April 15 for all users worldwide on macOS 15 and later, with Option + Space as the summon shortcut. Users can share their screen or local files with Gemini, and the app also supports image generation with Nano Banana and video generation with Veo. The key shift is desktop access plus live context sharing, not just another client.

Why it matters: Google shipping a native Gemini app for Mac clears HKR-H/K/R: the hook is desktop entry, the new facts are hotkey and context sharing, and the resonance is the desktop assistant race. Still a mid-weight product update, not a model leap, so it sits at the low end of featured.

Apr 14Tuesday

最佳拍档 (BestPartners)

Global GPU shortage worsens: H100 rental prices rose nearly 40% in five months

SemiAnalysis says Nvidia H100 one-year rental pricing rose from $1.70 to $2.35 per GPU-hour between Oct 2025 and Mar 2026, up nearly 40% in five months. The post attributes this to Anthropic-driven demand, multi-agent and media generation workloads, and memory cost spikes, with LPDDR5 and DDR5 contract prices up about 4x and 5x year over year; much new capacity is already prebooked. The key variable is the supply gap, not Blackwell refreshes alone.

Why it matters: Strong HKR-H/K/R: the story has a sharp price-shock hook, concrete market data, and clear resonance with compute-cost anxiety. It stays below P1 because this is a secondary video synthesis of a SemiAnalysis report, not a primary company or product announcement.

Apr 11Saturday

QbitAI · WeChat

OpenClaw-style methods reach multimodal generation, with a 6B model beating Nano Banana 2 on some tasks

A team led by Shanghai AI Laboratory introduced GEMS, adding Agent Loop, Memory, and Skills to multimodal generation, and reports that 6B Z-Image-Turbo beats Nano Banana 2 on some tasks. The post reports +14.22 average gains on 5 mainstream tasks and +8.92 over the best baseline on 4 downstream tasks; the paper and code are public, but the post does not disclose Nano Banana 2's full setup.

Why it matters: Strong HKR-H/K/R: the hook is a 6B multimodal model beating Nano Banana 2, and the post includes mechanism plus testable deltas (+14.22 / +8.92) with paper and code. It stays below P1 because the article does not disclose the full Nano Banana 2 comparison setup.

QbitAI · WeChat

Liu Zhuang and Danqi Chen team open-source Vero, a general visual reasoning RL framework, reaching SOTA with zero thinking data

Princeton researchers including Liu Zhuang and Danqi Chen open-sourced Vero, an RL framework for visual reasoning, and report beating Qwen3-VL-8B-Thinking on 23 of 30 benchmarks. The post says Vero uses 600K samples filtered from 59 datasets, task-routed rewards, and single-stage RL across six task groups. The key point is the mechanism mix: no private thinking data, but the post does not disclose training cost or base model configuration.

Why it matters: Featured on HKR-H/K/R: the zero-thinking-data claim is a strong hook, and the post includes concrete benchmark and method details. I keep it in the low 80s because training cost, base model choice, and full reproduction conditions are not disclosed.

QbitAI · WeChat

A Chinese embodied model reached global No.1 as a 100,000-hour human dataset for robots was released

Psibot says it released a 100,889-hour human-plus-robot manipulation dataset, and that Psi-R2 ranked first on AllenAI’s MolmoSpace benchmark. The post lists 95,472 hours of human data, 5,417 hours of robot data, 1,000 open-sourced hours, 294 scenes, 4,821 tasks, and 1,382 objects; Psi-W0 adds 30% failure samples, and Psi-R2 latency drops from 2.2s to under 100ms. The key point is the data loop and benchmark framing: the post claims nearly 10x higher success, but does not disclose task setup, full baselines, or statistics.

Why it matters: HKR-H/K/R all pass: the data scale, failure-sample mix, and latency cut are concrete and discussable. I keep it at 80 because the No.1 ranking and near-10x success claim lack task setup, full baselines, and statistical detail in the body.

Apr 10Friday

QbitAI · WeChat

Tencent open-sources 3B SVG model HiVG to make tokens geometry-aware

Tencent Hunyuan open-sourced the 3B-parameter HiVG, claiming 62.7%-63.8% shorter SVG sequences via hierarchical tokenization and better SVG generation metrics than GPT-5.2, Claude-4.5-Sonnet, and some 8B open models. The post reports 0.896 SSIM, 0.114 LPIPS, and 0.957 CLIP-S on Image-to-SVG; the core method packs drawing commands plus coordinates into segment tokens and uses HMN to initialize coordinate embeddings. The part to watch is token design, not parameter count; paper, code, and project page are public.

Why it matters: Tencent's HiVG earns HKR-H and HKR-K: a 3B open model claims GPT/Claude-level SVG results, and the article includes 62.7%-63.8% token compression plus SSIM 0.896, LPIPS 0.114, and CLIP-S 0.957. HKR-R is weaker because SVG generation remains niche, so it lands at the low end of `f

Apr 9Thursday

QbitAI · WeChat

Beyond MoE, Tencent introduces MoT: a 2B embodied model ranks first in 16 of 22 evaluations

Tencent Hunyuan and Robotics X released HY-Embodied-0.5; its MoT-2B uses 4B total params with 2B active and ranks first in 16 of 22 embodied evaluations. The post says it uses 100M+ embodied data, 600B+ pretraining tokens, 30M+ mid-training samples, plus visual latent tokens, bidirectional attention, RFT, RL, and online distillation. The key point is a rebuilt edge-oriented embodied stack, not a simple VLM fine-tune.

Why it matters: Strong on HKR-H/K/R: the headline has a real hook, the body includes concrete numbers and training mechanisms, and the edge-robotics angle lands with practitioners. I keep it at 83, not 85+, because this is a high-quality embodied-model release, not a broad same-day industry-def

X · @op7418

Meta releases Muse Spark model

Meta released the Muse Spark model with native multimodal reasoning, tool use, visual chain-of-thought, and multi-agent orchestration, but it is only available in the Meta AI app and is not open source for now. The snippet says its Contemplating mode coordinates multiple parallel agents for reasoning, and its Artificial Analysis score is below Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. The post does not disclose model size, pricing, or rollout timing.

Why it matters: A major-lab model launch plus the “poached team’s first output” angle lands HKR-H/K/R. The score stays near the featured floor because the post offers capability claims and relative benchmark placement only; params, pricing, rollout timing, and access scope are not disclosed.