Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

381–400 of 514

Apr 27Monday

QbitAI · WeChat

Meshy tops 10M users and moves into 3D printing as ARR rises 14x

Meshy says it passed 10M registered users, reached $40M ARR, and grew 2025 revenue 14x year over year. Meshy Creative Lab supports keychain, magnet, and keycap design; physical ordering is not live yet. The key signal is print fit: 97% slice-pass rate in Bambu Studio across 75 tested models.

Why it matters: HKR-H/K/R all pass: the hook, revenue metrics, and print-readiness test are concrete. This is a vertical 3D AI product update from company disclosure, so it lands at the lower featured band.

Synced · WeChat

Apple paper asks: What do your logits know?

Apple researchers posted an arXiv paper testing whether VLM top-k logits leak image details. Using CLEVR, MSCOCO, and probes, 30–80 logits recover noise, target traits, and some background attributes. The key risk is gray-box APIs exposing top-k log probabilities.

Why it matters: HKR-H/K/R all pass: the Apple paper turns VLM logit outputs into a concrete privacy risk, with CLEVR/MSCOCO probes and a 30–80 logit range. It is strong research, not a same-day platform event, so it stays in 78–84.

Synced · WeChat

From 99 Lines of Frozen Code to Meshy AI’s 3D Momentum in the West

Meshy AI released Meshy 6 and claims over 60% share in developed Western markets. The post says it has 10M+ users, $40M+ ARR, and 100M+ AI-generated 3D models in three years. The key signal is workflow fit: 37Games reports 30–40% less base sculpting work.

Why it matters: HKR-H/K/R pass: Meshy 6 has a clear founder/product hook, concrete traction metrics, and a production-labor angle. Kept in the low featured band because the market-share claim is company-sourced and no independent benchmark is disclosed.

Apr 25Saturday

Hacker News front page

Google Flow Music

Google Flow Music launched a web creation entry with six sections: songs, playlists, Spaces, videos, projects, and Turntable. The page says Producer creates full songs with Lyria 3, and AI music videos use Veo. Pricing, regions, model specs, and rights terms are not disclosed.

Why it matters: HKR-H/K/R pass: a Google AI music web product tying Lyria 3 and Veo is clickable, concrete, and competitive. Score stays in 72–77 because price, regions, rights, and model specs are not disclosed.

TechCrunch · AI

ComfyUI hits $500M valuation as creators seek more control over AI-generated media

ComfyUI raised $30 million at a $500 million valuation. The RSS snippet says its tools give creators more control over AI image, video, and audio generation; the post does not disclose investors, round stage, pricing, or release timing. The real signal is workflow control, not another model vendor.

Why it matters: TechCrunch reports a $30M raise at a $500M valuation, making controllable media workflows a real market signal rather than a hobbyist niche. HKR-H/K/R all pass, but missing investors, round stage, pricing, and roadmap keep it in the low featured band.

Hacker News front page

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Google DeepMind released TIPSv2 with 3 pretraining changes for CVPR 2026. iBOT++ applies self-distillation to masked and visible patches, adding 14.1 mIoU on ADE150; Head-only EMA cuts training parameters by 42%. The key signal is visible-token supervision, not a larger teacher model.

Why it matters: HKR-K is strong: ADE150 gains 14.1 mIoU and trainable params drop 42%. HKR-H/R pass, but this is still a VLM research release, not a same-day model launch.

The Verge · AI

How Project Maven taught the military to love AI

In the first 24 hours of the assault on Iran, the US military struck more than 1,000 targets, with targeting accelerated by AI systems including Maven Smart System. The snippet says this was nearly 2x the scale of Iraq's “shock and awe” attack over 20 years ago, and Katrina Manson's new book traces Project Maven from its 2017 start in computer vision for drone footage; the post does not disclose model details, later contractors, or current deployment scope.

Why it matters: HKR-H/K/R all pass: the angle is military AI adoption at strike scale, with a concrete number (1,000+ targets in 24 hours) and a named system. Kept at 74 because the piece does not disclose current models, vendor changes, or deployment scope.

Apr 24Friday

Synced · WeChat

Remember more, answer faster, use less: HERMES speeds real-time streaming video understanding by 10x

Fudan University, Shanghai Academy of AI for Science, and NUS proposed HERMES, a training-free framework that turns KV cache into hierarchical memory for streaming video understanding and cuts TTFT by up to 10x. The post lists three mechanisms: hierarchical cache management, cross-layer memory smoothing, and position re-indexing; it reports 68% fewer video tokens with comparable or better results, and Qwen2.5-VL-7B on StreamingBench rising from 73.31% to 79.44%. What matters for practitioners: it answers without external retrieval, with TTFT around 27/29/28 ms at 16/64/256 frames.

Why it matters: Strong HKR-H/K/R: the 10x speed claim is a real hook, and the article includes concrete mechanisms and numbers, including 68% fewer video tokens and 27-29 ms TTFT. It stays below major product-news bands because this is an academic research release, not a market-moving launch.

Synced · WeChat

After robots beat humans in marathon times: hardware nears its limit, intelligence becomes the second half

Honor's humanoid robot Lightning ran 50:26 at the 2026 Beijing Yizhuang half marathon, faster than the men's human world record of 57:20; the post also says Unitree H1 did a 1.9 km winding course in 4:13. The post cites nearly 200 embodied-AI financings and over RMB 30 billion in Q1 2026, plus Spirit AI's $455 million Pre-A on April 16. The real signal is capital shifting from robot hardware to model-centric 'brains.'

Why it matters: Strong HKR-H/K/R: the human-vs-robot race result is a real hook, and the piece adds concrete funding numbers plus a clear thesis on value shifting from hardware to intelligence. It remains secondary commentary rather than a primary product, research, or company release, so it is

Xinzhiyuan · WeChat

Google's Vision Banana aims to unify vision tasks with a single pixel-generation interface

Google DeepMind and collaborators including Kaiming He introduced Vision Banana, claiming one pixel-generation interface can cover detection, segmentation, generation, and editing. The RSS snippet gives two head-to-head numbers versus Nano Banana Pro: 53.5% human win rate on GenAI-Bench and 47.8% on ImgEdit; it says only a small amount of reversible-format task data was mixed in, while data scale and full benchmark tables are not disclosed in the post.

Why it matters: HKR-H/K/R all pass: the story is a unified pixel-output interface spanning detection, segmentation, generation, and editing, with 53.5% and 47.8% benchmark figures. It stays in the 78-84 band because training scale and full benchmark coverage are not disclosed.

Hacker News front page

GPT-5.5: Mythos-Like Hacking, Open to All

XBOW says GPT-5.5 cut miss rate to 10% on its real-vulnerability benchmark, versus 40% for GPT-5 and 18% for Opus 4.6. It scored 97.5% on visual acuity and used about half the login iterations of the next-best model. The key point is black-box testing: GPT-5.5 without source beat GPT-5 with source.

Why it matters: HKR-H/K/R all pass: a major OpenAI model claim, concrete security benchmark numbers, and a clear practitioner safety nerve. The source is XBOW rather than an OpenAI launch post, so it stays below 95.

Apr 23Thursday

QbitAI · WeChat

Qwen3.6-27B open-weights, beats its 397B flagship predecessor on agentic coding

Qwen released Qwen3.6-27B and says it beats Qwen3.5-397B on 4 agentic coding benchmarks with about 1/15 the parameters. The post cites SkillsBench rising from 30.0 to 48.2, GPQA Diamond at 87.8, and AIME26 at 94.1; it uses a dense architecture, Thinking Preservation, and Gated DeltaNet, with weights on Hugging Face and ModelScope.

Why it matters: This is a substantive Qwen open-source model release with concrete agent-coding and reasoning scores, so HKR-H/K/R all pass. I keep it at 84, not higher, because the post gives strong benchmarks but no pricing, context window, or independent reproduction yet.

Xinzhiyuan · WeChat

Tashi Zhihang raises $455.0 million in a Pre-A round, with Sequoia China and Hillhouse jointly leading

Tashi Zhihang said on April 16 it closed a $455.0 million Pre-A round led by Sequoia China, Hillhouse Ventures, and Meituan, which the post says set China records for embodied AI single-round and Pre-A financing. The post also says its AWE3.0 four-modal model lifted unseen-view task success by 3x and cut execution jitter by about 45%, and that its A1 robot set a Guinness record in sub-millimeter wire-harness assembly within one hour. What matters is whether model, data, and deployment keep reproducing; the post does not disclose valuation or deal terms.

Why it matters: HKR-H/K/R all pass: the round size and investor mix are compelling, and the post includes concrete model and robot metrics. I keep it at 83, not P1, because key facts remain company-supplied; valuation, deal terms, and third-party validation are not disclosed.

Hacker News front page

Website streamed live directly from a model

Flipbook generates an entire clickable website in real time with an image model, where each page is a pixel image and every click spawns a deeper image. The post says all on-screen text is drawn by the image model with no HTML or text overlays, and content comes from agentic web search plus model knowledge. The key point is the interaction model, not a standard generative UI; the live video stream remains an experimental, resource-heavy toggle.

Why it matters: HKR-H/K/R all pass: a live, clickable site rendered entirely as model-generated pixels is a strong hook, and the post explains the mechanism (no HTML/text overlay, agentic web search). Kept at 76 because latency, cost, model stack, and usage are not disclosed.

Apr 22Wednesday

Hacker News front page

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Qwen released the open-weight 27B dense model Qwen3.6-27B and made it available in Qwen Studio. It scores 77.2 on SWE-bench Verified vs. 76.2 for Qwen3.5-397B-A17B, and 59.3 on Terminal-Bench 2.0 under a 256K context and 3-hour timeout. The real takeaway is deployment: this is not a larger MoE, but a denser 27B model with stronger coding results.

Why it matters: Qwen3.6-27B is a substantive flagship-model release with open weights, concrete coding benchmarks, and a practical dense-deployment angle. HKR-H/K/R all pass, and per policy a major Chinese model launch should score on par with an equivalent US-lab release.

QbitAI · WeChat

SenseAuto's Sage with 3B active params claims to beat GPT-5.4 and Opus 4.6 in cars

SenseAuto released Sage, an in-car multimodal edge model with 32B total params and 3B active params, and says it scored 94% on PinchBench, above Claude Opus 4.6 at 93.3% and GPT-5.4 at 90.5%. The post says Sage runs on Nvidia OrinX with about 0.5s TTFT, 0.03s TPOT, and 80 tok/s throughput; its SCOUT training method cuts GPU hours by about 60%, and ERL raises complex-task completion by 20%. The key point is not the headline race but whether a 3B-active model can sustain multi-step tool use on device.

Why it matters: HKR-H/K/R all pass: the 3B-active-vs-GPT hook is strong, and the post gives concrete OrinX latency, throughput, and benchmark numbers. I keep it at 79 because the evidence is self-reported and the impact is narrower than a general model launch.

Latent Space

OpenAI launches GPT-Image-2

OpenAI shipped GPT-Image-2 in ChatGPT, Codex, and the API. It has thinking and non-thinking variants, with stronger text, layout, editing, multilingual output, and QR codes. Arena ranks it first on 3 Image Arena boards, with 1512 Elo in text-to-image and a +242 lead.

Why it matters: OpenAI shipped GPT-Image-2 across ChatGPT/API/Codex with Arena #1 claims and 1512 T2I Elo. HKR-H/K/R all pass, so this lands in the 85–94 same-day band.

X · @dotey

OpenAI launches ChatGPT Images 2.0, available to all ChatGPT and Codex users starting today

OpenAI made ChatGPT Images 2.0 available today to all ChatGPT and Codex users, and also opened the gpt-image-2 API. The RSS snippet says it supports up to 2K output, aspect ratios from 3:1 to 1:3, and more reliable non-English text rendering. In thinking mode, it can search the web, generate multiple styles, and self-check outputs; that tier is limited to Plus, Pro, and Business, with Enterprise not yet available.

Why it matters: This is a substantive OpenAI product update: ChatGPT Images 2.0 rolls into ChatGPT, Codex, and the gpt-image-2 API, with concrete facts on resolution, aspect ratios, and thinking-mode limits. HKR-H/K/R all pass, but the source is a short repost-style summary and omits pricing and

X · @OpenAI

Introducing ChatGPT Images 2.0

OpenAI introduced ChatGPT Images 2.0 as an image model for complex visual tasks and directly usable visuals. The RSS snippet cites sharper editing, richer layouts, and “thinking-level intelligence,” but the post does not disclose model size, pricing, latency, or rollout scope.

Why it matters: OpenAI’s official post makes this a source-authoritative product update, and the “Images 2.0” framing gives it HKR-H plus HKR-R. I kept it near the featured floor because the post lacks model details, pricing, latency, benchmarks, and rollout scope, so HKR-K fails.

Bloomberg Technology

OpenAI unveils new image model that is better at charts and diagrams

OpenAI released an update to its image generation software to produce more accurate, complex charts and scientific diagrams. The RSS snippet does not disclose the model name, launch timing, pricing, benchmarks, or technical method. The real signal is a push into professional use cases, not generic image quality.

Why it matters: Bloomberg gives this a source-authority tiebreak: OpenAI is targeting a high-value weakness in image generation, so HKR-H and HKR-R pass. HKR-K misses because the snippet lacks the model name, rollout, price, benchmarks, and mechanism, keeping it at the featured floor.