Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

441–460 of 514

Apr 8Wednesday

QbitAI · WeChat

Xiaomi unveils two AI audio frameworks: Any2Speech and Midasheng-audio-generate

Xiaomi's large-model application team introduced Xiaomi Any2Speech and Midasheng-audio-generate. Any2Speech generates up to about 10 minutes per inference, while the other model turns one text prompt into mixed audio with speech, music, and ambient sound. The post names GST labeling, dual-path planning with dimension dropout, Flow Matching, and five-field structured labels; benchmark scores, training scale, and commercial terms are not disclosed.

Why it matters: Xiaomi released two audio-generation frameworks with a clear hook and concrete mechanisms, so HKR-H and HKR-K pass. HKR-R is weaker because benchmark results, training data scale, open-source status, and commercial terms are not disclosed, so this sits at the low end of featured.

QbitAI · WeChat

After a late-night update, DeepSeek reportedly said: I am V4?

DeepSeek added Fast and Expert modes on its web app and started gray-testing a Vision model; the claim that Expert mode is V4 comes only from user probes and the model’s own replies. The post gives one concrete detail: Expert mode focuses on code, web, and harder generation tasks, is supply-limited, does not support multimodal or file upload, and one user reported a length cap at about 133K tokens. What matters is the official model ID and context spec; the post does not disclose them, pricing, or a release timeline.

Why it matters: HKR-H is strong on the 'I am V4' hook. HKR-K and HKR-R pass because the post gives testable mode behavior and a ~133K token limit, and DeepSeek silent swaps are highly discussable. The score stays in the mid-70s because the model name, price, and context window remain unconfirmed

Apr 7Tuesday

Latent Space

[AINews] Gemma 4 crosses 2 million downloads

Google’s Gemma 4 reached about 2 million downloads in its first week. The post compares that with Gemma 3 at 6.7 million over the past year, Gemma 2 at 1.4 million since June 2024, and Qwen 3.5 at about 27 million in roughly 1.5 months. The signal for practitioners is local deployment: one iPhone 17 Pro demo ran Gemma 4 E2B at about 40 tok/s via MLX, with support across Hugging Face, vLLM, llama.cpp, Ollama, and NVIDIA.

Why it matters: HKR-H/K/R all pass: the story has a clean hook, concrete comparative download data, and a real open-model adoption nerve. It stays low-featured because this is a secondary-source uptake snapshot, not a primary Google release or a substantive capability update.

Apr 3Friday

X · @op7418

Alibaba released the Qwen 3.6 Plus model

Alibaba released Qwen 3.6 Plus with a 1M context window, 64K input, and nearly 991K max output. The RSS snippet says it improves over Qwen 3.5 on agents, coding, image, and document understanding, priced at RMB 2 per 1M input tokens and RMB 12 per 1M output tokens; benchmark scores and test conditions are not disclosed.

Why it matters: Alibaba shipping Qwen 3.6 Plus is a substantive domestic model update. HKR-H/K/R all pass on the 1M-context plus pricing combo, but it stays below P1 because benchmark scores, baselines, and test conditions are not disclosed in the body.

X · @op7418

Google releases Gemma 4 for on-device use under Apache 2.0

Google released Gemma 4 in four variants—E2B, E4B, 26B MoE, and 31B Dense—targeting phones, edge devices, and up to single-H100 workstations. The RSS snippet says the 26B MoE activates 3.8B parameters and adds native function calling, JSON output, multimodal I/O, speech-to-text, and Apache 2.0 licensing; the post does not disclose benchmarks, context length, or rollout details.

Why it matters: Google releasing Gemma 4 is a substantive open-model update. HKR-H/K/R all pass on the size spread, 3.8B-active MoE detail, and deployment-cost relevance; it stays at 81 because benchmarks, context window, and test conditions are not disclosed here.

X · @dotey

LatePost on DeepSeek before V4: traits, organization, and Liang Wenfeng's goals

LatePost says DeepSeek has confirmed 4 core departures, and V4's large model slipped from around Lunar New Year to April; the report says it will likely remain open source. The snippet cites 2x-3x recruiting offers, some 8-digit packages, a 100-plus research team, and a shift from CUDA/Triton to TileLang for domestic GPU adaptation. The real signal is strategy: DeepSeek had spent less on agents and coding, but now names an agent product role; the post does not disclose V4's size, price, or benchmarks.

Why it matters: This is not the V4 launch, but it carries real signal: four confirmed departures, an April delay, a 100+ research team, and partial migration from CUDA/Triton to TileLang. HKR-H/K/R all pass; missing V4 specs, price, and benchmarks keeps it below launch-tier or p1.

X · @dotey

Google releases the Gemma 4 open model family under Apache 2.0

Google released the Gemma 4 family and switched the full line to Apache 2.0. The post says it includes 31B Dense, 26B MoE, E4B, and E2B; 31B and 26B support 256K context, and 31B fits on one 80GB H100. The key change is distribution terms: fewer limits on commercial use, modification, and redistribution, plus native function calling and structured JSON for agent workflows.

Why it matters: This is a substantive Google model release, with the Apache 2.0 switch carrying as much weight as the model specs. HKR-H/K/R all pass on novelty, concrete deploy details, and commercial relevance; it stays below P1 because the post lacks formal eval links and direct head-to-heads

Google DeepMind

Google DeepMind releases the Gemma 4 open model family

Google DeepMind released Gemma 4, which it calls its most intelligent open model yet, aimed at advanced reasoning and agentic workflows under an Apache 2.0 license. The family comes in four sizes: E2B, E4B, 26B MoE and 31B Dense. The 31B ranks 3rd among open models on the Arena AI text leaderboard, and the 26B ranks 6th.

Why it matters: Gemma 4 is Apache 2.0 and spans four sizes from on-device to workstation, so you can weigh deployment and fine-tuning options for open models.

Apr 1Wednesday

MIT Technology Review · AI

The gig workers who are training humanoid robots at home

Micro1 hires thousands of contractors across 50+ countries to film chores at home with iPhones and sell that real-world data to humanoid robotics companies. The piece cites $15/hour pay for one worker, says robotics firms spend over $100 million a year on such data, and notes $6 billion+ went into humanoids in 2025. The real issue is data governance: workers know the footage trains robots, but the post shows they often do not know how it is stored, shared, or deleted.

Why it matters: This clears HKR-H/K/R: at-home chore videos are a strong hook, and the piece adds numbers on scale, pay, and spend. The sharper industry signal is the hidden data pipeline and weak governance on storage, sharing, and deletion, so it merits featured, not p1.

TheValley101 (硅谷101)

E231 | From B2B to A2A: What Agent Infrastructure Could Do for a One-Person Global Business

Alibaba International president Zhang Kuo said procurement agent product Accio reached 10 million MAU in March and is still growing quickly month over month. The interview’s clearest metric: AI cuts procurement communication time to one-fifth, from about one week to one day, by chaining research, design-pack generation, cross-language communication, and supplier screening into an agent workflow. The real point is A2A: the post frames it as agents restructuring buyer, seller, and platform flows, not just a better chat box.

Why it matters: This is not a major launch, but it is a primary-source exec interview with concrete numbers: 10M MAU and a 1 week→1 day cycle cut. HKR-H/K/R all pass, yet the event is still below a model release or major product update, so it lands in featured, not p1.

Mar 24Tuesday

Mistral AI

Mistral AI releases Voxtral TTS, a 4B-parameter speech model

Mistral AI released Voxtral TTS, its first text-to-speech model. It has 4B parameters and supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It handles emotional expression and zero-shot cross-lingual voice adaptation.

Why it matters: The 4B size, nine languages, 70ms latency and pricing give readers a basis for judging cost and model choice in enterprise voice agents.

Mar 18Wednesday

MIT Technology Review · AI

The Pentagon plans to let AI companies train models on classified data, defense official says

The Pentagon is discussing secure facilities where AI firms can train military-specific models on classified data. The post says training would follow tests on nonclassified data; the DoD keeps data ownership, and company staff would access it only rarely with clearance. The key issue is leakage: one shared model may resurface classified information across groups with different access levels.

Why it matters: HKR-H lands on the unusual classified-data-training angle; HKR-K lands on concrete guardrails and ownership terms; HKR-R lands on defense procurement and leakage risk. Score stays below 85 because this is a planning-stage report, not a signed program, budget, or deployment.

Mar 17Tuesday

OpenAI News

Introducing GPT-5.4 mini and nano

OpenAI released GPT-5.4 mini and nano on March 17, 2026 for coding and subagents; mini runs over 2x faster than GPT-5 mini. In the API, mini has a 400k context window and costs $0.75/$4.50 per 1M input/output tokens, while nano is API-only at $0.20/$1.25. The key signal is performance per latency: mini scores 54.4% on SWE-Bench Pro versus GPT-5.4 at 57.7%.

Why it matters: This is an official OpenAI model launch, not a routine patch. It includes concrete numbers—>2x speed, 400k context, API pricing, and 54.4% vs 57.7% on SWE-Bench Pro—so HKR-H/K/R all pass; scored at the low end of the 85–94 band.

Mistral AI

Mistral releases Mistral Small 4, unifying reasoning, multimodal and coding

Mistral AI released Mistral Small 4, the first Mistral model to unify Magistral reasoning, Pixtral multimodal and Devstral coding-agent abilities in a single model. It ships under the Apache 2.0 license.

Why it matters: Merging reasoning, multimodal and coding agents into one open model is a direct test of what unified models do to deployment cost.

MIT Technology Review · AI

Where OpenAI’s technology could show up in Iran

Just over two weeks after OpenAI’s classified-use deal with the Pentagon, MIT Technology Review outlined three places its tech could surface in Iran-related conflict. The post names target prioritization, Anduril counter-drone analysis, and GenAI.mil back-office use; it does not disclose when classified integration will finish or confirm deployment in Iran.

Why it matters: MIT Technology Review maps OpenAI’s classified-defense deal to 3 Iran-linked scenarios, giving it strong HKR-H and HKR-R. HKR-K is weaker because the piece does not confirm deployment, integration timing, or system limits, so it lands at the featured floor.

Mar 13Friday

MIT Technology Review · AI

A defense official reveals how AI chatbots could be used for targeting decisions

A US defense official said the Pentagon can feed target lists into generative AI, have the model rank them using factors like aircraft location, and send strike recommendations for human review. The post says this chatbot layer may sit on top of Maven to speed search and analysis, but it does not disclose the speed gain, and the official did not confirm current operational use. The key issue is verification: chat outputs are easier to use than Maven’s map UI but harder to check.

Why it matters: Full HKR: the headline's hook is a chatbot in target ranking, and the body gives a concrete workflow tied to Maven plus human review. I keep it at 80, not higher, because the official describes a possible use case; speed gains and combat deployment are not confirmed.

Mar 9Monday

MIT Technology Review · AI

How AI Is Turning the Iran Conflict Into Theater

The author reviewed more than a dozen Iran-war dashboards in one week and argues they turn satellite data, ship tracking, AI summaries, and betting links into a real-time war spectator interface. The post cites a dashboard built by two Andreessen Horowitz staffers that pulls in Kalshi bets, while Craig Silverman has logged 20 similar dashboards. The point to watch is information quality: the piece cites Financial Times reporting on AI-generated satellite images spreading online, while these dashboards lack the human vetting and historical context used by intelligence agencies.

Why it matters: HKR-H lands on the war-dashboard-plus-betting hook; HKR-K lands on the named examples, counts, and Kalshi mechanism; HKR-R lands on reliability and ethics nerves for AI builders. Strong reported commentary, but not a product, model, or research milestone, so it ranks as featured,

Mar 5Thursday

36Kr (direct RSS)

Embodied AI company Pascini raises over RMB 1 billion in Series B, valuation tops RMB 10 billion

Pascini said it closed a Series B round of over RMB 1 billion, bringing its valuation above RMB 10 billion. Lead investors include Huangpujiang Capital, Kaitai Capital, and Xinan Capital; the post says Pascini will use 10-billion-scale real-world multimodal data to train its VTLA model, but does not disclose model details.

Why it matters: HKR-H/K/R all pass: the round size and valuation are the hook, and the brief gives concrete funding and data numbers. Score stays at 76 because this is a single-company funding flash; model capability, customers, and deployment progress are not disclosed.

Feb 28Saturday

36Kr (direct RSS)

Qwen plans AI glasses, earbuds, and rings as tech giants race for a new AI entry point

A report says Alibaba's Qwen plans AI glasses, earbuds, and rings for a global launch in 2026; the glasses are slated for MWC 2026, with reservations opening on March 2. The post adds that Qwen app functions like food delivery and ride hailing will move to these devices, and cites Qwen3.5-Plus with 60% lower memory use, up to 19x inference throughput, and RMB 0.8 per million tokens. The real point is distribution: if the hardware connects Alipay, Amap, and Taobao, Alibaba is chasing the consumer AI entry layer, not just device sales.

Why it matters: This is a distribution-entry story for Alibaba/Qwen, not a routine accessory refresh. HKR-H/K/R all pass: the multi-device bet is a strong hook, the report includes launch timing and model economics, and it hits the ecosystem-front-end nerve; but it is still a media exclusive, so

Feb 27Friday

36Kr (direct RSS)

Embodied AI startup Zhongke Diwuji, which supplies the "brain" for Unitree, raised hundreds of millions of yuan

Zhongke Diwuji completed Pre-A and Pre-A+ rounds worth hundreds of millions of yuan within one month, and became a Unitree core ecosystem partner in Jan 2026. Since 2025, it has supplied the "brain" model for Unitree robots; the company says its FAM models use secondary pretraining and heatmap alignment to learn new tasks from 3-5 real-robot demos, with 97% success on basic tasks. The signal to watch is commercialization: it is moving from POC to power inspection, industrial handling, and retail deployments, charging robot OEMs per-device license.

Why it matters: Embodied AI plus a Unitree supplier angle gives HKR-H and HKR-R. The story adds company-reported facts—3-5 real-robot demos, 97% base-task success, per-robot licensing—so HKR-K passes; it stays below 85 because the funding size is vague and no third-party replication is disclosed