Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

461–480 of 514

Feb 26Thursday

New York Times Chinese

Where Is the U.S. Losing to China in AI?

The piece argues China has embedded AI into manufacturing, with 30,000+ smart factories, and over half of all industrial robots installed globally in 2024 going to Chinese plants. It cites shop-floor data: Zeekr's Ningbo plant uses 800+ robots, Xiaomi says its Beijing factory produces one car every 76 seconds, while only 18% of U.S. manufacturers report a formal AI strategy and two-thirds struggle to scale pilots. The real point is not frontier models but AI deployment in factory automation, scheduling, and inspection.

Why it matters: Data-backed commentary with all three HKR axes: a strong US-vs-China hook, concrete factory metrics, and direct resonance on AI deployment and competitiveness. Not a new product, model, or research release, so it stays in the low featured band.

Feb 14Saturday

MIT Technology Review · AI

ALS stole this musician’s voice. AI let him sing again.

Patrick Darling, 32, returned to the stage on February 11 in London after two years without singing, using an AI voice clone rebuilt from old recordings. The post says speech cloning typically needs about 10 minutes of clean audio; his singing clone was built from noisy phone clips and kitchen recordings, then refined with Eleven Music over about six weeks. The practical signal is access, not sentiment: ElevenLabs offers the tools free to people who lost their voices to ALS and similar conditions, but the post does not disclose model details.

Why it matters: HKR-H/K/R all land: the hook is strong, the story gives concrete reproducible details, and the use case hits accessibility plus voice-rights nerves. Still, this is a strong application story, not a major model, product, or research release, so it stays in low featured.

Feb 13Friday

OpenAI News

Beyond rate limits: scaling access to Codex and Sora

OpenAI says in the headline it will scale access to Codex and Sora beyond current rate limits. The body is empty and does not disclose quota changes, eligible users, pricing, or rollout timing. The key missing fact is the access mechanism, not the headline claim.

Why it matters: This is an official OpenAI product update, so HKR-H and HKR-R pass: the rate-limit angle is clickable and quota pain resonates with users. HKR-K fails because the body discloses no quota delta, eligible tiers, pricing, or rollout date, so it stays at the featured floor.

Feb 12Thursday

MIT Technology Review · AI

AI is already making online crimes easier. It could get much worse.

Microsoft said it blocked $4 billion in scams and fraudulent transactions in the year to April 2025, with many likely aided by AI-generated content. The article cites research estimating at least half of spam email is now LLM-generated, and LLM use in targeted email attacks rose from 7.6% in April 2024 to 14% in April 2025. Don’t overread “fully automated AI hackers”: the immediate issue is AI scaling phishing, deepfakes, and malware support, while the post does not disclose total attack growth.

Why it matters: HKR-H/K/R all pass: the swindle angle is strong, and the article adds concrete abuse metrics ($4B blocked, half of spam, 7.6%→14%). Featured, not p1, because this is a solid trend report on AI-enabled fraud, not a same-day industry-moving release or incident.

Feb 10Tuesday

36Kr (direct RSS)

Alibaba Qwen launches the new image generation foundation model Qwen-Image-2.0

Alibaba Qwen announced Qwen-Image-2.0, an image generation foundation model, and opened API invite testing on Alibaba Cloud Bailian. Developers can also try it for free in Qwen Chat; the post does not disclose model size, pricing, eval results, or a general release date.

Feb 3Tuesday

MIT Technology Review · AI

What We’ve Been Getting Wrong About AI’s Truth Crisis

MIT Technology Review says the US Department of Homeland Security has confirmed using Google and Adobe AI video generators for public-facing content, reported last Thursday. The post cites two failure points: Adobe auto-labels only fully AI-made content, mixed edits are opt-in, and X can remove or hide labels. The key issue is influence after exposure: a new Communications Psychology paper found participants still used a fake confession deepfake to judge guilt even after being told it was fake.

Why it matters: This is not zero-sourcing commentary: it ties confirmed DHS usage to concrete labeling gaps at Adobe and X, then adds a named study showing disclosure did not reset judgment. HKR-H/K/R all pass, but it is still commentary plus one study, not a same-day industry-moving event.

Jan 31Saturday

MIT Technology Review · AI

Inside the marketplace powering bespoke AI deepfakes of real women

Researchers from Stanford and Indiana University found that on Civitai, 90% of deepfake bounty requests targeted women and 86% asked for custom LoRAs between mid-2023 and late 2024. Bounties paid $0.50 to $5 and nearly 92% were fulfilled; MIT Technology Review confirmed that even after Civitai's May 2025 deepfake ban, many older requests and purchasable outputs remained live. The key point is that the platform hosts tutorials, payment rails, and matching infrastructure, not just user uploads.

Why it matters: HKR-H lands because the story turns abuse into a visible market. HKR-K lands on four concrete stats and a post-ban moderation gap; HKR-R lands on safety and governance anxiety around open image platforms. Strong featured, not p1.

Jan 30Friday

MIT Technology Review · AI

DHS is using Google and Adobe AI to make videos

A DHS document says the agency uses Google Veo 3, Google Flow, and Adobe Firefly for public-facing content, with an estimated 100 to 1,000 licenses. It also says DHS uses Microsoft Copilot Chat for drafting and summarization and Poolside for coding; the post does not disclose which specific videos used which tool. The key point for practitioners is that commercial video generators are now inside a federal public-communications workflow, while watermark retention and attribution remain unverifiable across platforms.

Bloomberg Technology

Apple Buys Israeli AI Startup Q.ai That Interprets Facial Movements

Apple has acquired Israeli AI startup Q.ai, which builds tech to read facial movements and interpret silent communication. The RSS snippet confirms the deal and focus, but the post does not disclose price, team size, or Apple integration plans. The key question is whether Apple folds this vision capability into accessibility, AirPods, or Vision products.

Why it matters: Bloomberg gives this enough source authority for the featured floor: Apple buying silent-communication vision tech lands HKR-H and HKR-R. HKR-K is weaker because the report discloses no price, team size, accuracy, or integration plan.

Jan 29Thursday

Ruan YiFeng's Weblog

Kimi’s integrated stack vs. Manus’s layered approach

Kimi released the K2.5 model and K2.5 Agent together, with an agent mode already available on its website. The post cites 1,500-step long-horizon actions, up to 100 agents in parallel, and visual coding from design files or web videos; pricing, context window, and API terms are not disclosed. The key point is product shape: not just a model launch, but a bundled model-plus-agent release.

Why it matters: HKR-H lands on the integrated release angle; HKR-K lands on the 1,500-step, 100-agent, visual-programming details; HKR-R lands on the stack-design debate. Missing price, context window, and API terms, plus a commentary source, keep it below p1.

Jan 20Tuesday

MIT Technology Review · AI

The UK government is backing AI scientists that can run their own lab experiments

UK agency ARIA selected 12 AI scientist projects from 245 proposals, doubled its planned funding, and will give each team about £500,000 for nine months. ARIA defines an AI scientist as a system that hypothesizes, runs experiments, analyzes results, and iterates; the funded projects still rely on existing tools. The key signal is reproducible lab-loop execution, not press-release heat: one cited external study reports LLM agents failed to complete a scientific workflow 3 out of 4 times.

Why it matters: HKR-H/K/R all pass: 'AI runs its own lab experiments' is a strong hook, and the piece includes 12 teams, 245 proposals, ~£500k each, a 9-month term, and a cited 75% failure rate. Important for agentic science, but this is funding for early systems, not a proven breakthrough.

Jan 16Friday

Ruan YiFeng's Weblog

Technology Enthusiast Weekly (Issue 381): What China's AI Foundation Model Leaders Are Thinking

Ruan Yifeng’s Issue 381 excerpts talks from Beijing’s AGI-Next summit on Jan 10, covering views from Zhipu, Alibaba Qwen, and Tencent AI leaders on China’s model roadmap. The post cites Lin Junyang saying US compute is 1-2 orders of magnitude larger, Yao Shunyu calling the odds of a China-led top AI company in 3-5 years high, while Lin puts it at 20%. The key split is strategic: Tang Jie points to RLVR in 2025, Lin bets on multimodal foundation agents, and Yao says B2B buyers pay a $200/month premium for stronger models.

Why it matters: It clears all three HKR axes: public strategic disagreement gives it a strong hook, and the post includes concrete numbers and testable claims. The score stops short of the high bands because this is a secondary synthesis of summit remarks, not a primary release or original scoop

Jan 13Tuesday

MIT Technology Review · AI

CES showed me why Chinese tech companies feel so optimistic

CES 2026 drew 148,000+ attendees and 4,100+ exhibitors, with Chinese companies making up nearly a quarter and standing out in AI hardware and robotics. The post ties their optimism to manufacturing-led iteration speed, not one breakthrough; Lenovo Qira, Nvidia Vera Rubin, and AMD Helios show the race is shifting to cloud and hybrid AI.

Why it matters: This is on-the-ground CES reporting with a competition thesis: Chinese optimism comes from manufacturing and supply-chain iteration, supported by 148k attendees, 4,100 exhibitors, and roughly one-quarter from China. HKR-H/K/R pass, but shipment, revenue, and order data are not in

Jan 12Monday

36Kr (direct RSS)

He Xiaopeng: The best AI companies in the future will build their own chips

He Xiaopeng said XPeng's four 2026 vehicle models will use its Turing AI chip, and Ultra SE and Ultra trims will run a second-gen VLA model for entry-level L4-assisted driving. The post says MAX uses one 750 TOPS chip, Ultra SE uses two, and Ultra uses three; XPeng has entered 60 countries and regions, and VLA 2.0 is already being road-tested in Europe. The real signal is that automakers are pulling chips, models, and deployment in-house as a ceiling-on-performance play, not just a cost move.

Why it matters: The signal is not the slogan but the concrete roadmap: 4 cars, 750 TOPS per chip, 1/2/3-chip trims, and VLA 2.0 road tests. HKR-H/K/R all pass, but this is still a roadmap disclosure rather than a shipped AI-industry event, so it sits at the low end of featured.

Jan 6Tuesday

NVIDIA Blog

NVIDIA RTX Accelerates 4K AI Video Generation on PC With LTX-2 and ComfyUI Upgrades

NVIDIA said GeForce RTX and related devices can run LTX-2 and updated ComfyUI for local AI video generation up to 3x faster with up to 60% lower VRAM use. The post attributes this to PyTorch-CUDA optimizations, native NVFP4/FP8 support in ComfyUI, and an RTX Video 4K upscaling node due next month; LTX-2 open weights are available now and the workflow ships next month. The real signal for AI builders is that local 4K video is shifting from VRAM-bound demos to usable RTX workflows.

Why it matters: HKR-H/K/R all pass: the story has a sharp hook, concrete mechanisms, and clear resonance for local-inference users. I keep it at 76 because this is a vendor-blog ecosystem optimization update, not a major model launch or broad platform shift.

NVIDIA Blog

NVIDIA unveils new open models, data and tools across agents, robotics, AVs and biomedicine

NVIDIA released open models, datasets and training tools spanning Nemotron, Cosmos, Alpamayo, Isaac GR00T and Clara, plus 10T language tokens, 500K robotics trajectories, 455K protein structures and 100TB of vehicle sensor data. Newly disclosed items include Nemotron Speech/RAG/Safety, Cosmos Reason 2, Transfer 2.5, Predict 2.5, GR00T N1.6 and Alpamayo 1; the key signal is that NVIDIA is opening the data stack across agents, physical AI, AVs and biomedicine.

Jan 4Sunday

36Kr (direct RSS)

Huawei Cloud embodied robotics lead left to start a company using brain cognition to redesign robot brains

Former Huawei Cloud embodied robotics lead Zhu Senhua left in Oct. 2025 to found Julao Panshi, which has raised a seed round worth tens of millions of RMB. The company says it uses brain-inspired methods to modify VLA for embodied AI; prototype tests showed 40% higher deployment efficiency in open environments and a 90% cut in data needs for few-shot manipulation. The key point is that it starts as a VLA add-on, while targeting Asia-Pacific service and industrial use cases where overseas customers accept robots that replace only 50%-70% of human labor.

Why it matters: A solid featured story: founder spinout + seed funding + a concrete VLA add-on thesis with +40%/-90% prototype claims. Not higher because the evidence is still company-reported; the piece does not disclose a public benchmark, customer count, or scaled deployment data.

Dec 17, 2025Wednesday

Mistral AI

Mistral releases OCR 3 with better forms and handwriting, plus Document AI Playground

Mistral released Mistral OCR 3, which wins 74% of head-to-head comparisons against Mistral OCR 2 on forms, scanned documents, complex tables and handwriting. Mistral says its accuracy beats enterprise document-processing tools and AI-native OCR options.

Why it matters: Mistral OCR 3's upgrades on forms, handwriting and complex tables, plus its $2 per 1,000 pages pricing, are useful when evaluating document parsing options.

Dec 16, 2025Tuesday

OpenAI News

The new ChatGPT Images is here

OpenAI says the new ChatGPT Images is now available, and the only confirmed fact is a product availability update. The body is empty; the post does not disclose model name, quality, pricing, quotas, or rollout scope.

Why it matters: An official OpenAI launch post makes HKR-H and HKR-R pass: a new ChatGPT image feature is a real product event people will discuss. HKR-K fails because the body here discloses no model name, pricing, quotas, rollout scope, or examples, so it stays near the featured floor.

Dec 11, 2025Thursday

OpenAI News

The Walt Disney Company and OpenAI reach agreement to bring beloved characters to Sora

The Walt Disney Company and OpenAI reached an agreement to bring Disney characters to Sora; only the title is available and the body is empty. The title confirms the parties and Sora as the target, but the post does not disclose scope, character list, launch timing, or licensing terms.

Why it matters: This official partnership clears HKR-H and HKR-R: Disney characters entering Sora is inherently clickable and will spark discussion on licensing, compliance, and video-model competition. HKR-K fails because the post discloses the deal only; scope, rollout, and terms are missing,,