Skip to content

#多模态

0 today

Jun 5Friday

QbitAI · WeChat

Yao Shunyu Responds to Whether Tencent Is Behind in AI

Yao Shunyu said at Tencent Cloud’s AI industry application conference that Hunyuan 3 rebuilt pretraining and reinforcement-learning infrastructure, changed data and evaluation, and assigned its strongest post-training staff to improve Yuanbao first; he named coding agents, multimodality, and embodied AI as Tencent’s next focus areas.

Why it matters: HKR-H/K/R all pass, but the facts are conference remarks and roadmap signals, not a new model release with specs, benchmarks, or launch date. This fits the lower featured band for a major Chinese tech AI strategy update.

May 29Friday

AI HOT (Curated Pool)

Google DeepMind CEO Demis Hassabis Says AGI Could Arrive Within Three Years

Demis Hassabis predicts AGI could arrive around 2029 to 2030, with mature multimodal capabilities and autonomous decision-making as key conditions, while warning that society remains underprepared and needs rules and safeguards before deployment.

Why it matters: HKR-H/K/R all pass: Hassabis gives a 2029-2030 AGI window and names multimodal plus autonomous decision-making as conditions. High-interest commentary, but thinner than a model release or major product update.

May 22Friday

AI HOT (Curated Pool)

Plastic Interfaces: The Future Shape of AI-Driven Software

Salesforce has adopted a headless architecture that lets salespeople update data through AI; the post says MCPs, HTML, audio, and web interfaces can be generated dynamically by context, but it does not disclose implementation metrics or adoption numbers.

Why it matters: HKR-H/K/R all pass, but this is a software-form thesis without user metrics, launch timing, or a reproducible test. It fits the insightful-commentary band, not a must-write release.

May 18Monday

Latent Space

The Autonomous Drone Tech Stack and Economics of Drones — Yaroslav Azhnyuk

Latent Space interviewed The Fourth Law founder Yaroslav Azhnyuk for a two-hour episode covering FPV drones, five levels of autonomy, eight dimensions of the autonomous battlefield, and China’s manufacturing advantage; the transcript claims Ukraine produced 4 million FPV drones last year and discusses a hypothetical Chinese capacity of 4 billion.

Why it matters: HKR-H/K/R all pass: the Latent Space interview offers concrete autonomy and battlefield frameworks. It is still commentary, not a model release, product update, or research artifact, so it stays just above the featured threshold.

May 17Sunday

Financial Times · Technology

Chinese AI Groups Pull Ahead of US Rivals in Video Generation Race

FT says Chinese AI groups have moved ahead of US rivals in video generation; the RSS snippet names ByteDance and Kuaishou and says they outshine western competitors in advertising and entertainment quality, but the post does not disclose benchmark metrics or model details.

Why it matters: FT authority plus a China-vs-US video-generation lead claim clears HKR-H and HKR-R. HKR-K fails because the body lacks metrics, samples, and eval method, so it sits at the low featured threshold.

Synced · WeChat

What Are World Models? Their History and the $10 Billion Bet

Jiqizhixin translated a MoE Capital blog tracing two world-model lineages. The article says more than $10 billion entered the category over 18 months, and cites DreamDojo as using 44,711 hours of first-person video pretraining to reach r=0.995 correlation with real-world robot policy outcomes.

Why it matters: HKR-H/K/R all pass: the hook is strong and the article gives concrete figures, but it is a compiled explainer rather than a new release. It fits the featured-threshold band for a strong commentary/tutorial.

May 12Tuesday

AI HOT (Curated Pool)

The Evolution of Human-Computer Interfaces: From Text to Interactive Neural Video

Karpathy argues that LLM output is moving from Markdown toward richer HTML, while interactive neural video still has an open problem: how to combine neural generation with precise traditional software.

Why it matters: HKR-H/K/R pass: Karpathy gives a fresh UI frame, a concrete Markdown→HTML→neural-video path, and a builder-facing product question. Single X post with no data keeps it at the featured floor.

May 10Sunday

Synced · WeChat

Ted Xiao Reviews Three Eras of Robot Learning, from RT-1/RT-2 to Scaling

Ted Xiao divides nearly a decade of robot learning into three eras: Google’s team trained RT-1 on 87,000 teleoperation trajectories, then adapted 5B to 55B VLMs into VLA policies for RT-2.

Why it matters: HKR-H/K/R all pass: a named Google robotics insider, concrete RT-1/RT-2 numbers, and strong embodied-AI resonance. It is retrospective commentary, not a launch, so it stays in the 72–77 featured band.

Apr 25Saturday

The Verge · AI

How Project Maven taught the military to love AI

In the first 24 hours of the assault on Iran, the US military struck more than 1,000 targets, with targeting accelerated by AI systems including Maven Smart System. The snippet says this was nearly 2x the scale of Iraq's “shock and awe” attack over 20 years ago, and Katrina Manson's new book traces Project Maven from its 2017 start in computer vision for drone footage; the post does not disclose model details, later contractors, or current deployment scope.

Why it matters: HKR-H/K/R all pass: the angle is military AI adoption at strike scale, with a concrete number (1,000+ targets in 24 hours) and a named system. Kept at 74 because the piece does not disclose current models, vendor changes, or deployment scope.

Apr 24Friday

Synced · WeChat

After robots beat humans in marathon times: hardware nears its limit, intelligence becomes the second half

Honor's humanoid robot Lightning ran 50:26 at the 2026 Beijing Yizhuang half marathon, faster than the men's human world record of 57:20; the post also says Unitree H1 did a 1.9 km winding course in 4:13. The post cites nearly 200 embodied-AI financings and over RMB 30 billion in Q1 2026, plus Spirit AI's $455 million Pre-A on April 16. The real signal is capital shifting from robot hardware to model-centric 'brains.'

Why it matters: Strong HKR-H/K/R: the human-vs-robot race result is a real hook, and the piece adds concrete funding numbers plus a clear thesis on value shifting from hardware to intelligence. It remains secondary commentary rather than a primary product, research, or company release, so it is

Apr 21Tuesday

Hacker News front page

Expansion Artifacts

Matt Ström-Awn argues that flaws in LLM outputs are “expansion artifacts,” not compression artifacts, and cites 2024 evidence that they can be tracked. He notes Stanford researchers estimated AI-drafted text in 17.5% of recent CS papers and 16.9% of peer reviews from post-ChatGPT word-frequency shifts, and contrasts this with a JPG after 10,000 recompressions reaching PSNR 14.59. The point for practitioners is forensic: these artifacts expose both model aesthetics and generation provenance.

Why it matters: HKR-H lands on the “expansion artifacts” hook; HKR-K adds concrete numbers and a testable provenance claim; HKR-R hits peer-review trust and detection anxiety. It stays at 73 because this is personal-blog commentary, not a primary research or product release event.

Apr 19Sunday

QbitAI · WeChat

Did Musk Really Sell Lao Gan Ma on Douyin?

QbitAI says the shown “Musk selling Lao Gan Ma on Douyin” and “GTA-6 crossover” images were generated by OpenAI GPT Image 2; the claimed 100K+ live viewers were part of fake visuals. The post argues Image 2 can render realistic posters, game screenshots, and readable long text, and links that to Codex-style UI workflows; the post does not disclose pricing, rollout scope, or launch timing. The real issue is verification: image realism is eroding “photo as evidence.”

Why it matters: HKR-H/K/R all pass: the hook is novel, the article shows a concrete capability jump, and the trust/verification angle resonates with practitioners. It stops short of p1 because the body does not disclose rollout, pricing, or an official launch scope.

Apr 17Friday

MIT Technology Review · AI

How robots learn: A brief, contemporary history

Companies and investors put $6.1 billion into humanoid robots in 2025, 4x 2024, and MIT Technology Review attributes the surge to a shift in how robots learn. The piece highlights two mechanisms: around 2015, simulation plus reward signals enabled millions of trial-and-error runs; after ChatGPT in 2022, robotics models took images, sensors, and joint states to predict dozens of motor commands per second. The key change is data-driven learning over hand-written rules; the provided text is truncated, so later examples are not fully disclosed.

Why it matters: HKR-H/K/R all pass: the $6.1B and 4x funding jump provide the hook, and the piece maps the shift from sim+RL to multimodal action models. It stays in the lower featured band because this is commentary rather than a new release, and the excerpt is truncated on company-level detail

Apr 14Tuesday

最佳拍档 (BestPartners)

Global GPU shortage worsens: H100 rental prices rose nearly 40% in five months

SemiAnalysis says Nvidia H100 one-year rental pricing rose from $1.70 to $2.35 per GPU-hour between Oct 2025 and Mar 2026, up nearly 40% in five months. The post attributes this to Anthropic-driven demand, multi-agent and media generation workloads, and memory cost spikes, with LPDDR5 and DDR5 contract prices up about 4x and 5x year over year; much new capacity is already prebooked. The key variable is the supply gap, not Blackwell refreshes alone.

Why it matters: Strong HKR-H/K/R: the story has a sharp price-shock hook, concrete market data, and clear resonance with compute-cost anxiety. It stays below P1 because this is a secondary video synthesis of a SemiAnalysis report, not a primary company or product announcement.

Apr 3Friday

X · @dotey

LatePost on DeepSeek before V4: traits, organization, and Liang Wenfeng's goals

LatePost says DeepSeek has confirmed 4 core departures, and V4's large model slipped from around Lunar New Year to April; the report says it will likely remain open source. The snippet cites 2x-3x recruiting offers, some 8-digit packages, a 100-plus research team, and a shift from CUDA/Triton to TileLang for domestic GPU adaptation. The real signal is strategy: DeepSeek had spent less on agents and coding, but now names an agent product role; the post does not disclose V4's size, price, or benchmarks.

Why it matters: This is not the V4 launch, but it carries real signal: four confirmed departures, an April delay, a 100+ research team, and partial migration from CUDA/Triton to TileLang. HKR-H/K/R all pass; missing V4 specs, price, and benchmarks keeps it below launch-tier or p1.

Apr 1Wednesday

MIT Technology Review · AI

The gig workers who are training humanoid robots at home

Micro1 hires thousands of contractors across 50+ countries to film chores at home with iPhones and sell that real-world data to humanoid robotics companies. The piece cites $15/hour pay for one worker, says robotics firms spend over $100 million a year on such data, and notes $6 billion+ went into humanoids in 2025. The real issue is data governance: workers know the footage trains robots, but the post shows they often do not know how it is stored, shared, or deleted.

Why it matters: This clears HKR-H/K/R: at-home chore videos are a strong hook, and the piece adds numbers on scale, pay, and spend. The sharper industry signal is the hidden data pipeline and weak governance on storage, sharing, and deletion, so it merits featured, not p1.

TheValley101 (硅谷101)

E231 | From B2B to A2A: What Agent Infrastructure Could Do for a One-Person Global Business

Alibaba International president Zhang Kuo said procurement agent product Accio reached 10 million MAU in March and is still growing quickly month over month. The interview’s clearest metric: AI cuts procurement communication time to one-fifth, from about one week to one day, by chaining research, design-pack generation, cross-language communication, and supplier screening into an agent workflow. The real point is A2A: the post frames it as agents restructuring buyer, seller, and platform flows, not just a better chat box.

Why it matters: This is not a major launch, but it is a primary-source exec interview with concrete numbers: 10M MAU and a 1 week→1 day cycle cut. HKR-H/K/R all pass, yet the event is still below a model release or major product update, so it lands in featured, not p1.

Mar 17Tuesday

MIT Technology Review · AI

Where OpenAI’s technology could show up in Iran

Just over two weeks after OpenAI’s classified-use deal with the Pentagon, MIT Technology Review outlined three places its tech could surface in Iran-related conflict. The post names target prioritization, Anduril counter-drone analysis, and GenAI.mil back-office use; it does not disclose when classified integration will finish or confirm deployment in Iran.

Why it matters: MIT Technology Review maps OpenAI’s classified-defense deal to 3 Iran-linked scenarios, giving it strong HKR-H and HKR-R. HKR-K is weaker because the piece does not confirm deployment, integration timing, or system limits, so it lands at the featured floor.

Mar 9Monday

MIT Technology Review · AI

How AI Is Turning the Iran Conflict Into Theater

The author reviewed more than a dozen Iran-war dashboards in one week and argues they turn satellite data, ship tracking, AI summaries, and betting links into a real-time war spectator interface. The post cites a dashboard built by two Andreessen Horowitz staffers that pulls in Kalshi bets, while Craig Silverman has logged 20 similar dashboards. The point to watch is information quality: the piece cites Financial Times reporting on AI-generated satellite images spreading online, while these dashboards lack the human vetting and historical context used by intelligence agencies.

Why it matters: HKR-H lands on the war-dashboard-plus-betting hook; HKR-K lands on the named examples, counts, and Kalshi mechanism; HKR-R lands on reliability and ethics nerves for AI builders. Strong reported commentary, but not a product, model, or research milestone, so it ranks as featured,

Feb 26Thursday

New York Times Chinese

Where Is the U.S. Losing to China in AI?

The piece argues China has embedded AI into manufacturing, with 30,000+ smart factories, and over half of all industrial robots installed globally in 2024 going to Chinese plants. It cites shop-floor data: Zeekr's Ningbo plant uses 800+ robots, Xiaomi says its Beijing factory produces one car every 76 seconds, while only 18% of U.S. manufacturers report a formal AI strategy and two-thirds struggle to scale pilots. The real point is not frontier models but AI deployment in factory automation, scheduling, and inspection.

Why it matters: Data-backed commentary with all three HKR axes: a strong US-vs-China hook, concrete factory metrics, and direct resonance on AI deployment and competitiveness. Not a new product, model, or research release, so it stays in the low featured band.

Feb 12Thursday

MIT Technology Review · AI

AI is already making online crimes easier. It could get much worse.

Microsoft said it blocked $4 billion in scams and fraudulent transactions in the year to April 2025, with many likely aided by AI-generated content. The article cites research estimating at least half of spam email is now LLM-generated, and LLM use in targeted email attacks rose from 7.6% in April 2024 to 14% in April 2025. Don’t overread “fully automated AI hackers”: the immediate issue is AI scaling phishing, deepfakes, and malware support, while the post does not disclose total attack growth.

Why it matters: HKR-H/K/R all pass: the swindle angle is strong, and the article adds concrete abuse metrics ($4B blocked, half of spam, 7.6%→14%). Featured, not p1, because this is a solid trend report on AI-enabled fraud, not a same-day industry-moving release or incident.

Feb 3Tuesday

MIT Technology Review · AI

What We’ve Been Getting Wrong About AI’s Truth Crisis

MIT Technology Review says the US Department of Homeland Security has confirmed using Google and Adobe AI video generators for public-facing content, reported last Thursday. The post cites two failure points: Adobe auto-labels only fully AI-made content, mixed edits are opt-in, and X can remove or hide labels. The key issue is influence after exposure: a new Communications Psychology paper found participants still used a fake confession deepfake to judge guilt even after being told it was fake.

Why it matters: This is not zero-sourcing commentary: it ties confirmed DHS usage to concrete labeling gaps at Adobe and X, then adds a named study showing disclosure did not reset judgment. HKR-H/K/R all pass, but it is still commentary plus one study, not a same-day industry-moving event.

Jan 16Friday

Ruan YiFeng's Weblog

Technology Enthusiast Weekly (Issue 381): What China's AI Foundation Model Leaders Are Thinking

Ruan Yifeng’s Issue 381 excerpts talks from Beijing’s AGI-Next summit on Jan 10, covering views from Zhipu, Alibaba Qwen, and Tencent AI leaders on China’s model roadmap. The post cites Lin Junyang saying US compute is 1-2 orders of magnitude larger, Yao Shunyu calling the odds of a China-led top AI company in 3-5 years high, while Lin puts it at 20%. The key split is strategic: Tang Jie points to RLVR in 2025, Lin bets on multimodal foundation agents, and Yao says B2B buyers pay a $200/month premium for stronger models.

Why it matters: It clears all three HKR axes: public strategic disagreement gives it a strong hook, and the post includes concrete numbers and testable claims. The score stops short of the high bands because this is a secondary synthesis of summit remarks, not a primary release or original scoop

Jan 13Tuesday

MIT Technology Review · AI

CES showed me why Chinese tech companies feel so optimistic

CES 2026 drew 148,000+ attendees and 4,100+ exhibitors, with Chinese companies making up nearly a quarter and standing out in AI hardware and robotics. The post ties their optimism to manufacturing-led iteration speed, not one breakthrough; Lenovo Qira, Nvidia Vera Rubin, and AMD Helios show the race is shifting to cloud and hybrid AI.

Why it matters: This is on-the-ground CES reporting with a competition thesis: Chinese optimism comes from manufacturing and supply-chain iteration, supported by 148k attendees, 4,100 exhibitors, and roughly one-quarter from China. HKR-H/K/R pass, but shipment, revenue, and order data are not in