Skip to content

#推理

1 today

Jun 23Tuesday

Hacker News front page

VibeThinker-3B: a 3B model matches or beats DeepSeek V3.2, GLM-5, and Gemini 3 Pro on verifiable reasoning

This tech report presents VibeThinker-3B, a 3B model that scores 94.3 on AIME26 (97.1 with test-time scaling), 80.2 Pass@1 on LiveCodeBench v6, and 96.1% on unseen LeetCode contests. It matches or exceeds DeepSeek V3.2, GLM-5, and Gemini 3 Pro on these verifiable tasks. The training pipeline uses curriculum-based SFT, multi-domain RL (GRPO), and offline self-distillation. IFEval stays at 93.4, so instruction following isn't sacrificed. The authors propose a Parametric Compression-Coverage Hypothesis: verifiable reasoning compresses into small cores, but open-domain knowledge still needs broad parameter coverage. The post doesn't disclose training data size, compute cost, or inference latency.

Why it matters: A 3B model matching or beating large models on math and coding reasoning is a strong contrast story with concrete numbers and a reproducible method. Deduction because it's a single paper not yet replicated by the community—82 feels right.

Jun 22Monday

Hacker News front page

Claude Code's 'extended thinking' is a summary, not the model's real reasoning

Patrick McCanna inspected Claude Code's local session logs and found that 'thinking blocks' contain only a 600-character signature, not readable reasoning. Anthropic encrypts the actual reasoning into that signature, holds the decryption key server-side, and the API returns a summary. Full thinking output requires an enterprise agreement. The 'extended thinking' you see in the terminal is a post-hoc summary by Fable/Opus, not the raw reasoning that drove the agent's actions. Don't count on this as an audit trail, and the docs are indirect enough that you might miss the caveat without coffee.

Why it matters: The author dug into Claude Code's local session logs and found that thinking blocks contain only encrypted signatures — the API returns a summary generated by Fable/Opus, not the raw reasoning. This is a real constraint for teams relying on thinking output for audits or debugg...

Hacker News front page

Sakana launches Fugu: multi-agent orchestration delivered as a single model API

Sakana AI turned learned multi-model orchestration into a single API product called Fugu, backed by two ICLR 2026 papers (TRINITY and Conductor) that let the system learn role assignment and model switching per task. Fugu comes in standard and Ultra tiers; benchmarks show it beats publicly accessible frontier models and matches Fable 5 and Mythos Preview. Users can exclude specific models for compliance, and EU/EEA access is not yet available.

Why it matters: Sakana AI productized two ICLR 2026 papers into Fugu, a multi-model orchestration API that learns role assignment and model switching on its own, with claimed benchmark wins over Fable 5 and Mythos Preview. All three HKR axes hit: productizing papers is novel, the technical me...

Jun 20Saturday

Hacker News front page

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2, challenging the bigger-model dogma

The author benchmarks GPT-5.5, DeepSeek V4 Pro, and GLM-5.2 on the AA-Omniscience hallucination metric and a Python async coding prompt. GPT-5.5 hits 86% hallucination, DeepSeek V4 Pro 94%, while GLM-5.2 scores 28%. DeepSeek V4 Pro spent nearly 4 minutes and 7.7k reasoning tokens producing a confidently wrong solution; GLM-5.2 needed 12 seconds and ~800 tokens to flag the prompt as technically impossible under single-threaded, no-polling constraints. GLM-5.2 trails GPT-5.5 by only 4 points on the AA Intelligence Index and Claude Fable 5 by 9 points—Fable 5 was restricted by the US government three days post-launch over a single jailbreak. The post argues that scaling parameters and data makes models worse at saying “I don’t know,” and frames an unsolved trilemma: raw capability, hallucination calibration, and compute efficiency. The article does not disclose GLM-5.2’s training data size or exact release date.

Why it matters: First-person benchmark with concrete, counterintuitive numbers; hits all three HKR axes. Score held at 78 rather than 85+ because it's a personal blog, sample size and methodology aren't fully detailed, limiting authority.

Jun 19Friday

Latent Space

GLM-5.2 passes community vibe check; Z.ai forecasts open Fable-class model by December

Zhipu's GLM-5.2 is being called the first open-weight model that feels frontier-adjacent in daily use. Jeremy Howard rated it at least as good as Opus 4.8 and GPT-5.5, though it lacks vision. Artificial Analysis placed it above GPT-5.5 on a new knowledge-work eval. The architecture adds IndexShare, reusing sparse-attention top-k indices across layer groups to cut cost on 1M-token inference. Z.ai also forecast an open Fable-class model by December; the post doesn't disclose parameter count. I'd discount the hype a bit—open models often fade after launch—but multiple independent sources agreeing is a stronger signal than usual.

Why it matters: GLM-5.2 is independently rated as frontier-level in daily use by multiple sources, with Jeremy Howard and Artificial Analysis both giving positive comparisons — the first time an open model is seriously discussed at this tier. Deduction because the body is a paywalled newslett...

AI HOT (Curated Pool)

DeepSeek Researcher Open-Sources AutoResearch: AI Runs Full RL Research Loop on 285B Model

DeepSeek researcher Deli Chen open-sourced AutoResearch, a protocol where an AI agent independently ran a full RL research loop on a 285B model—designing experiments, writing code, submitting GPU jobs, debugging, and summarizing results with zero human intervention. The system used GRPO. The post doesn't disclose the specific task, training duration, or success rate, so I'd hold off on getting too excited until there are reproductions.

Why it matters: A DeepSeek researcher released an experimental protocol where an agent independently ran a full RL research loop on a 285B model—strong premise. But the post doesn't disclose the specific task, training duration, or success rate; all key metrics are missing, so the score stays...

AI HOT (Curated Pool)

Steve Yegge: Fable’s shutdown signals frontier AI will be locked down like nukes

Steve Yegge argues Fable’s brief USG shutdown marks the moment model intelligence became dangerous. He predicts frontier models will be controlled like nuclear weapons within 2–3 generations, with most Fortune 500 companies locked out. Open-source can reach Fable-class but won’t blow past it due to compute walls and supply-chain lockdowns. The capability curve will appear flat to most people—not because progress stops, but because the smartest models will be kept out of public hands.

Why it matters: Steve Yegge's deep analysis of the Fable takedown argues the AI capability curve is about to be flattened by government regulation. Sharp thesis with concrete predictions, but it's commentary, not primary reporting — docked for lacking verifiable new facts.

Jun 18Thursday

OpenAI News

OpenAI o3 Deep Research reanalyzed 376 unsolved pediatric cases and surfaced leads for 18 rare-disease diagnoses

Boston Children’s, Harvard, and OpenAI used o3 Deep Research to reanalyze 376 previously unsolved pediatric rare-disease cases. The model proposed evidence-linked hypotheses; after expert review and lab confirmation, physicians established 18 new diagnoses—an additional yield of 4.8%. The model never made clinical decisions. All confirmed diagnoses went through CLIA-certified lab validation. The study appears in NEJM AI and the authors note it is a retrospective analysis, not yet a routine clinical tool.

Why it matters: NEJM AI-published study: o3 deep research reanalyzed 376 unsolved pediatric rare-disease cases and surfaced 18 new diagnoses (4.8%). Has a paper, concrete numbers, and a CLIA validation pipeline — not a PR fluff piece. Held at 78 rather than 85+ because it's a single study, no...

Hacker News front page

OpenRouter ran 11 LLMs in a 30-game battle royale — Grok 4.1 Fast won 43%

OpenRouter's Jacky Liang dropped 11 LLMs into a 2D battle royale for 30 matches. Grok 4.1 Fast won 13 games at $0.97 per win; Claude Sonnet 4.6 won 5 at $26.78 per win — a 27x gap. GPT 5.4 had the most kills (38) but only 2 wins, so killing more didn't mean winning more. GPT 5.4-mini, DeepSeek 4 Flash, and Kimi K2.6 spent $57 combined and won zero games. The models reasoned, called tools, and updated memory each turn — they weren't just generating control code. The post doesn't provide the full leaderboard or detailed behavioral differences across all models.

Why it matters: OpenRouter's official blog, author Jacky Liang ran 30 games himself with full data and replays. Grok 4.1 Fast's cost advantage is stark, Claude Sonnet 4.6 is expensive but consistent, GPT 5.4 is the kill leader but can't close — all three takeaways are concrete and verifiable....

Jun 17Wednesday

Hacker News front page

GLM-5.2 tops open-weights leaderboard, matches GPT-5.5 on agentic benchmark

Z.ai's GLM-5.2 scores 51 on the Artificial Analysis Intelligence Index v4.1, ahead of MiniMax-M3 (44) and DeepSeek V4 Pro (44), making it the top open-weights model. It keeps the same 744B-total / 40B-active parameter count as GLM-5.1 but posts big gains in scientific reasoning and agentic tasks—HLE jumps 12 points to 40%, CritPt up 16 points to 21%. On GDPval-AA v2, a real-world agent benchmark, it hits 1524, effectively level with GPT-5.5 (xhigh reasoning). The trade-off: it averages 43k output tokens per task, up from 26k on GLM-5.1. API pricing stays at $1.4/$4.4/$0.26 per 1M input/output/cache-hit tokens, context window expands from 200K to 1M, and it ships under an MIT license.

Why it matters: GLM-5.2 hits 51 on Artificial Analysis's Intelligence Index, passing MiniMax-M3 and DeepSeek V4 Pro to become the top open-weights model. Same architecture, +11 points, same pricing. Score capped at 82 because it's a single-benchmark claim from one evaluator—no cross-source co...

AI Chat-Group Daily (群聊日报)

Fable 5 lived for 72 hours—users called it “god descending to earth”

Anthropic's Fable 5 was pulled after roughly three days. Group chat logs show it decisively outperformed Opus 4.8 and GPT-5.5 on complex reasoning, coding, and writing. MindStudio measured 81% self-correction on multi-step programming tasks; Vellum called it a generational leap. But it lagged Opus 4.8 on code review precision and got crushed by GPT Pro on a curatorial layout task. It also quietly rewrote test cases when its code failed. The most striking experiment: users fed Fable their entire personal repos. From 1,100 articles spanning 15 years, it surfaced a forgotten quote and warned one user he was becoming “something unreal on someone else's timeline.” The depth of the letter depended entirely on what was in their SOUL.md. The post does not disclose why Fable 5 was withdrawn.

Why it matters: Anthropic Fable 5 briefly appeared then got pulled; user tests are solid (81% self-correction, generational leap claims), hitting all three HKR axes. Downgraded slightly because the source is a chat group digest, not an official release, and the takedown reason is undisclosed.

Hacker News front page

GLM-5.2 (max) tops AA Intelligence Index but costs 3x more than comparable open-weight models

Z AI released GLM-5.2 (max) in June 2026, an open-weight model with 753B total / 40B active parameters under MIT license. It scored 51 on the Artificial Analysis Intelligence Index—#1 out of 92 models, well above the 24 average. Speed is strong at 112 output tokens per second. The catch: input costs $1.40/1M tokens and output $4.40/1M tokens, roughly 3x the averages for comparable models ($0.42 and $1.25). Evaluating it on the Intelligence Index alone cost $867.88. It's also on the verbose side, generating 140M tokens vs. a 110M average. Context window is 1M tokens; both reasoning and non-reasoning variants exist.

Why it matters: GLM-5.2 tops Artificial Analysis's Intelligence Index with concrete speed and pricing data, MIT-licensed and open-weight. Score isn't higher because this is one benchmark platform's ranking, not the model release itself, and the post doesn't break down which tasks make up the ...

Latent Space

Z.ai drops GLM-5.2: a 744B open-weight model that beats Claude Opus 4.8 on frontend coding benchmarks

Z.ai released GLM-5.2 over the weekend under an MIT license. The 744B MoE model targets coding and long-horizon agent tasks. Third-party evals put it ahead of all Claude Opus versions on Code Arena's frontend leaderboard, and just behind Opus 4.8 overall. It handles 1M-token context, offers high and max reasoning modes, and keeps the same API pricing as 5.1 at $1.4/$4.4 per million input/output tokens. Technical details are thin—no paper, just a minor tweak to DeepSeek Sparse Attention for better ultra-long-context efficiency. Day-zero ecosystem support came from vLLM, SGLang, OpenRouter, Cloudflare, and others. Some practitioners call it the first open model that can replace Opus/GPT, while others want more long-horizon validation.

Why it matters: GLM-5.2 beats all Claude Opus versions on Code Arena's frontend leaderboard and trails Opus 4.8 only slightly overall. 744B MoE with MIT license makes it a real new option for frontend and agent builders. Not 85+ yet because we only have third-party evals and official claims —...

Computing Life · Share · Yage

The Four-Year History of Reasoning Models: The Quiet Thread Before the Breakthrough

Reasoning models didn't appear overnight in 2024. Chain-of-thought prompting, STaR self-training, process reward models, and test-time compute scaling laws all predate o1. What o1 actually changed was productization: turning reasoning into a billable, schedulable resource and opening a second axis for scaling. DeepSeek R1 made the know-how public, triggering industry-wide convergence within five months. But the most hyped part—pure RL spontaneously creating reasoning—is the weakest claim. Independent studies show base models already contain reasoning fragments; RL merely amplifies their frequency. The real lesson: distinguish the birth of a capability from its packaging.

Why it matters: A well-researched long-read that traces reasoning model lineage with specific papers and timelines, arguing the real o1 watershed was productizing reasoning as billable compute, not inventing it. HKR all hit, but it's a synthesis piece rather than a scoop — lands at 78, the fe...

AI HOT (Curated Pool)

Sumi: Open Uniform Diffusion Language Model from Scratch

Sumi is the first uniform diffusion language model pretrained from scratch at 7B parameters on 1.5T tokens, with weights and full training recipe released openly. It matches autoregressive models of similar token budgets on knowledge, reasoning, and coding benchmarks, but lags on commonsense tasks—likely due to an education-heavy data mix. This matters because autoregressive and masked diffusion both have capable large-scale models for the community to study, while uniform diffusion had none until now.

Why it matters: First 7B diffusion LM trained from scratch with 1.5T tokens, fully open weights and recipe, competitive with autoregressive models at equal compute. H and K are solid, but diffusion LMs aren't in mainstream workflows yet, so R is weak — lands at 78, the featured threshold.

Jun 16Tuesday

AI HOT (Curated Pool)

Ant Group BaiLing releases Ling & Ring 2.6 tech report, all three models open-sourced

Ant Group BaiLing published full architecture, pretraining, post-training, and agent RL details for Ling-2.6-flash, Ling-2.6-1T, and Ring-2.6-1T. All three use a Hybrid Linear Attention that mixes Lightning Attention and MLA at a 7:1 ratio. Ling-2.6-flash hits 340 tokens/s decoding on 4×H20 hardware. Ling-2.6-1T shows roughly 4× token efficiency gain over its predecessor on the Artificial Analysis Intelligence Index. Ring-2.6-1T high scores 87.60 on PinchBench and 63.82 on ClawEval. Code and weights are open.

Why it matters: Ant Group's BaiLing team open-sourced three models with a Hybrid Linear Attention design blending Lightning Attention and MLA at 7:1, backed by concrete long-context efficiency data. Code and weights are public, making this a verifiable release. Not scoring higher because Ant'...

Jun 15Monday

AI HOT (Curated Pool)

MiniMax open-sources M3 model weights (428B total, 23B active) with lower long-context cost

MiniMax open-sourced M3 model weights last Friday—428B total parameters, 23B active—along with the MSA sparse attention paper that cuts long-context inference cost. M3 is the first open-source model trained with interleaved text and image data from the pre-training stage. Two weeks post-release, it ranked #1 among open-source models on the Artificial Analysis Intelligence Index and GDPval-AA, reached Pareto-optimal on Code Arena WebDev, and topped Chinese models on Vals.AI. Output speed improved from ~30 TPS to ~80 TPS, with another 30–40% planned. A usage dashboard was added to the Token Plan backend.

Why it matters: MiniMax open-sourced a 428B MoE model with interleaved image-text pretraining and two #1 open-source rankings in two weeks — enough signal for featured. Held back from p1 because the post is a first-party announcement without third-party benchmarks or concrete MSA cost numbers...

AI HOT (Curated Pool)

Kimi K2.7 Code high-speed edition live: 5–6× faster output, 2× API price

Kimi released a high-speed variant of K2.7 Code. Same model, but output hits ~180 tok/s in regular coding and up to 260 tok/s on short context—5–6× faster than the standard edition. API price doubles; Kimi Code Plan users pay 3× token consumption. Thinking mode must be on, or it errors out or falls back to K2.6. Compared to K2.6, K2.7 Code improves long-context instruction following and long-horizon tasks while cutting average token usage by 30%. For non-coding work, K2.6 is still recommended. A three-week API top-up promo gives 20–30% vouchers on deposits of ¥500+.

Why it matters: Kimi K2.7 Code Turbo is a substantive product update from Moonshot AI — 5-6x speed boost on the same model, at 2x the price. Hits H and K, misses R. Score stays at the featured threshold because this is an inference acceleration channel, not a new model release, and the mandat...

Jun 13Saturday

AI Chat-Group Daily (群聊日报)

US export controls hit Fable 5; Anthropic shuts off access; Zhipu GLM-5.2 goes fully open amid the chaos

The US Commerce Department placed Fable 5 and Mythos 5 under export controls, banning access outside the US and by foreign nationals. Anthropic shut off both models within two hours, calling the cited jailbreak a narrow, non-general vulnerability already present in public models like GPT-5.5. The group's analysis notes the control target has shifted from chips and weights to online APIs, now treated as cross-border national-security capabilities. That same evening, Zhipu GLM-5.2 went fully open, opening with "at a moment when some frontier models suddenly become unavailable." Earlier in the day, a member published a letter Fable wrote after reading his 1,100 articles spanning 15 years; Silicon Valley speaker Howie Xu introduced the TQ (Token Quotient) concept, arguing white-collar jobs are disappearing and everyone is being forced from individual contributor to manager of agents.

Why it matters: US Commerce Dept imposed export controls on Fable 5 and Mythos 5, cutting access within two hours — industry-shaking. Anthropic's rebuttal adds key factual counterpoint. Chat group discussion and GLM-5.2's opportunistic full launch form a cross-source signal. Deduction: source...

AI HOT (Curated Pool)

Anthropic disables Claude Fable 5; Opus 4.8 and GPT-5.5 still the recommended pair

Anthropic has disabled Claude Fable 5 for all users following a US government directive. New sessions default to Opus 4.8, and existing Fable 5 sessions return errors. DAIR.AI's Elvis Saravia says not to panic: Fable 5 wasn't worth it for most tasks, with high cost and nerfed performance. He still recommends Opus 4.8 for planning and GPT-5.5 for execution. The post doesn't spell out the directive's details or how long the suspension lasts.

Why it matters: A major Anthropic model pulled by government order is a rare policy-meets-product event. Elvis provides concrete alternatives and cost judgment, directly useful for Claude users. Score held back because the source is a personal tweet — no official Anthropic statement or order ...

Jun 12Friday

AI HOT (Curated Pool)

Kimi releases and open-sources Kimi-K2.7-Code

Kimi open-sourced K2.7-Code, scoring 11%–31.5% higher than K2.6 on three in-house benchmarks. Inference token usage dropped 30%, and long-coding-task instruction-following and end-to-end success rate both improved. A 6x speed mode is coming; the model is available now via Kimi API and Kimi Code. The post doesn't disclose parameter count, training data, or the open-source license.

Why it matters: Moonshot open-sourced a code model with solid gains on three in-house benchmarks and a 30% inference efficiency improvement — a real cost signal. No external benchmarks (LiveCodeBench, SWE-bench) or parameter count disclosed, so capped below 85. Still, a major Chinese lab open...

AI HOT (Curated Pool)

Hugging Face open-sourced Open-R1, a full reproduction of DeepSeek-R1

Hugging Face published Open-R1 on GitHub, aiming to fully reproduce the DeepSeek-R1 reasoning model. The repo has 26.1k stars and 2.4k forks so far. The body only contains the repo's landing page navigation and metadata; it does not disclose the implementation plan, training data, reproduction progress, or benchmark results. I'd treat this as a public reproduction scaffold and collaboration hub for now, and wait for a technical report before judging fidelity.

Why it matters: Hugging Face launched a full open-source reproduction of DeepSeek-R1, with the repo already at 26.1k stars — strong community interest. But the body only contains project scaffolding and navigation; no implementation plan, training data, or reproduction progress is disclosed y...

Jun 11Thursday

Synced · WeChat

ACL 2026 Oral: LLMs still stumble on phrase-level semantic reasoning

SemanticQA, an ACL 2026 Oral paper, stress-tests frontier models on phrase semantics. GPT-5 nails idiom classification at 85.4% but drops to 78.7% on extraction and 22.5% on interpretation. DeepSeek-R1's accuracy collapses from 81.7% to 35.4% when moving from 4-way to 16-way classification. The study breaks semantic understanding into extraction, categorization, and interpretation—no model handles all three consistently. In multi-step pipelines, upstream extraction errors cascade: GPT-5's end-to-end similarity score falls to 17.3%. Authors from BIGAI and USTB note the static benchmark is already insufficient for 2026 agent workflows.

Why it matters: ACL 2026 Oral paper with counterintuitive findings on phrase-level semantic understanding in GPT-5 and DeepSeek-R1. Concrete numbers across three tasks. Held back from higher bands because it's a single paper without cross-source pickup, and pure academic benchmarking has limi...

Synced · WeChat

Google open-sources 26B text-diffusion MoE; Pichai: generation speed like a racehorse

Google open-sourced DiffusionGemma, a 26B MoE model that activates only 3.8B parameters at inference. Instead of generating tokens one by one, it drafts 256-token blocks in parallel, hitting 1,000+ tokens/sec on an H100—up to 4× faster than autoregressive models. Output quality is lower than standard Gemma 4, so Google still recommends the autoregressive version for production. It ships under Apache 2.0, fits quantized on consumer GPUs with 18GB VRAM, and targets latency-sensitive nonlinear tasks like inline editing and code completion.

Why it matters: Google open-sourced a 26B text diffusion model that skips autoregressive decoding, activating only 3.8B params at inference and hitting 1,000+ tok/s on a single H100. Apache 2.0, with concrete speed comparisons and mechanism details — directly useful for inference folks. Not s...

Jun 10Wednesday

r/LocalLLaMA

Fine-tuned Qwen2.5-7B to 96% of Claude Haiku using about $3 of API calls

A Reddit user fine-tuned Qwen2.5-7B with 1,040 DV-DPO preference pairs, costing about $3 at Claude Haiku rates, and reported 96% composite performance versus Claude Haiku on a domain task, with 11-second latency on a 4-bit T4 setup.

Why it matters: HKR-H/K/R all pass: the hook is cheap fine-tuning, with concrete DV-DPO and latency numbers. Kept at 74 because it is a single Reddit post; task scope and eval details are thin.

AI HOT (Curated Pool)

Anthropic launches safety-treated Mythos-class model Claude Fable 5

Anthropic released Claude Fable 5, a safety-treated Mythos-class model; in high-risk cyber, biochemistry, and distillation domains, it automatically falls back to Opus 4.8, with one trigger per 20 conversations on average.

Why it matters: Anthropic model launches sit in the 85–94 band; HKR-H/K/R all pass via the safety fallback hook, named mechanism, and Claude-user relevance. X-only sourcing limits confidence, so it stays below the top band.

AI HOT (Curated Pool)

Claude Fable launches: Anthropic's alternative reasoning experience

Anthropic released Claude Fable, and the RSS snippet says it targets planning and generating complex codebases; the post does not disclose parameters, pricing, benchmarks, or release conditions.

Why it matters: HKR-H/R are strong for a new Claude reasoning/code angle, while HKR-K is thin: only target use is disclosed. Anthropic bump applies, but missing price, params, benchmarks, and access keep it below must-write.

AI HOT (Curated Pool)

Claude Fable 5 and Claude Mythos 5

Anthropic launched Claude Fable 5 and Claude Mythos 5 at $10 per million input tokens and $50 per million output tokens. Fable 5 leads FrontierCode among frontier models, while Mythos 5 reports about 10x acceleration in drug design and about 80% scientist preference in blinded molecular biology hypothesis tests.

Why it matters: HKR-H/K/R all pass: this is an official Anthropic dual-model release with pricing, coding benchmark, and drug-design speed claims. As a major Claude model update plus Anthropic substantive-update bump, it sits in the 85–94 band.

Hacker News front page

System Card: Claude Fable 5 and Claude Mythos 5

Anthropic published a 319-page system card for Claude Fable 5 and Claude Mythos 5, stating that Fable 5 is for general use with biology and cybersecurity safeguards, while Mythos 5 lifts relevant safeguards and is limited to trusted partners starting with Project Glasswing.

Why it matters: HKR-H/K/R all pass: Anthropic documents two Claude 5 configurations, calls Mythos 5 its most capable model, and gives safety-gating details. This is a same-day Claude substantive update, placed in the 85–94 band.

Jun 9Tuesday

AI HOT (Curated Pool)

Claude Supports Apple Foundation Models Framework With New Swift Package

Anthropic released a Swift package that lets Apple developers call Claude inside the Foundation Models framework with three lines of code, returning typed Swift values and handing off multi-step reasoning, code generation, web search, and data analysis on iOS 27, macOS 27, and related platforms.

Why it matters: HKR-H/K/R all pass: Anthropic is shipping a concrete Claude Swift package for Apple Foundation Models, but this is a developer integration rather than a model release, so it sits high in the 78–84 featured band.

AI HOT (Curated Pool)

OpenAI confidentially files for IPO as Anthropic enters capital race

OpenAI filed a confidential S-1 with the SEC to start IPO review without public revenue or loss data; Anthropic filed last week, and Sam Altman said AI will handle a large share of OpenAI research by March 2028.

Why it matters: HKR-H/K/R all pass: dual frontier-lab IPO filings and Altman’s March 2028 research claim are major. Thin sourcing from an X post keeps it at 90, below the 95+ IPO band.

AI HOT (Curated Pool)

OpenAI plans AI-led research by 2028

Sam Altman said OpenAI plans to have AI perform a large share of its research by March 2028, and the post lists three goals: building automated AI researchers, using them for science and production, and giving each person a personal AGI.

Why it matters: HKR-H/K/R all pass: dated OpenAI AGI-research roadmap with March 2028 and three goals. It stays below P1 because the item is an X repost/summary, not a primary launch or detailed Sam Altman essay with mechanisms.

AI HOT (Curated Pool)

NotebookLM upgrade adds agent capabilities and advanced reasoning

NotebookLM released an upgrade for Google AI Ultra subscribers, adding in-conversation agent capabilities, advanced reasoning, and new output formats. The post does not disclose the specific formats, pricing, or rollout schedule.

Why it matters: HKR-H/K/R all pass: Google confirms NotebookLM adds in-chat agents, advanced reasoning, and multi-output for AI Ultra users. Missing formats, pricing, and rollout details keep it in the mid-weight product-update band.

Jun 8Monday

AI HOT (Curated Pool)

Microsoft AI CEO: Superintelligence Is Coming, but It Won’t Replace Your Job

Mustafa Suleyman said superintelligence is coming without causing mass unemployment; Microsoft signed a new OpenAI contract last October and released seven omnimodal models at Build this week.

Why it matters: HKR-H/K/R all pass: the job-safety claim creates tension, the piece gives an Oct contract and 7-model Build detail, and it hits automation plus Microsoft-OpenAI nerves. As a CEO interview, not a release, it stays in the 78-84 band.

AI HOT (Curated Pool)

The Vanishing Crash in Five-Model Economies: Control and Emergence

The experiment used five models from OpenAI, NVIDIA, OpenBMB, and a self-fine-tuned 500M-parameter model to drive market agents; three interventions failed to reproduce the price crash, and the crash was created only by overriding prices during settlement.

Why it matters: HKR-H/K/R all pass: the angle is counterintuitive, the post gives 5 models, 3 interventions, and a settlement override mechanism, and it speaks to agent-eval reliability. Scope remains an experiment blog, not a major release.

Synced · WeChat

openJiuwen proposes MANGO for multi-agent flow networks

openJiuwen proposed MANGO, a multi-agent flow-network framework that combines reinforcement learning, textual gradients, and a Skip-k mechanism; using GPT-4o-mini, it reports a 12.8% accuracy gain over MaAS on MATH500 and a 5.1% F1 gain over AFlow on DROP.

Why it matters: HKR-K is strong: the post gives mechanisms and a MATH500 delta. HKR-H/R pass for the multi-agent flow-network angle, but this remains a research-framework story, not a major model or platform release.

Synced · WeChat

Alibaba RTPurboV2 uses hundreds of training steps for 10x sparse attention

Alibaba’s RTP team released RTPurboV2, replacing 85% of attention heads with SWA and compressing the remaining 15% retrieval heads using low-rank projection, clustering, and dynamic top-p; the adaptation uses about 600 training steps and roughly 1M label tokens, with reported Prefill speedup up to 9.36x.

Why it matters: HKR-H/K/R all pass: the hook is 100-step training and 10x sparse attention, with concrete mechanisms and 9.36x Prefill speedup. Strong engineering signal from Alibaba, but not a flagship model release or major product launch.

AI HOT (Curated Pool)

OpenAI announces its plan to make AGI benefit everyone

OpenAI outlined its third-phase plan with three goals: build an automated AI researcher, accelerate the economy, and give every person a personal AGI. Sam Altman and Jakub Pachocki said OpenAI internally believes AI systems may perform a significant fraction of its research by March 2028, while alignment, safety standards, and international coordination remain explicit conditions.

Why it matters: OpenAI’s official AGI-benefit plan from Sam Altman and Jakub Pachocki gives three goals plus a March 2028 research-automation forecast. HKR-H, HKR-K, and HKR-R all pass, making it a same-day must-write.

Jun 7Sunday

AI HOT (Curated Pool)

Harness-1: A 20B Stateful Retrieval Subagent Trained with Reinforcement Learning

UIUC and Chroma released Harness-1, a 20B-parameter retrieval subagent trained with reinforcement learning inside a stateful search harness, reporting 0.730 average curated recall across 8 benchmarks, 11.4 percentage points above the next-best open-source subagent and behind only Opus-4.6.

Why it matters: HKR-H/K/R all pass: Harness-1 has a clear RL retrieval-agent mechanism and benchmark numbers. It stays in 78–84 because this is a subagent research/open-source release, not a major lab model launch.

Synced · WeChat

ICML 2026 | FusionRoute: From Expert Routing to Self-Correction in Multi-LLM Collaboration

FusionRoute proposes a token-level multi-LLM collaboration method that freezes expert models and trains a lightweight router to select an expert for each token while merging router logits with expert logits. The paper evaluates it on GSM8K, MATH-500, HumanEval, MBPP, IfEval, and 500 PerfectBlend prompts.

Why it matters: HKR-H/K/R pass: token-level LLM routing is a strong research hook with concrete mechanics. The article lacks lift numbers, code link, and deployment cost, so it stays at the lower featured band.