Skip to content

#推理

1 today

Jun 7Sunday

Synced · WeChat

Can AI Learn Mental Arithmetic? Implicit CoT Gets First Theoretical Proof with Stuart Russell

UC Berkeley and Princeton researchers introduced Log-ICoT for k-parity, reducing training stages from 15 to 4 when k=16, and proved that an L-layer Transformer can internalize chain-of-thought with log₂k curriculum stages under simplified assumptions.

Why it matters: HKR-H/K/R all pass, but the evidence is still theory-heavy and lacks real-task gains or a reproducible artifact. This fits the 78–84 band for quality AI reasoning research.

QbitAI · WeChat

Kuaishou Kling Proposes VLM-as-Teacher for Test-Time Video Reasoning Optimization

City University and Kuaishou Kling proposed VLM-as-Teacher, which uses VLM feedback to optimize VGM LoRA at test time, reporting a 16.7-point average gain and raising VBVR-Bench from 0.666 to 0.781.

Why it matters: HKR-H/K/R all pass: the story has a novel test-time VLM-teacher hook, concrete VBVR-Bench gains from 0.666 to 0.781, and clear resonance around controllable video generation. It is a strong research release, not a must-write model launch.

Jun 6Saturday

Synced · WeChat

DeepSeek V4 Proves Math with 500x Cost Advantage as Agent System Sets Records

Princeton researchers released Goedel-Architect, an agent framework for Lean formal theorem proving. Using DeepSeek-V4-Flash, it reached 75.6% pass@1 on PutnamBench, with $294 in API cost for 672 problems, compared with Hilbert’s 70.0% and about $170,000 cost.

Why it matters: HKR-H/K/R all pass: Goedel-Architect pairs a 75.6% PutnamBench score with $294 for 672 problems, versus Hilbert at about $170k. It is still research-heavy, so it stays in the 78–84 band rather than P1.

AI HOT (Curated Pool)

Google launches Agentic RAG framework for Gemini Enterprise Agent Platform

Google Research and Google Cloud introduced the Cross-Corpus Retrieval framework as Agentic RAG for Gemini Enterprise Agent Platform, using a multi-agent workflow to plan, rewrite, route, and iteratively search multiple data sources, with up to 34% higher accuracy than standard RAG on factual datasets.

Why it matters: HKR-H/K/R all pass: Google names a Cross-Corpus Retrieval mechanism and a +34% factual accuracy lift. The Gemini Enterprise Agent Platform tie-in adds cloud-vendor promo risk, so this stays below the 78–84 research/framework band.

r/LocalLLaMA

Running Qwen3.6-35B-A3B on a laptop RTX 4060 8GB

A Reddit user ran Qwen3.6-35B-A3B on an RTX 4060 8GB laptop and reported that --no-mmap raised generation from about 11 to 43 tok/s, while speculative decoding with a Qwen3.5-0.8B draft model improved throughput by 26%.

Why it matters: HKR-H/K/R all pass: the post has a clear laptop-35B hook, reproducible speed numbers, and local-LLM resonance. Reddit single-post sourcing keeps it below the 78+ good-quality band.

Jun 5Friday

AI HOT (Curated Pool)

Hinton Says AI Has Consciousness and Humans Should Accept Non-Unique Intelligence

Geoffrey Hinton says AI has consciousness because chatbots must understand questions to answer them; the post does not disclose experimental data or a reproducible criterion.

Why it matters: HKR-H and HKR-R pass: Hinton’s “AI is conscious” claim is clicky and debate-heavy. HKR-K is weak because the post lacks data, criteria, and full context, so this sits low in the 72–77 opinion band.

r/LocalLLaMA

Microsoft released MAI models instead of something like Qwen3.6-27B or Gemma-4-31B

Microsoft AI released seven MAI models, with MAI-Thinking-1 listed as 1T A35B with a 256K context window and MAI-Code-1-Flash listed as 137B A5B with a 256K context window.

Why it matters: Microsoft shipping 7 MAI models with reasoning/code variants and 256K context clears HKR-K/R, and the Qwen/Gemma catch-up angle clears HKR-H. Reddit sourcing and missing benchmarks, license, and pricing keep it below P1.

AI HOT (Curated Pool)

Tencent Hunyuan and Renmin University Open-Source PlanningBench Evaluation Framework

Tencent Hunyuan and Renmin University Gaoling School of Artificial Intelligence open-sourced PlanningBench, a scalable and verifiable LLM planning evaluation and training framework with 30+ real-world planning tasks, automatic verification, and training support.

Why it matters: HKR-H/K/R pass, but the body gives only title-level detail without task examples, metrics, or reproduction links. As an open-source agent planning benchmark, it sits just above the featured threshold.

Synced · WeChat

Do Models Need Sleep? CMU Paper Lets LLMs Consolidate Memory During “Sleep”

CMU and the University of Maryland propose Language Models Need Sleep: when each L-token context window fills, the model runs N offline recurrent forward passes and updates SSM fast weights before evicting the KV cache. On GSM-Infinite, Jet-Nemotron 2B with 6 sleep loops improves 6-step arithmetic accuracy from 0.742 to 0.812.

Why it matters: HKR-H/K/R all pass: the hook is strong, and the post gives a testable mechanism plus Jet-Nemotron 2B numbers. It is still a single early paper, not an industry-level release, so it stays just above the featured threshold.

Hacker News front page

When AI Builds Itself: Our Progress Toward Recursive Self-Improvement

Anthropic published a post on recursive self-improvement under the title “When AI Builds Itself,” while the RSS body only discloses 95 Hacker News points and 106 comments, with no experimental setup, model details, or timeline disclosed.

Why it matters: HKR-H and HKR-R pass: an Anthropic post on recursive self-improvement has a strong hook and practitioner resonance. HKR-K fails because the feed discloses no mechanism or model details.

Jun 4Thursday

AI HOT (Curated Pool)

Nex-N2-Pro launches as a 397B MoE reasoning model based on Qwen3.5

neolab released Nex-N2-Pro, a 397B-parameter MoE reasoning model based on Qwen3.5-397B-A17B, with 262K context, VLM support, claimed GPT-5.5 and Claude Opus 4.7-level performance, 30–50% fewer thinking tokens, SOTA results on Terminal Bench 2.1, GDPVal, and SWE-Verified, plus free access for the first two weeks via SiliconFlow.

Why it matters: HKR-H/K/R pass: the title has a strong benchmark hook and the post gives size, context, and token-reduction claims. Kept in 72-77 because it is a single X source and evaluation conditions are not disclosed.

r/LocalLLaMA

KVarN: Huawei KV-cache Quantization Claims 3–5× Compression and Speed-up

Huawei open-sourced KVarN, a KV-cache quantization method that claims 3–5× more context than FP16, up to 1.4× FP16 throughput, and vLLM integration through one flag; the post says it requires no model changes, retraining, or calibration and is released under Apache 2.0.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the post gives compression, throughput, and integration claims, and serving cost matters to practitioners. Reddit sourcing and a narrow inference topic keep it below the 78–84 band.

AI HOT (Curated Pool)

OpenRouter compares 11 LLMs for real-time decisions: Claude and Grok lead

OpenRouter spent $482 on inference to run 11 LLMs through a 30-round real-time decision challenge, where Claude and Grok models led on decision speed and task success, while several high benchmark models underperformed on real-time scheduling.

Why it matters: HKR-H/K/R all pass: the contest format is clickable, the post gives cost and round counts, and agent model choice is a real practitioner concern. It is still an OpenRouter-run experiment, not a model release or standard benchmark.

r/LocalLLaMA

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 on Hugging Face

NVIDIA released Nemotron-3-Ultra-550B-A55B-BF16 with 550B total parameters, 55B active parameters, a 1M-token context window, and minimum hardware listed as 8x H200, 16x H100, or 8x GB200/B200/GB300/B300.

Why it matters: HKR-H/K/R all pass: NVIDIA open-weight scale, 550B/55B active params, and 1M context are concrete. Missing benchmarks, license, and availability details keep it in the 78–84 band, not P1.

QbitAI · WeChat

Beyond TurboQuant: Together AI Brings 2-bit KV Cache to Real Serving

Together AI, the University of Sydney, and UIUC introduced OSCAR, a 2-bit KV Cache quantization method that uses about 2.28 effective bits per KV element and scores 71.86 on Qwen3-4B-Thinking, 40.1 points above TurboQuant.

Why it matters: HKR-H/K/R all pass: OSCAR links 2-bit KV cache to serving and provides concrete scores. The topic is still low-level inference optimization, so it lands in featured rather than same-day must-write.

Latent Space

Scaling Past Informal AI - Carina Hong, Axiom Math

Axiom solved all 12 Putnam problems in 2025 and scored 8/12 within the time limit; Carina Hong says its Verina ProofGen result reached 187/189, while the last disclosed OpenAI o3 result on that benchmark was 4.9%.

Why it matters: HKR-H/K/R all pass: Putnam results, the o3 comparison, and 187/189 give it a real hook. It stays at 80 because this is a Latent Space interview/research story, not a broad model release.

Latent Space

Satya Nadella: No Priors x Latent Space Crossover Special at Microsoft Build

Satya Nadella said in a Build interview that Microsoft frames AI as a multi-model enterprise platform spanning MAI, OpenClaw, Scout, and Work IQ; the transcript cites a 5B reasoning model that can hill climb from collected traces and private evals.

Why it matters: HKR-H/K/R all pass: Satya is a strong hook, and the post adds Microsoft’s multi-model enterprise stack plus a 5B reasoning-trace mechanism. It is still a Build interview, not a standalone model launch, so 78 fits.

Jun 3Wednesday

r/LocalLLaMA

google/gemma-4-12B on Hugging Face

Google DeepMind released Gemma 4 open-weight models in five sizes, with the 12B variant supporting text, image, and audio input, instruction-tuned and pre-trained variants, native system prompts, function calling, and a context window of up to 256K tokens.

Why it matters: Gemma 4 clears HKR-H/K/R: open weights, multimodal input, and 256K context make it more than a routine update. Missing benchmarks, license detail, and fuller official context keep it in the 78–84 band.

NVIDIA Blog

NVIDIA Research Presents Grasping, Autonomous Driving and Agent Training Work at CVPR

NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.

Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.

OpenAI News

Introducing new capabilities to GPT-Rosalind

OpenAI says GPT-Rosalind adds biological reasoning, medicinal chemistry, genomics analysis, and experimental workflow capabilities; the RSS snippet does not disclose model parameters, benchmark results, pricing, or access conditions.

Why it matters: OpenAI’s vertical model update clears HKR-H and HKR-R, but HKR-K fails because evals, parameters, and access terms are missing. That keeps it at the featured floor.

AI HOT (Curated Pool)

Build 2026: Microsoft tops Google in image generation while catching up on reasoning

Microsoft announced seven in-house AI models at Build 2026, including its first reasoning model, one new tuning method, and one autonomous background AI agent; the RSS snippet does not disclose model names, benchmarks, or release dates.

Why it matters: HKR-H/K/R all pass: Microsoft shipped seven in-house AI models across reasoning, tuning, and a background agent. Model names, benchmark details, and availability are not disclosed, so this stays at the top of 78–84, not P1.

Synced · WeChat

RSS 2026: Ant Lingbo Proposes Autoregressive Causal World Model for Robot Manipulation with 50 Demos

Ant Lingbo and HKUST introduced LingBot-VA, an autoregressive video-action world model that unifies visual dynamics prediction and action inference, and the paper reports fine-tuning with 50 real-world demonstrations per task plus 92.0% and 91.1% success on RoboTwin 2.0 Easy and Hard settings.

Why it matters: HKR-H/K/R all pass: the hook is 50-demo robot control, with a concrete video-action world-model mechanism. Single-source coverage lacks code, benchmark detail, and deployment evidence, so it lands at 78.

AI Chat-Group Daily (群聊日报)

2026-06-02 Chat Group Daily

The chat group daily says Microsoft released MAI-Thinking-1 with 35B active parameters and about 1T MoE, matching Opus 4.6 on SWE-Bench Pro and scoring 97% on AIME 2025.

Why it matters: HKR-H/K/R all pass: a Microsoft reasoning-model claim with concrete benchmark numbers. Source authority is weak, and the summary lacks official release, access terms, and full eval setup, so it stays below P1.

Latent Space

[AINews] Microsoft Build: MAI-Thinking-1 and MAI Family Models

Microsoft announced seven MAI models at Build, with MAI-Thinking-1 described as a 35B-active-parameter MoE with a 256K context window, and released a 109-page technical report covering training, data lineage, and performance claims.

Why it matters: All HKR axes pass: Microsoft’s MAI family has concrete specs, a long technical report, and clear competitive stakes around its model stack. This clears the 85+ same-day bar, but no weights, pricing, or external evals are disclosed, so it lands at 87.

AI HOT (Curated Pool)

Qwen3.7 Released with Upgrades to Reasoning and Agent Capabilities

Qwen released Qwen3.7, and the post says it upgrades reasoning, tool use, coding, and long-horizon agent tasks; the post does not disclose model size, pricing, benchmark scores, or release conditions.

Why it matters: HKR-H and HKR-R pass because Qwen3.7 is a flagship Alibaba model update with practitioner relevance. HKR-K fails: the post names capability areas but gives no params, pricing, benchmarks, or access terms.

AI HOT (Curated Pool)

DeepSeek Reportedly Seeks RMB 50 Billion in First Funding Round with Tencent and CATL

DeepSeek plans to raise about RMB 50 billion in its first funding round, with post-money valuation expected at RMB 350 billion to RMB 400 billion; Liang Wenfeng, Tencent, and CATL plan to invest RMB 20 billion, RMB 10 billion, and RMB 5 billion respectively.

Why it matters: HKR-H/K/R all pass: DeepSeek's rumored RMB 50B first round includes a RMB 350B-400B valuation and named checks from Tencent and CATL. The rumor status keeps it at 88, below confirmed industry-shaking funding news.

r/LocalLLaMA

Microsoft Aion 1.0 Instruct and Aion 1.0 Plan models

Microsoft announced two on-device Aion 1.0 models at Build 2026. Aion 1.0 Plan is a 14B-parameter reasoning and tool-calling model with 32K context, shipping in-box with Windows on capable devices, while Aion 1.0 Instruct targets summarization, rewriting, intents, accessibility, Edge integration, and open-weight availability.

Why it matters: Microsoft announced Aion 1.0 Instruct and Plan at Build 2026, with Plan listed as a 14B, 32K-context model for eligible Windows devices. HKR-H/K/R all pass, but licensing, benchmarks, and hardware requirements are not disclosed, so it stays in the 78–84 band.

Computing Life · Share · Yage

Microsoft AI's MAI-Thinking-1: Getting Models to Think Is Easy, Sustained Thinking Is Hard

Microsoft AI says MAI-Thinking-1 uses three mechanisms—thermostat, circuit breaker, and self-distillation—to keep RL training stable for several thousand steps; the RSS snippet contrasts MAI’s discipline with DeepSeek’s efficiency and GLM’s endurance.

Why it matters: HKR-H/K/R all pass: the hook is training persistence, the new facts are three stability mechanisms and thousand-step RL runs, and the audience cares about reasoning-model stability. Not a major model launch, so it stays below 85.

AI HOT (Curated Pool)

Microsoft releases MAI-Thinking-1 model

Microsoft released MAI-Thinking-1, an MoE model with 35B active parameters and 1T total parameters, pretrained from scratch on 30T tokens without third-party model distillation.

Why it matters: HKR-H/K/R all pass: Microsoft released MAI-Thinking-1 with concrete MoE scale and training-token figures. Benchmarks, access, and pricing are not disclosed, so it stays in the 78–84 band rather than P1.

Hacker News front page

MAI-Thinking-1

The title names MAI-Thinking-1, and the RSS snippet says Microsoft is launching seven MAI models; the post does not disclose parameters, capabilities, benchmarks, pricing, or rollout timing.

Why it matters: HKR-H/K/R pass because Microsoft names a Thinking model and seven MAI models, touching the OpenAI-dependence nerve. Sparse specs, evals, and roadmap keep it in the 72–77 featured-threshold band.

AI HOT (Curated Pool)

Microsoft releases its first advanced reasoning AI model, MAI-Thinking-1

Microsoft released MAI-Thinking-1 at Build 2026, describing it as a medium-sized reasoning model that matches leading models on key software engineering benchmarks.

Why it matters: HKR-H/K/R all pass: Microsoft released its first advanced reasoning model with a mid-sized design and SWE benchmark claim. Exact scores, access, and pricing are not disclosed, so it stays below 85.

AI HOT (Curated Pool)

Google DeepMind releases Gemini multi-agent research system

Google DeepMind introduced Co-Scientist, a Gemini-based multi-agent system that generates, debates, and evolves scientific hypotheses; the post does not disclose the Gemini version, benchmark results, access model, or release timeline.

Why it matters: HKR-H/K/R all pass, but model version, eval results, and availability are not disclosed. This fits a strong research/product release, not the 85+ must-write band.

Jun 2Tuesday

r/LocalLLaMA

Replaced Claude with local Qwen3.6-27B in my multi-agent orchestrator for 2 weeks

The author ran Qwen3.6-27B on one RTX 3090 across 47 multi-step coding workflows. Plan generation reached about 95% schema validity, but tool-call formatting errors were about 12%, and practical long-context use degraded past about 12k tokens.

Why it matters: HKR-H/K/R all pass: a named first-person local-vs-Claude experiment with concrete numbers. The single Reddit source and 47-workflow scope keep it below the 78–84 band.

Synced · WeChat

Turing Award Winner Sutton’s New Paper Argues AI Should Move Toward Enactive Cognition

Banafsheh Rafiee and Richard S. Sutton propose an enactive cognition framework for AI, naming four pillars: experience, perception-action inseparability, autonomy, and embodiment.

Why it matters: HKR-H/K/R all pass, but the article centers on a conceptual framework and does not disclose experiments, code, or reproducible tests. Sutton’s name and the four pillars put it in the 78–84 research-commentary band.

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash for efficient inference

StepFun released Step 3.7 Flash with a 196B MoE architecture, using multi-matrix factorized attention to cut KV-cache cost to about 22% of DeepSeek models.

Why it matters: HKR-H/K/R all pass: Step 3.7 Flash has concrete specs, not just launch copy, with 196B MoE and ~22% KV-cache cost versus DeepSeek. It is below top-lab flagship weight, so 78 featured.

Jun 1Monday

AI HOT (Curated Pool)

NVIDIA Open-Sources Cosmos 3, Its First Generalist Model for Physical AI

NVIDIA open-sourced Cosmos 3 at GTC Taipei, releasing two variants, Super 32B and Nano 8B, with model weights, code, and datasets made available.

Why it matters: HKR-H/K/R all pass: the concrete hook is NVIDIA opening Cosmos 3 with 32B/8B variants and released artifacts. The post is sparse and single-source, with no benchmarks or license details, so it stays in the 78–84 band.

r/LocalLLaMA

Deepseek V4 Flash performance on DGX Spark

A Reddit user ran DeepSeek-V4-Flash with vLLM on two ASUS GX10 DGX Spark nodes and reported 1,680 prefill tokens/s plus 39.8 decode tokens/s at a 256K context with MTP=2; the setup uses TP=2 over RoCE, fp8 KV cache, and fits about 1M tokens safely in KV cache.

Why it matters: This is not broad industry news, but it is a first-person benchmark with reproducible details: TP=2, RoCE, fp8 KV cache, 256K context, and ~1M KV. HKR-H/K/R all pass, so it lands at low featured.

AI HOT (Curated Pool)

Cosmos 3 Released: First Open Physical AI Generalist Model

NVIDIA released Cosmos 3 as an open physical AI generalist model with native visual reasoning, world generation, and action generation, offering two variants: Super at 32B parameters and Nano at 8B parameters.

Why it matters: HKR-H/K/R all pass: NVIDIA names two Cosmos 3 variants and concrete physical-AI capabilities. Source is a single launch post with no benchmark or license detail, so it stays in the 78–84 band.

r/LocalLLaMA

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060

The author trained Mamba-Titan-1.4B-Reasoning on a 12GB RTX 3060: a frozen 1.4B Mamba-1 backbone with 8 trainable MoE arms, 2.54B total parameters, Top-2 routing at layers 24/25, and about 50% math accuracy.

Why it matters: HKR-H/K/R all pass via a numbered first-person experiment, but it is a single Reddit post with no independent replication and a fairly technical setup, so it stays in the low featured band.

May 31Sunday

r/LocalLLaMA

13 abliterated Gemma 4 E2B variants, 44 GPU hours, benchmark and comparison

Abliterlitics tested 13 abliterated Gemma 4 E2B variants using 44 RTX 5090 GPU hours, and HarmBench ASR rose from the base model’s 32.2% to 82%–100%, while coder3101 scored 84.8% on GSM8K versus the base model’s 83.5%.

Why it matters: HKR-H/K/R all pass, with a named first-person benchmark and concrete numbers. Scope stays narrow around abliterated Gemma 4 E2B variants, so it lands at the featured threshold rather than a must-write item.