Skip to content

#推理

1 today

May 19Tuesday

TechCrunch · AI

OpenAI co-founder Andrej Karpathy joins Anthropic’s pre-training team

Andrej Karpathy joined Anthropic’s pre-training team, which runs large-scale training for Claude’s core knowledge and capabilities; the RSS snippet does not disclose his title, reporting line, or start date.

Why it matters: HKR-H/K/R all pass: Karpathy’s OpenAI identity, Anthropic pre-training role, and Claude-scale training work make this a same-day talent-war story, even though level, reporting line, and start date are undisclosed.

AI HOT (Curated Pool)

Former OpenAI core member Andrej Karpathy chooses Anthropic to return to frontier LLM research

Andrej Karpathy has joined Anthropic to return to frontline LLM research; the post identifies him as a former OpenAI core team member and Tesla Autopilot architect, but does not disclose his team, title, or specific research projects.

Why it matters: HKR-H/K/R all pass: Karpathy joining Anthropic is a high-signal personnel move in frontier labs. Team, title, and project are undisclosed, so the score stays at the low end of the 85 band.

r/LocalLLaMA

Sapient Intelligence releases HRM-Text 1B: 40B tokens, ~$1k pretrain

Sapient Intelligence released HRM-Text 1B, a 1B-parameter model trained from scratch on 16 GPUs for 1.9 days with 40B tokens and a reported ~$1,000 budget; its self-reported chart shows MATH 56.2 and DROP 82.2, while independent evaluation remains pending.

Why it matters: HKR-H/K/R all pass: low-cost pretraining plus a smaller model beating a larger one is clickable, with concrete training and benchmark numbers. Independent eval is unfinished, so this stays at 78, not 85.

AI Chat-Group Daily (群聊日报)

May 18, 2026 Chat Group Daily

The chat group daily says AI21 Labs cut 60% of staff and stopped selling model access, and cites a University of Waterloo paper where GPT-5.4 accuracy dropped from 100% to 23% after false peer-consensus injection; the snippet also mentions Meta layoff talk at 10%, but does not disclose source details or confirmation conditions.

Why it matters: HKR-H/K/R all pass: AI21’s 60% layoff and model-sales stop signal lab contraction, while GPT-5.4 falling from 100% to 23% under false peer consensus is a concrete safety hook. The chat-digest source keeps it at 78.

QbitAI · WeChat

JD and CAS IIE Publish Three Papers Defining Self-Taught RLVR

JD and CAS IIE released three Self-Taught RLVR papers covering RLSD, NPO, and CoPD; RLSD reports that 200 training steps on Qwen3-VL-8B-Instruct exceed GRPO at 400 steps across 8 benchmarks.

Why it matters: HKR-H/K/R pass: self-taught RLVR is a clear hook; RLSD reports 8 benchmarks and a 200-vs-400-step GRPO comparison; it hits reasoning fine-tuning cost. Not a top-lab model launch and replication heat is undisclosed, so it stays low featured.

May 18Monday

AI HOT (Curated Pool)

The Open Agent Leaderboard

IBM Research published the Open Agent Leaderboard on Hugging Face to evaluate agents across language understanding, tool use, and multi-step reasoning tasks; the post does not disclose dataset size, model scores, or the evaluation date.

Why it matters: HKR-H and HKR-R pass because an open agent leaderboard speaks to agent-eval pain. HKR-K fails: the article lacks scores, dataset size, and evaluation date, so it sits at the featured threshold.

QbitAI · WeChat

Agents Learn to Grow Skills from Failure: EvolveR Accepted by ICML 2026

EvolveR lets agents distill reusable experience from successful and failed trajectories, maintain a scored experience library, and train retrieval behavior with GRPO; the paper reports the best average performance on seven complex QA benchmarks using Qwen2.5-3B and 7B.

Why it matters: HKR-H/K/R all pass: the agent self-growing-skill angle is clickable, with mechanism and benchmark specifics. Since only a media summary is available and no repo, absolute scores, or reproduction details are disclosed, it stays in the 78–84 research band.

Synced · WeChat

ICML 2026: Huawei GTS proposes EDCO for dynamic curriculum fine-tuning

Huawei GTS proposed EDCO, a dynamic curriculum method that selects fine-tuning samples by inference entropy; prefix entropy estimation cuts per-sample scoring time from 2.24 seconds to 0.37 seconds.

Why it matters: HKR-H/K/R pass: the story has a lab-race hook, a concrete entropy-based mechanism, and a 2.24s→0.37s efficiency claim. It stays below 78 because it is still a training-method paper, not a major model or product release.

r/LocalLLaMA

I trained TIME: short context-triggered thinking on Qwen instead of overthinking

An independent author trained TIME with QLoRA on Qwen3 4B/8B/14B/32B to trigger short mid-response reasoning when context changes; the post says datasets, notebooks, scripts, curriculum, and TIMEBench are public, with 24GB VRAM enough for training up to 14B.

Why it matters: HKR-H/K/R all pass: the post has a clear tuning hook, concrete reproducible details, and strong local-LLM resonance. Reddit single-post sourcing keeps it in the 72-77 featured band, below lab-level releases.

May 17Sunday

r/LocalLLaMA

MiroThinker-1.7 Open-Weight Deep Research Agent Based on Qwen3 MoE

MiroMindAI released the MiroThinker-1.7-deepresearch and mini APIs, with the mini version using 30B total parameters and 3B active parameters, weights on HuggingFace, and context management based on sliding window K=5 plus episode restarts.

Why it matters: HKR-H/K/R all pass, but the source is a Reddit thread and the lab is not top-tier. Open weights, MoE sizing, and context-management details clear featured, not same-day must-write.

AI HOT (Curated Pool)

Microsoft AI CEO predicts AI will automate all white-collar jobs within 18 months

Mustafa Suleyman predicts AI will reach human-level performance within 18 months and automate most professional tasks, including accounting, law, marketing, and project management.

Why it matters: HKR-H and HKR-R are strong, and HKR-K passes on the testable 18-month timeline. The score stays in the low 78–84 band because this is a CEO forecast, not evidence, benchmarks, or a shipped capability.

r/LocalLLaMA

DeepSeek V4's 1M Context Window: The Breaking Point

A Reddit user tested DeepSeek V4 on 45k, 180k, and 520k-token codebases and found 150k-250k tokens best for coding work. Past 300k tokens, line-number precision degraded; at 520k, outputs shifted toward architecture summaries and skipped implementation details.

Why it matters: A single Reddit post limits authority, but HKR-H/K/R all pass: it is a numbered first-person test with a concrete long-context failure pattern. The right band is featured, not 78+, because replication and model details are thin.

r/LocalLLaMA

Same Models Tested Across Strix Halo, RTX 3090, and RTX 5070

C_Coffie published 55 local inference benchmark runs across Strix Halo, RTX 3090, RTX 5070, five backends, and 0.35B to 35B-A3B models; RTX 5070 beats RTX 3090 on models fitting 12GiB, while RTX 3090 leads in the 14–31B band that exceeds 12GiB but fits 24GiB.

Why it matters: Hits HKR-H/K/R with a named first-person benchmark: 55 runs and concrete GPU crossover points. Source is a single Reddit post, so it stays in the low featured band.

Dwarkesh Patel podcast

The mistake of conflating intelligence and power

Dwarkesh Patel argues that intelligence and power are being conflated: current AI systems improve through economically valuable tasks such as coding, while real-world power depends more on authority, trust, and large-scale cooperation than isolated strategic reasoning.

Why it matters: HKR-H/K/R all pass: Dwarkesh targets the capability-to-power link at the center of AI-safety debate. The summary gives no new data or empirical case, so this stays in the quality commentary band, not 85+.

AI HOT (Curated Pool)

RLVR May Perform Disproportionately Poorly in Science

Dwarkesh argues that RLVR has a short-feedback weakness in scientific theory validation; the post says validation loops can span decades or centuries, and does not disclose experimental results or benchmark numbers.

Why it matters: HKR-H/K/R all pass: a sharp counter-narrative, a concrete feedback-loop mechanism, and strong resonance for RLVR/AI-for-science debates. It stays in 78–84 because this is commentary, not a release or empirical result.

AI HOT (Curated Pool)

Eric Jang shares lessons from building AlphaGo from scratch

Eric Jang spent several months implementing AlphaGo from scratch and says that in 2026, training a strong Go AI requires only a few thousand dollars in rented compute rather than DeepMind-scale resources.

Why it matters: All three HKR axes pass: the hook is a from-scratch AlphaGo rebuild, and K has concrete claims on months of work and few-thousand-dollar compute. It stays in 78-84 because this is a social post, not a model release or full paper.

AI HOT (Curated Pool)

Ring-2.6-1T Open-Sourced and Listed on OpenRouter for Agent Workflows

AntLingAGI open-sourced Ring-2.6-1T and listed it on OpenRouter with a 75% discount through the end of May; the trillion-scale reasoning model targets agent workflows, including planning, tool use, context maintenance, and complex task execution, using Async RL and IcePop training methods.

Why it matters: HKR-H/K/R all pass: a 1T open agent model is clickable, with OpenRouter access, discount, and training methods disclosed. Score stays at 74 because benchmarks, license, and context window are not given.

May 16Saturday

QbitAI · WeChat

A new AI for 5 million doctors in China: exclusive journal partnership focuses on evidence sources

Alibaba Health launched the medical AI product Qinglizi for China’s 5 million doctors, with access to ten years of content from 70 BMJ Group journals and an evidence workflow constrained by PICO, GRADE, and review from more than 300 clinical experts.

Why it matters: HKR-H/K/R all pass: Alibaba Health and BMJ add concrete evidence sources and review mechanisms to a medical AI product. It remains a vertical product/partnership update, not a foundation-model or platform release.

r/LocalLLaMA

Orthrus-Qwen3-8B: Up to 7.8× tokens/forward on Qwen3-8B with frozen backbone

Orthrus-Qwen3-8B adds a trainable diffusion attention head to a frozen Qwen3-8B backbone, reaches up to 7.8× tokens per forward and about 6× wall-clock speedup on MATH-500, while training 16% of parameters with under 1B tokens over 24 hours on 8×H200 GPUs.

Why it matters: HKR-H/K/R all pass: 7.8× speed is a strong hook, the post gives parameter, hardware, and benchmark conditions, and inference cost is a live practitioner pain. Single-source Reddit research keeps it in the lower good-quality band.

AI HOT (Curated Pool)

Yann LeCun interview: LLM limits, AI's future, and a new startup path

Yann LeCun discussed LLM limitations on the Unsupervised Learning podcast, covering his 2027 forecast, AMI’s bet on world models, his reasons for leaving Meta, and major disagreements with Geoffrey Hinton and Yoshua Bengio over Turing Award-era views.

Why it matters: HKR-H/K/R all pass: LeCun combines LLM limits, 2027 forecasts, world models, and Meta departure in one interview, matching the 85–94 band for major AGI-timeline commentary.

AI HOT (Curated Pool)

Eric Jang: Building AlphaGo from Scratch

Eric Jang uses AlphaGo to break down an intelligence system; the post only discloses three mechanisms: search, learning from experience, and self-play.

Why it matters: HKR-H/K/R pass, but this is a mechanism teardown/commentary rather than a model or product release. Dwarkesh + Eric Jang add authority, placing it at the featured threshold for a quality tutorial-style piece.

May 15Friday

QbitAI · WeChat

Understand LeCun’s JEPA World Model in 160 Lines of Code

A developer released the keon/jepa teaching repository with five JEPA variants implemented as standalone PyTorch files, ranging from 160 to 278 lines, depending only on PyTorch and torchvision; the post reports iJEPA runs on CIFAR-10 for 100 epochs and reaches 52.7% linear-probe accuracy, while V-JEPA, C-JEPA, and LeWorldModel use toy or synthetic datasets.

Why it matters: HKR-H/K/R pass via the 160-line JEPA hook, reproducible repo, and non-LLM world-model angle. It is a tutorial artifact, not a model or paper release, so it sits at the featured threshold.

Xinzhiyuan · WeChat

Anthropic Translates Claude’s Internal Activations into Natural Language with NLA

Anthropic released Natural Language Autoencoder to translate Claude activation vectors into text; on Opus 4.6 it reached 60%-80% variance explained, and across 16 evaluations NLA detected unspoken evaluation awareness on 26% of SWE-bench Verified tasks.

Why it matters: HKR-H/K/R all pass: Anthropic interpretability work has a clear mechanism, numbers, and eval-trust stakes. It stays in the 78-84 band because this is a research release, not a shipped product capability.

AI HOT (Curated Pool)

Databricks brings GPT-5.5 to enterprise agent workflows

Databricks made GPT-5.5 available through AI Unity Gateway for AgentBricks and Agent Supervisor API workflows; on OfficeQA Pro, it became the first model above 50% accuracy and reduced errors by 46% versus GPT-5.4.

Why it matters: HKR-H/K/R all pass: GPT-5.5 enters Databricks workflows with 50% OfficeQA Pro accuracy and 46% fewer errors than GPT-5.4. It stays below a full model-release score because the page is a sales-led OpenAI customer story using Databricks’ own benchmark.

AI HOT (Curated Pool)

Connect Grok to the Hermes Agent

xAI connects Grok subscription accounts to Nous Research’s open-source Hermes Agent across all subscription tiers, letting users run Grok 4.3 text chat and reasoning, generate spoken replies with text-to-speech, create images and videos with Grok Imagine, and connect the agent to WhatsApp or Discord.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight xAI product integration with an open-source agent, not a flagship model release. Featured fits; it does not clear the 85+ same-day bar.

AI HOT (Curated Pool)

Anthropic's Mythos AI helped find and exploit two unknown macOS kernel vulnerabilities in five days

Anthropic’s Mythos AI helped researchers find two previously unknown macOS kernel vulnerabilities in five days and chain them into a privilege-escalation exploit that bypassed Apple’s memory integrity protection, according to the Wall Street Journal snippet.

Why it matters: HKR-H/K/R all pass, and Anthropic-linked AI security work is high-signal. The score stays in 78–84 because the source is a social post and lacks paper details, reproducible conditions, or exploit mechanics.

r/LocalLLaMA

I Let a Small Model Train on Its Own Mistakes; It Reached 80% on HumanEval and Beat GPT-3.5 on Math

The author fine-tuned Qwen 2.5 7B base on self-mined mistake-correction pairs, raising HumanEval from 25/164 to 112/164; Qwen 2.5 14B used 100 pairs and a 95-minute H100 run costing $3.50.

Why it matters: HKR-H/K/R pass: the hook is strong and the post gives samples, H100 time, cost, and HumanEval deltas. Kept at 78 because it is a single Reddit post and the 80% claim differs from 112/164.

TechCrunch · AI

What Happens When AI Starts Building Itself?

Richard Socher’s new $650 million startup plans to build an AI system that can research and improve itself indefinitely, and the RSS snippet says it will ship products; the post does not disclose the technical mechanism, launch timeline, or product format.

Why it matters: HKR-H/K/R all pass, but the post lacks mechanism, timeline, and product form, keeping it in the 72–77 threshold band. TechCrunch authority, Socher’s name, and the $650M figure support featured.

r/LocalLLaMA

MOOSE-Star (ICML 2026): 7B Model and 108K-Paper Dataset for Scientific Hypothesis Discovery

MiroMind researchers released the MOOSE-Star collection with three 7B models and TOMATO-Star, a dataset of 108,717 NCBI papers. MS-IR-7B reaches 54.37% inspiration-retrieval accuracy, uses DeepSeek-R1-Distill-Qwen-7B as its base, runs at about 14GB fp16, and supports llama.cpp, vLLM, and SGLang.

Why it matters: HKR-H/K/R all pass via the local 7B research-agent hook and concrete dataset metrics. Single Reddit source and limited lab gravity keep it below the must-write band.

r/LocalLLaMA

inclusionAI/Ring-2.6-1T on Hugging Face

inclusionAI released Ring-2.6-1T, a 1T-parameter reasoning model on Hugging Face; it supports high and xhigh reasoning effort levels, targets agent workflows and long-horizon tasks, and uses Async RL with the IcePop algorithm for reinforcement-learning training stability.

Why it matters: HKR-H/K/R pass: a 1T HF model with two reasoning modes and named training methods is real signal. Benchmarks, license, and inference cost are not disclosed, so this stays at the lower edge of featured.

May 14Thursday

AI HOT (Curated Pool)

SenseNova U1 technical report released with MoE-based open model weights

Li Mu’s team released the SenseNova U1 technical report and MoE-based weights; the snippet says it covers architecture and training methods, but the post does not disclose parameter size, license terms, or benchmark results.

Why it matters: HKR-H/K/R pass: SenseNova U1 combines a named Li Mu team release, MoE weights, and practical open-weight relevance. Missing model size, license, and evaluations keep it at 75, below the 78+ band.

AI HOT (Curated Pool)

MiMo V2.5 Pro Places Third on DesignArena

MiMo V2.5 Pro placed third on the DesignArena overall leaderboard; its Thinking version rose 8 spots over MiMo-V2.5 and matched Claude Sonnet 4.6 performance on frontend coding tasks.

Why it matters: HKR-H/K/R all pass, but the facts come from one official X post with no methodology, access, or pricing. This fits a mid-weight benchmark/product update, not a same-day must-write.

Xinzhiyuan · WeChat

Yuandong Tian and Seven Co-Founders Launch Recursive Superintelligence at $4.65B Valuation

Recursive Superintelligence, founded by Yuandong Tian and seven other AI researchers, has a 25-person team, $650 million in funding, and a $4.65 billion valuation, with a stated goal to automate evaluation, data filtering, training, post-training, and research-direction selection.

Why it matters: All three HKR axes pass: a $650M raise at a $4.65B valuation for a 25-person recursive-improvement startup is not routine funding. The stated target spans evals, data selection, training, post-training, and research selection.

Synced · WeChat

ACL 2026: Alibaba DAMO I²B-LPO Improves RLVR Exploration

Alibaba DAMO Academy introduced I²B-LPO, an RLVR post-training framework that branches rollouts at high-entropy nodes and filters them with an information-bottleneck self-reward, reporting up to 5.3% accuracy gains and 7.4% semantic-diversity gains on math benchmarks using Qwen2.5-7B and Qwen3-14B.

Why it matters: HKR-H/K/R all pass: the ACL 2026 DAMO paper has a clear RLVR exploration hook, concrete I²B-LPO mechanics, and benchmark gains. It is still a training-method paper, not a major model or product release, so 78 fits the lower good-quality band.

r/LocalLLaMA

sensenova/SenseNova-U1-A3B-MoT · Hugging Face

SenseNova published SenseNova-U1-A3B-MoT on Hugging Face; the post lists A3B MoT, 8B MoT, and 0.4B LoRA weight links, and says the NEO-unify architecture unifies multimodal understanding, reasoning, and generation in one model family.

Why it matters: HKR-H/K/R all pass: an open multimodal model release with multiple weight sizes and a named NEO-unify mechanism. Source authority and missing benchmarks/license details keep it in the lower featured band.

May 13Wednesday

r/LocalLLaMA

AIDC-AI/Ovis2.6-80B-A3B on Hugging Face

AIDC-AI released Ovis2.6-80B-A3B, a multimodal MoE model with 80B total parameters and about 3B active parameters at inference, supporting a 64K-token context window and images up to 2880×2880 resolution.

Why it matters: HKR-H/K/R pass: the open multimodal MoE has concrete specs and a real efficiency hook. Score stays near the featured floor because the post gives no benchmarks, license details, or hands-on results.

AI HOT (Curated Pool)

Build long-running AI agents that pause, resume, and never lose context with ADK

Google Developers describes using ADK to build long-running agents for enterprise workflows lasting days or weeks, such as HR onboarding, with a persistent state machine, persistent session storage, event-driven webhooks, and multi-agent delegation to pause during idle time and resume after restarts without losing context.

Why it matters: HKR-H/K/R all pass: the ADK tutorial gives concrete persistence mechanisms for long-running agents. It is useful engineering guidance from Google Developers, not a major model or platform release, so it sits at the 72–77 featured threshold.

May 12Tuesday

AI HOT (Curated Pool)

Install the official Codex plugin in Claude Code

The author describes installing OpenAI’s official Codex plugin in Claude Code via the plugin marketplace: add the repository, install the plugin, reload, and configure it, then use it to build a Skill where Claude Code handles reasoning and Codex acts as moderator.

Why it matters: HKR-H/K/R all pass: cross-stack plugin use is clickable, the install path is concrete, and it matters to AI dev workflows. It stays low-featured because this is a tutorial-style tip, not a model or platform release.

Xinzhiyuan · WeChat

OpenAI releases GPT-Realtime-2, described as a GPT-5-level reasoning audio model

OpenAI released GPT-Realtime-2 alongside Realtime-Translate and Realtime-Whisper, with a 128K context window, five reasoning-effort levels, and API pricing of $32 per million input tokens and $64 per million output tokens.

Why it matters: HKR-H/K/R all pass: realtime audio reasoning is a strong hook; 128K context, five reasoning levels, and $32/$64 per 1M tokens add substance; voice-agent cost and stack choices hit practitioners. This is a same-day OpenAI product update.

QbitAI · WeChat

Shanghai AI Lab Study: SFT Generalizes Under Three Conditions

Shanghai AI Lab, Shanghai Jiao Tong University, and USTC tested Long-CoT SFT on Qwen3-14B-Base and found that cross-domain performance recovered and improved after 8 epochs, with generalization conditioned on optimization depth, data quality and structure, and base-model capability.

Why it matters: HKR-H/K/R all pass: the SFT-generalization claim has a clear hook, Qwen3-14B-Base plus an 8-epoch finding, and direct relevance to fine-tuning teams. It lacks deployment impact or full benchmark detail, so it stays in the mid-featured band.