Skip to content

#Agent

0 today

Aug 21Friday

Jun 16Tuesday

Google DeepMind

Google DeepMind publishes AI Control Roadmap for internal AI agents

Google DeepMind published an AI Control Roadmap, a framework for building and managing advanced AI deployed inside Google. It takes a defense-in-depth approach, adding system-level safety layers on top of model alignment so protections hold even when alignment is imperfect.

Why it matters: DeepMind made its internal AI Control Roadmap public, laying out a layered way to monitor and block agents as if they were insider threats.

Jun 8Monday

AI HOT (Curated Pool)

The Vanishing Crash in Five-Model Economies: Control and Emergence

The experiment used five models from OpenAI, NVIDIA, OpenBMB, and a self-fine-tuned 500M-parameter model to drive market agents; three interventions failed to reproduce the price crash, and the crash was created only by overriding prices during settlement.

Why it matters: HKR-H/K/R all pass: the angle is counterintuitive, the post gives 5 models, 3 interventions, and a settlement override mechanism, and it speaks to agent-eval reliability. Scope remains an experiment blog, not a major release.

Synced · WeChat

openJiuwen proposes MANGO for multi-agent flow networks

openJiuwen proposed MANGO, a multi-agent flow-network framework that combines reinforcement learning, textual gradients, and a Skip-k mechanism; using GPT-4o-mini, it reports a 12.8% accuracy gain over MaAS on MATH500 and a 5.1% F1 gain over AFlow on DROP.

Why it matters: HKR-K is strong: the post gives mechanisms and a MATH500 delta. HKR-H/R pass for the multi-agent flow-network angle, but this remains a research-framework story, not a major model or platform release.

Jun 7Sunday

AI HOT (Curated Pool)

Harness-1: A 20B Stateful Retrieval Subagent Trained with Reinforcement Learning

UIUC and Chroma released Harness-1, a 20B-parameter retrieval subagent trained with reinforcement learning inside a stateful search harness, reporting 0.730 average curated recall across 8 benchmarks, 11.4 percentage points above the next-best open-source subagent and behind only Opus-4.6.

Why it matters: HKR-H/K/R all pass: Harness-1 has a clear RL retrieval-agent mechanism and benchmark numbers. It stays in 78–84 because this is a subagent research/open-source release, not a major lab model launch.

Synced · WeChat

ICML 2026 | FusionRoute: From Expert Routing to Self-Correction in Multi-LLM Collaboration

FusionRoute proposes a token-level multi-LLM collaboration method that freezes expert models and trains a lightweight router to select an expert for each token while merging router logits with expert logits. The paper evaluates it on GSM8K, MATH-500, HumanEval, MBPP, IfEval, and 500 PerfectBlend prompts.

Why it matters: HKR-H/K/R pass: token-level LLM routing is a strong research hook with concrete mechanics. The article lacks lift numbers, code link, and deployment cost, so it stays at the lower featured band.

AI HOT (Curated Pool)

Five Labs, Five Minds: Building a Multi-Model Financial Drama Game with Small Models

Thousand Token Wood v2 uses four small models from different labs to drive agents in a financial simulation game, with vLLM 0.22.1’s CUDA toolkit dependency identified as the main serving friction, while a fine-tuned 0.5B Qwen reached 0% self-trading and 100% valid quotes.

Why it matters: HKR-H/K/R all pass: the small-model finance game is a real hook, with vLLM and 0.5B Qwen metrics, plus agent-engineering resonance. Scope remains an experiment, so it sits in low featured.

Jun 6Saturday

Synced · WeChat

DeepSeek V4 Proves Math with 500x Cost Advantage as Agent System Sets Records

Princeton researchers released Goedel-Architect, an agent framework for Lean formal theorem proving. Using DeepSeek-V4-Flash, it reached 75.6% pass@1 on PutnamBench, with $294 in API cost for 672 problems, compared with Hilbert’s 70.0% and about $170,000 cost.

Why it matters: HKR-H/K/R all pass: Goedel-Architect pairs a 75.6% PutnamBench score with $294 for 672 problems, versus Hilbert at about $170k. It is still research-heavy, so it stays in the 78–84 band rather than P1.

AI HOT (Curated Pool)

Building a Multi-Agent Economy with Qwen2.5-3B: Engineering Report

A developer used Qwen2.5-3B to build a five-agent forest economy, and across 15 simulation rounds honey prices fell from 10 to 3, firewood rose from 4 to 7, and the Gini coefficient increased from 0.14 to 0.38.

Why it matters: HKR-H/K/R pass: the 3B multi-agent economy has a hook and concrete price/Gini results. It remains a single engineering experiment, not a product or framework launch, so it stays at the featured floor.

AI HOT (Curated Pool)

Google launches Agentic RAG framework for Gemini Enterprise Agent Platform

Google Research and Google Cloud introduced the Cross-Corpus Retrieval framework as Agentic RAG for Gemini Enterprise Agent Platform, using a multi-agent workflow to plan, rewrite, route, and iteratively search multiple data sources, with up to 34% higher accuracy than standard RAG on factual datasets.

Why it matters: HKR-H/K/R all pass: Google names a Cross-Corpus Retrieval mechanism and a +34% factual accuracy lift. The Gemini Enterprise Agent Platform tie-in adds cloud-vendor promo risk, so this stays below the 78–84 research/framework band.

Jun 5Friday

Synced · WeChat

Do Models Need Sleep? CMU Paper Lets LLMs Consolidate Memory During “Sleep”

CMU and the University of Maryland propose Language Models Need Sleep: when each L-token context window fills, the model runs N offline recurrent forward passes and updates SSM fast weights before evicting the KV cache. On GSM-Infinite, Jet-Nemotron 2B with 6 sleep loops improves 6-step arithmetic accuracy from 0.742 to 0.812.

Why it matters: HKR-H/K/R all pass: the hook is strong, and the post gives a testable mechanism plus Jet-Nemotron 2B numbers. It is still a single early paper, not an industry-level release, so it stays just above the featured threshold.

Hacker News front page

When AI Builds Itself: Our Progress Toward Recursive Self-Improvement

Anthropic published a post on recursive self-improvement under the title “When AI Builds Itself,” while the RSS body only discloses 95 Hacker News points and 106 comments, with no experimental setup, model details, or timeline disclosed.

Why it matters: HKR-H and HKR-R pass: an Anthropic post on recursive self-improvement has a strong hook and practitioner resonance. HKR-K fails because the feed discloses no mechanism or model details.

Jun 3Wednesday

NVIDIA Blog

NVIDIA Research Presents Grasping, Autonomous Driving and Agent Training Work at CVPR

NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.

Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.

Computing Life · Share · Yage

After vibe coding: the industrialization of AI programming

MAI filtered 265,000 trainable tasks from 4.87 million open-source PRs and built a three-layer judging system. The key change after vibe coding is the industrialization of training infrastructure.

Why it matters: HKR-H/K/R all pass via the post-vibe-coding angle, 4.87M PR corpus, 265K tasks, and code-agent infra stakes. No model scores, open-source scope, or product access are disclosed, so it stays below P1.

Jun 2Tuesday

Synced · WeChat

DataMaster: When AI Becomes Its Own Data Engineer

DataMaster searches, cleans, and combines data while keeping the model and training algorithm fixed; on MLE-Bench Lite, it raised the medal rate from 35.91% to 68.18%.

Why it matters: HKR-H/K/R all pass: DataMaster changes the data pipeline under fixed model and training code, lifting MLE-Bench Lite medal rate from 35.91% to 68.18%. This is still a single research release without production validation, so it lands at 78 featured.

Synced · WeChat

Turing Award Winner Sutton’s New Paper Argues AI Should Move Toward Enactive Cognition

Banafsheh Rafiee and Richard S. Sutton propose an enactive cognition framework for AI, naming four pillars: experience, perception-action inseparability, autonomy, and embodiment.

Why it matters: HKR-H/K/R all pass, but the article centers on a conceptual framework and does not disclose experiments, code, or reproducible tests. Sutton’s name and the four pillars put it in the 78–84 research-commentary band.

May 31Sunday

Synced · WeChat

Rubrics Survey: How to Define a Good Answer in the Agent Era

Renmin University Gaoling School of Artificial Intelligence released a 40-page survey on rubrics for LLMs, organizing the topic into five parts: definitions, construction methods, training uses, evaluation scenarios, and open challenges.

Why it matters: HKR-H/K/R all pass, but this is a survey rather than a model or product launch. The 40-page rubric framework is useful for agent evaluation, placing it at the featured threshold.

QbitAI · WeChat

Fudan and Tongyi introduce ToolCUA for GUI-Tool path selection in agents

Fudan University and Tongyi Lab introduced ToolCUA-8B, which reaches 46.85% accuracy on OSWorld-MCP after training with about 4k synthetic tools and 180k interleaved GUI-Tool trajectory steps.

Why it matters: HKR-H/K/R all pass: the tool-selection failure hook is concrete, with OSWorld-MCP 46.85% and 180k steps. It stays in the 78–84 band because this is a research release, not a major model or product launch.

May 30Saturday

Synced · WeChat

NVIDIA and Tsinghua Team's Gamma-World Tops Hugging Face Daily Chart

NVIDIA, Tsinghua, University of Toronto, and Vector Institute released Gamma-World, a multi-agent world model using simplex-based positional encoding and hub tokens to cut interaction cost from quadratic to linear, with 8-player latency dropping from 17.6 ms to 4.5 ms.

Why it matters: HKR-H/K/R all pass: Gamma-World has a concrete mechanism and latency claim from NVIDIA/Tsinghua. Scope remains multi-agent world-model research, so it sits in the 78–84 good-quality band rather than must-write.

May 29Friday

Synced · WeChat

Meta Uses 183B Tokens to Turn Math Textbooks into a Large Lean Library

Meta released ATLAS, a Lean 4 formalization library covering 26 math textbooks and 46,203 declarations, using 183.157 billion tokens to generate 630,999 lines of code, with 42,837 completed proofs and a 92.7% proof pass rate.

Why it matters: HKR-H/K/R all pass: the token scale, Lean corpus size, and verified-proof count are concrete. It stays below P1 because this is a specialized research/open-source release, not a broad model or product launch.

AI HOT (Curated Pool)

Cursor team releases Developer Habits Report

Cursor’s report says developers’ weekly code output rose from about 3.6K to 8.6K lines, while AI agents increased tool calls per session by roughly 30%.

Why it matters: HKR-H/K/R all pass: Cursor’s own report gives concrete 3.6K→8.6K and +30% figures for AI coding work. It is not a product launch or cross-source event, so 78–84 fits better than the must-write band.

May 28Thursday

QbitAI · WeChat

A New Paradigm for GUI Agent Trajectories: FSMs Generate Trajectories at $0.04 Each

AutoWebWorld synthesized 29 web environments, 875 pages, and 11,663 verified trajectories at about $0.04 per trajectory, using FSM-defined states, preconditions, and transitions to verify GUI agent tasks instead of human labeling or an LLM judge.

Why it matters: HKR-H/K/R all pass: $0.04 per trace, 11,663 verified traces, and FSM state checks give concrete hooks for GUI-agent data and eval cost. The source is not a top lab release, so it stays in the 78–84 research-tool band.

NVIDIA Blog

NVIDIA Research Advances Robotics From Simulation to the Real World

NVIDIA Research presented 8 ICRA papers on sim-to-real robotics: ScheduleStream delivered a 3x speedup for multi-arm planning, COMPASS reached about 80% success across 20 real-world navigation trials, and Grasp-MPC achieved about 75% real-robot grasping success.

Why it matters: HKR-K and HKR-R are strong: the post gives concrete sim-to-real numbers from ICRA and addresses robot deployment reliability. HKR-H is moderate but passes on the real-world success-rate hook.

Synced · WeChat

ICML 2026: AutoMoT reaches SOTA on Bench2Drive and nuScenes

NTU AutoMan Lab, Harvard, and Xiaomi Auto proposed AutoMoT, a unified VLA driving model using a 4B Qwen3-VL Understanding Expert and a 1.6B Action Expert with asynchronous inference, reaching 89.42 DS and 74.09% SR on Bench2Drive with AutoMoT+.

Why it matters: HKR-H/K pass via the async VLM-driving setup and concrete Bench2Drive numbers. The autonomy focus narrows HKR-R, so this sits at the featured threshold rather than the 78+ research tier.

AI HOT (Curated Pool)

NVIDIA Releases AI Framework Polar, Raising Codex Benchmark Score by 594.74%

NVIDIA’s research team open-sourced Polar, an agent reinforcement learning framework that connects GRPO training at the model API boundary without rewriting Codex CLI, Claude Code, Qwen Code, or Pi; on Qwen3.5-4B, Polar raised Codex pass@1 on SWE-Bench Verified from 3.8% to 26.4%, while prefix_merging cut training steps from 1,185 to 218.

Why it matters: HKR-H/K/R all pass: NVIDIA open-sourced Polar with a concrete GRPO mechanism and SWE-Bench Verified numbers. This is a strong research/open-source item, not a major model or product release, so it stays in the 78–84 band.

May 27Wednesday

r/LocalLLaMA

I ran 8 open-weight models as agents in a persistent MMO for 10 days

Firespawn Studios ran 25 agents across 8 open-weight models for 10 days in Null Epoch Season 0 and released about 93,000 logged events, with roughly 70% of actions including the model’s reasoning or justification.

Why it matters: HKR-H/K/R all pass: a concrete 10-day MMO agent trial with 25 agents and 93k events. Reddit sourcing limits reach, so it lands in the 78–84 good-quality band, not P1.

QbitAI · WeChat

7B Medical AI Agent Beats o3 and GPT-5 by Learning Where and How to Look

Shanghai Innovation Institute’s LeapQuest and three universities released Ophiuchus and MedScope, applying Think with Images and Think with Videos to medical AI; Ophiuchus-7B scored 68.0 on eight VQA benchmarks, above OpenAI-o3 at 62.2, Gemini 2.5 Pro at 61.8, and GPT-5 at 59.9.

Why it matters: HKR-H/K/R all pass: a 7B model beating o3/GPT-5 is a strong hook, 8 VQA benchmarks with 68.0 vs 62.2 add a testable claim, and medical specialist evaluation will trigger debate. Not a frontier-lab general model release, so it stays in 78–84.

QbitAI · WeChat

Language Models Need Sleep: Let AI Nap Before Continuing Inference

Carnegie Mellon University and the University of Maryland propose a “sleep” mechanism for language models: when the context window is nearly full, the model stops accepting new tokens, runs multiple offline recursive forward passes to compress accumulated context into fast weights, clears the KV cache, and then resumes inference; tests cover cellular automata, multi-hop graph retrieval, and GSM-Infinite reasoning tasks.

Why it matters: HKR-H/K/R all pass: the sleep metaphor is clickable, and the mechanism is concrete. Score stays below 78 because the provided body lacks benchmark gains, code, or deployment evidence.

Synced · WeChat

From Foundation Models to Physical AI, Samsung Moves Into the Core LLM Race

Samsung disclosed three AI efforts—Meki, M2RL, and LiveClawBench—covering a memory-based edge architecture, multi-domain reinforcement learning, and Physical AI evaluation; the article also says Samsung has purchased tens of thousands of GPUs for AI infrastructure, but does not provide model size, training budget, or deployment timelines.

Why it matters: HKR-H, HKR-K, and HKR-R pass, but this is a Samsung research bundle plus strategy signal, not a flagship model or product launch. It fits the 72–77 featured band, below same-day must-write.

May 26Tuesday

r/LocalLLaMA

SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery

SkillOpt uses a frontier model to propose add, delete, and replace edits to markdown skill files, then accepts only strict gains on a held-out validation set; the best skills usually converge after 1 to 4 accepted edits.

Why it matters: HKR-H/K/R all pass: the hook is trainable markdown skills, with held-out validation and 1-4 accepted edits. Single Reddit/project source and no broad adoption data keep it at 78, featured not p1.

Synced · WeChat

ACL 2026 Main: Spatial-Agent Generates Executable Geospatial Analysis Workflows for LLMs

Spatial-Agent inserts a GeoFlow Graph between natural-language questions and map tools, and Spatial-Agent with GPT-4o-mini reaches 45.15% accuracy on MapEval-API versus a 23.00% API baseline.

Why it matters: ACL Main gives a concrete mechanism and testable numbers, so HKR-H/K pass. The GIS focus limits HKR-R, placing it at the featured threshold rather than a must-write item.

May 24Sunday

Xinzhiyuan · WeChat

AI Agent Completes Chip Design from 219 Words to 7nm GDSII Without Engineer Input

Verkor’s Design Conductor generated an ASAP7 7nm GDSII layout for the VerCore RISC-V CPU from a 219-word English spec in 12 hours, with no engineer in the design loop; the reported result scored 3,261 CoreMark at 1.48GHz, but it has not been fabricated and lacks cache implementation.

Why it matters: HKR-H/K/R all pass, but VerCore is not taped out and lacks cache, so the claim stays at demo-and-benchmark level. Concrete numbers and test conditions put it in the 78–84 recommendation band.

May 23Saturday

AI HOT (Curated Pool)

Project Glasswing: Initial Update

Anthropic says Project Glasswing used Claude Mythos Preview with about 50 partners to find more than 10,000 high or critical vulnerabilities in global critical systems, with independently verified accuracy of 90.6%.

Why it matters: HKR-H/K/R all pass: Anthropic gives concrete numbers—~50 partners, 10,000+ high/critical bugs, 90.6% validation—and the story hits AI-agent security automation and critical-system risk.

May 22Friday

Xinzhiyuan · WeChat

OpenClaw Case: Routine Chats Can Poison an Agent’s Long-Term Memory

Researchers from The Hong Kong Polytechnic University and HKUST (Guangzhou) introduced ULSPB with 350 settings; routine conversations can poison an agent’s long-term state without malicious prompts, while StateGuard audits state diffs before persistence and reduces Harm Score to near zero in Targeted-Ensemble settings.

Why it matters: HKR-H/K/R all pass: the story has a non-malicious agent corruption hook and a concrete ULSPB benchmark with 350 settings. It is useful agent-safety research, not a top-lab product release.

Synced · WeChat

CVPR 2026 | HiF-VLA: A Motion-Centric World Action Model

Westlake University and collaborators introduced HiF-VLA, a motion-centric VLA framework that extracts compact Motion vectors with codecs such as H.264 and uses a joint expert to predict future visual motion and generate action sequences, reporting 31.4GB peak memory and 117.7ms latency under the cited history-window setting.

Why it matters: HKR-H/K/R all pass: the H.264-motion angle, concrete VRAM/latency numbers, and robotics deployment pressure are clear. It remains a single research item without adoption or cross-source heat, so it sits in the lower featured band.

May 20Wednesday

AI HOT (Curated Pool)

Empirical Research Assistant ERA: From Nature Publication to Computational Discovery

Google Research published its Gemini-based Empirical Research Assistant in Nature and opened early access through the Google Labs trusted tester program.

Why it matters: HKR-H/K/R all pass: Google moves Gemini-based ERA from a Nature paper to a Labs trusted-tester trial. Score stays at 78 because the provided text lacks metrics, benchmark setup, or reproducible workflow details.

May 19Tuesday

Xinzhiyuan · WeChat

CUHK and Zhejiang University Question Whether AI Agent Memory Is Just a Memo

CUHK and Zhejiang University researchers argue that mainstream Agent memory is retrieval-based memo storage, not true memory, citing an Ω(k²) case requirement for compositional tasks and a PoisonedRAG result where 5 adversarial texts reached a 90% attack success rate.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the summary gives Ω(k²) and 90% attack success, and the issue matters to agent-memory and RAG-security builders. Strong research signal, not a same-day model-release event.

Synced · WeChat

Recent LLM Architecture Changes: From Gemma 4 to DeepSeek V4

Jiqizhixin translated Sebastian Raschka’s blog on recent LLM architecture changes, covering long-context cost reductions in Gemma 4, Laguna XS.2, and ZAYA1-8B; the article states that Gemma 4 E2B saves about 2.7GB of KV cache at 128K context with bfloat16 precision.

Why it matters: HKR-H/K/R pass: notable model names, a concrete 128K bf16 KV-cache saving, and inference-cost relevance. As a translated survey rather than a release, it stays in the 72–77 featured band.

AI HOT (Curated Pool)

First real-time multi-agent world model released, humans interact with AI on the same screen

Odyssey Labs released Agora-1, described as the first real-time multi-agent world model, using a GoldenEye deathmatch demo where multiple humans and AI agents interact in the same simulated world; the post says a playable research preview is available now, but does not disclose model architecture or latency figures.

Why it matters: HKR-H/K/R all pass: Agora-1 combines multi-agent world modeling with a live human-AI preview. Sparse details on architecture, latency, cost, and benchmarks keep it in the 78–84 band.

May 18Monday

QbitAI · WeChat

Agents Learn to Grow Skills from Failure: EvolveR Accepted by ICML 2026

EvolveR lets agents distill reusable experience from successful and failed trajectories, maintain a scored experience library, and train retrieval behavior with GRPO; the paper reports the best average performance on seven complex QA benchmarks using Qwen2.5-3B and 7B.

Why it matters: HKR-H/K/R all pass: the agent self-growing-skill angle is clickable, with mechanism and benchmark specifics. Since only a media summary is available and no repo, absolute scores, or reproduction details are disclosed, it stays in the 78–84 research band.