From Atari to EVE Online: Building on 15 Years of AI Research in Games
Google DeepMind 宣布与游戏开发商合作,用 AI 原型化全新游戏体验,并回顾了从 DQN 玩 49 款 Atari 游戏、AlphaGo、AlphaZero、MuZero、AlphaStar 到 SIMA 的游戏 AI 研究历程。
Google DeepMind 宣布与游戏开发商合作,用 AI 原型化全新游戏体验,并回顾了从 DQN 玩 49 款 Atari 游戏、AlphaGo、AlphaZero、MuZero、AlphaStar 到 SIMA 的游戏 AI 研究历程。
Google DeepMind published an AI Control Roadmap, a framework for building and managing advanced AI deployed inside Google. It takes a defense-in-depth approach, adding system-level safety layers on top of model alignment so protections hold even when alignment is imperfect.
Why it matters: DeepMind made its internal AI Control Roadmap public, laying out a layered way to monitor and block agents as if they were insider threats.
The experiment used five models from OpenAI, NVIDIA, OpenBMB, and a self-fine-tuned 500M-parameter model to drive market agents; three interventions failed to reproduce the price crash, and the crash was created only by overriding prices during settlement.
Why it matters: HKR-H/K/R all pass: the angle is counterintuitive, the post gives 5 models, 3 interventions, and a settlement override mechanism, and it speaks to agent-eval reliability. Scope remains an experiment blog, not a major release.
openJiuwen proposed MANGO, a multi-agent flow-network framework that combines reinforcement learning, textual gradients, and a Skip-k mechanism; using GPT-4o-mini, it reports a 12.8% accuracy gain over MaAS on MATH500 and a 5.1% F1 gain over AFlow on DROP.
Why it matters: HKR-K is strong: the post gives mechanisms and a MATH500 delta. HKR-H/R pass for the multi-agent flow-network angle, but this remains a research-framework story, not a major model or platform release.
UIUC and Chroma released Harness-1, a 20B-parameter retrieval subagent trained with reinforcement learning inside a stateful search harness, reporting 0.730 average curated recall across 8 benchmarks, 11.4 percentage points above the next-best open-source subagent and behind only Opus-4.6.
Why it matters: HKR-H/K/R all pass: Harness-1 has a clear RL retrieval-agent mechanism and benchmark numbers. It stays in 78–84 because this is a subagent research/open-source release, not a major lab model launch.
FusionRoute proposes a token-level multi-LLM collaboration method that freezes expert models and trains a lightweight router to select an expert for each token while merging router logits with expert logits. The paper evaluates it on GSM8K, MATH-500, HumanEval, MBPP, IfEval, and 500 PerfectBlend prompts.
Why it matters: HKR-H/K/R pass: token-level LLM routing is a strong research hook with concrete mechanics. The article lacks lift numbers, code link, and deployment cost, so it stays at the lower featured band.
Thousand Token Wood v2 uses four small models from different labs to drive agents in a financial simulation game, with vLLM 0.22.1’s CUDA toolkit dependency identified as the main serving friction, while a fine-tuned 0.5B Qwen reached 0% self-trading and 100% valid quotes.
Why it matters: HKR-H/K/R all pass: the small-model finance game is a real hook, with vLLM and 0.5B Qwen metrics, plus agent-engineering resonance. Scope remains an experiment, so it sits in low featured.
Princeton researchers released Goedel-Architect, an agent framework for Lean formal theorem proving. Using DeepSeek-V4-Flash, it reached 75.6% pass@1 on PutnamBench, with $294 in API cost for 672 problems, compared with Hilbert’s 70.0% and about $170,000 cost.
Why it matters: HKR-H/K/R all pass: Goedel-Architect pairs a 75.6% PutnamBench score with $294 for 672 problems, versus Hilbert at about $170k. It is still research-heavy, so it stays in the 78–84 band rather than P1.
A developer used Qwen2.5-3B to build a five-agent forest economy, and across 15 simulation rounds honey prices fell from 10 to 3, firewood rose from 4 to 7, and the Gini coefficient increased from 0.14 to 0.38.
Why it matters: HKR-H/K/R pass: the 3B multi-agent economy has a hook and concrete price/Gini results. It remains a single engineering experiment, not a product or framework launch, so it stays at the featured floor.
Google Research and Google Cloud introduced the Cross-Corpus Retrieval framework as Agentic RAG for Gemini Enterprise Agent Platform, using a multi-agent workflow to plan, rewrite, route, and iteratively search multiple data sources, with up to 34% higher accuracy than standard RAG on factual datasets.
Why it matters: HKR-H/K/R all pass: Google names a Cross-Corpus Retrieval mechanism and a +34% factual accuracy lift. The Gemini Enterprise Agent Platform tie-in adds cloud-vendor promo risk, so this stays below the 78–84 research/framework band.
CMU and the University of Maryland propose Language Models Need Sleep: when each L-token context window fills, the model runs N offline recurrent forward passes and updates SSM fast weights before evicting the KV cache. On GSM-Infinite, Jet-Nemotron 2B with 6 sleep loops improves 6-step arithmetic accuracy from 0.742 to 0.812.
Why it matters: HKR-H/K/R all pass: the hook is strong, and the post gives a testable mechanism plus Jet-Nemotron 2B numbers. It is still a single early paper, not an industry-level release, so it stays just above the featured threshold.
Anthropic published a post on recursive self-improvement under the title “When AI Builds Itself,” while the RSS body only discloses 95 Hacker News points and 106 comments, with no experimental setup, model details, or timeline disclosed.
Why it matters: HKR-H and HKR-R pass: an Anthropic post on recursive self-improvement has a strong hook and practitioner resonance. HKR-K fails because the feed discloses no mechanism or model details.
NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.
Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.
MAI filtered 265,000 trainable tasks from 4.87 million open-source PRs and built a three-layer judging system. The key change after vibe coding is the industrialization of training infrastructure.
Why it matters: HKR-H/K/R all pass via the post-vibe-coding angle, 4.87M PR corpus, 265K tasks, and code-agent infra stakes. No model scores, open-source scope, or product access are disclosed, so it stays below P1.
DataMaster searches, cleans, and combines data while keeping the model and training algorithm fixed; on MLE-Bench Lite, it raised the medal rate from 35.91% to 68.18%.
Why it matters: HKR-H/K/R all pass: DataMaster changes the data pipeline under fixed model and training code, lifting MLE-Bench Lite medal rate from 35.91% to 68.18%. This is still a single research release without production validation, so it lands at 78 featured.
Banafsheh Rafiee and Richard S. Sutton propose an enactive cognition framework for AI, naming four pillars: experience, perception-action inseparability, autonomy, and embodiment.
Why it matters: HKR-H/K/R all pass, but the article centers on a conceptual framework and does not disclose experiments, code, or reproducible tests. Sutton’s name and the four pillars put it in the 78–84 research-commentary band.
Renmin University Gaoling School of Artificial Intelligence released a 40-page survey on rubrics for LLMs, organizing the topic into five parts: definitions, construction methods, training uses, evaluation scenarios, and open challenges.
Why it matters: HKR-H/K/R all pass, but this is a survey rather than a model or product launch. The 40-page rubric framework is useful for agent evaluation, placing it at the featured threshold.
Fudan University and Tongyi Lab introduced ToolCUA-8B, which reaches 46.85% accuracy on OSWorld-MCP after training with about 4k synthetic tools and 180k interleaved GUI-Tool trajectory steps.
Why it matters: HKR-H/K/R all pass: the tool-selection failure hook is concrete, with OSWorld-MCP 46.85% and 180k steps. It stays in the 78–84 band because this is a research release, not a major model or product launch.
NVIDIA, Tsinghua, University of Toronto, and Vector Institute released Gamma-World, a multi-agent world model using simplex-based positional encoding and hub tokens to cut interaction cost from quadratic to linear, with 8-player latency dropping from 17.6 ms to 4.5 ms.
Why it matters: HKR-H/K/R all pass: Gamma-World has a concrete mechanism and latency claim from NVIDIA/Tsinghua. Scope remains multi-agent world-model research, so it sits in the 78–84 good-quality band rather than must-write.
Meta released ATLAS, a Lean 4 formalization library covering 26 math textbooks and 46,203 declarations, using 183.157 billion tokens to generate 630,999 lines of code, with 42,837 completed proofs and a 92.7% proof pass rate.
Why it matters: HKR-H/K/R all pass: the token scale, Lean corpus size, and verified-proof count are concrete. It stays below P1 because this is a specialized research/open-source release, not a broad model or product launch.
Cursor’s report says developers’ weekly code output rose from about 3.6K to 8.6K lines, while AI agents increased tool calls per session by roughly 30%.
Why it matters: HKR-H/K/R all pass: Cursor’s own report gives concrete 3.6K→8.6K and +30% figures for AI coding work. It is not a product launch or cross-source event, so 78–84 fits better than the must-write band.
AutoWebWorld synthesized 29 web environments, 875 pages, and 11,663 verified trajectories at about $0.04 per trajectory, using FSM-defined states, preconditions, and transitions to verify GUI agent tasks instead of human labeling or an LLM judge.
Why it matters: HKR-H/K/R all pass: $0.04 per trace, 11,663 verified traces, and FSM state checks give concrete hooks for GUI-agent data and eval cost. The source is not a top lab release, so it stays in the 78–84 research-tool band.
NVIDIA Research presented 8 ICRA papers on sim-to-real robotics: ScheduleStream delivered a 3x speedup for multi-arm planning, COMPASS reached about 80% success across 20 real-world navigation trials, and Grasp-MPC achieved about 75% real-robot grasping success.
Why it matters: HKR-K and HKR-R are strong: the post gives concrete sim-to-real numbers from ICRA and addresses robot deployment reliability. HKR-H is moderate but passes on the real-world success-rate hook.
NTU AutoMan Lab, Harvard, and Xiaomi Auto proposed AutoMoT, a unified VLA driving model using a 4B Qwen3-VL Understanding Expert and a 1.6B Action Expert with asynchronous inference, reaching 89.42 DS and 74.09% SR on Bench2Drive with AutoMoT+.
Why it matters: HKR-H/K pass via the async VLM-driving setup and concrete Bench2Drive numbers. The autonomy focus narrows HKR-R, so this sits at the featured threshold rather than the 78+ research tier.
NVIDIA’s research team open-sourced Polar, an agent reinforcement learning framework that connects GRPO training at the model API boundary without rewriting Codex CLI, Claude Code, Qwen Code, or Pi; on Qwen3.5-4B, Polar raised Codex pass@1 on SWE-Bench Verified from 3.8% to 26.4%, while prefix_merging cut training steps from 1,185 to 218.
Why it matters: HKR-H/K/R all pass: NVIDIA open-sourced Polar with a concrete GRPO mechanism and SWE-Bench Verified numbers. This is a strong research/open-source item, not a major model or product release, so it stays in the 78–84 band.
Firespawn Studios ran 25 agents across 8 open-weight models for 10 days in Null Epoch Season 0 and released about 93,000 logged events, with roughly 70% of actions including the model’s reasoning or justification.
Why it matters: HKR-H/K/R all pass: a concrete 10-day MMO agent trial with 25 agents and 93k events. Reddit sourcing limits reach, so it lands in the 78–84 good-quality band, not P1.
Shanghai Innovation Institute’s LeapQuest and three universities released Ophiuchus and MedScope, applying Think with Images and Think with Videos to medical AI; Ophiuchus-7B scored 68.0 on eight VQA benchmarks, above OpenAI-o3 at 62.2, Gemini 2.5 Pro at 61.8, and GPT-5 at 59.9.
Why it matters: HKR-H/K/R all pass: a 7B model beating o3/GPT-5 is a strong hook, 8 VQA benchmarks with 68.0 vs 62.2 add a testable claim, and medical specialist evaluation will trigger debate. Not a frontier-lab general model release, so it stays in 78–84.
Carnegie Mellon University and the University of Maryland propose a “sleep” mechanism for language models: when the context window is nearly full, the model stops accepting new tokens, runs multiple offline recursive forward passes to compress accumulated context into fast weights, clears the KV cache, and then resumes inference; tests cover cellular automata, multi-hop graph retrieval, and GSM-Infinite reasoning tasks.
Why it matters: HKR-H/K/R all pass: the sleep metaphor is clickable, and the mechanism is concrete. Score stays below 78 because the provided body lacks benchmark gains, code, or deployment evidence.
Samsung disclosed three AI efforts—Meki, M2RL, and LiveClawBench—covering a memory-based edge architecture, multi-domain reinforcement learning, and Physical AI evaluation; the article also says Samsung has purchased tens of thousands of GPUs for AI infrastructure, but does not provide model size, training budget, or deployment timelines.
Why it matters: HKR-H, HKR-K, and HKR-R pass, but this is a Samsung research bundle plus strategy signal, not a flagship model or product launch. It fits the 72–77 featured band, below same-day must-write.
SkillOpt uses a frontier model to propose add, delete, and replace edits to markdown skill files, then accepts only strict gains on a held-out validation set; the best skills usually converge after 1 to 4 accepted edits.
Why it matters: HKR-H/K/R all pass: the hook is trainable markdown skills, with held-out validation and 1-4 accepted edits. Single Reddit/project source and no broad adoption data keep it at 78, featured not p1.
Spatial-Agent inserts a GeoFlow Graph between natural-language questions and map tools, and Spatial-Agent with GPT-4o-mini reaches 45.15% accuracy on MapEval-API versus a 23.00% API baseline.
Why it matters: ACL Main gives a concrete mechanism and testable numbers, so HKR-H/K pass. The GIS focus limits HKR-R, placing it at the featured threshold rather than a must-write item.
Verkor’s Design Conductor generated an ASAP7 7nm GDSII layout for the VerCore RISC-V CPU from a 219-word English spec in 12 hours, with no engineer in the design loop; the reported result scored 3,261 CoreMark at 1.48GHz, but it has not been fabricated and lacks cache implementation.
Why it matters: HKR-H/K/R all pass, but VerCore is not taped out and lacks cache, so the claim stays at demo-and-benchmark level. Concrete numbers and test conditions put it in the 78–84 recommendation band.
Anthropic says Project Glasswing used Claude Mythos Preview with about 50 partners to find more than 10,000 high or critical vulnerabilities in global critical systems, with independently verified accuracy of 90.6%.
Why it matters: HKR-H/K/R all pass: Anthropic gives concrete numbers—~50 partners, 10,000+ high/critical bugs, 90.6% validation—and the story hits AI-agent security automation and critical-system risk.
Researchers from The Hong Kong Polytechnic University and HKUST (Guangzhou) introduced ULSPB with 350 settings; routine conversations can poison an agent’s long-term state without malicious prompts, while StateGuard audits state diffs before persistence and reduces Harm Score to near zero in Targeted-Ensemble settings.
Why it matters: HKR-H/K/R all pass: the story has a non-malicious agent corruption hook and a concrete ULSPB benchmark with 350 settings. It is useful agent-safety research, not a top-lab product release.
Westlake University and collaborators introduced HiF-VLA, a motion-centric VLA framework that extracts compact Motion vectors with codecs such as H.264 and uses a joint expert to predict future visual motion and generate action sequences, reporting 31.4GB peak memory and 117.7ms latency under the cited history-window setting.
Why it matters: HKR-H/K/R all pass: the H.264-motion angle, concrete VRAM/latency numbers, and robotics deployment pressure are clear. It remains a single research item without adoption or cross-source heat, so it sits in the lower featured band.
Google Research published its Gemini-based Empirical Research Assistant in Nature and opened early access through the Google Labs trusted tester program.
Why it matters: HKR-H/K/R all pass: Google moves Gemini-based ERA from a Nature paper to a Labs trusted-tester trial. Score stays at 78 because the provided text lacks metrics, benchmark setup, or reproducible workflow details.
CUHK and Zhejiang University researchers argue that mainstream Agent memory is retrieval-based memo storage, not true memory, citing an Ω(k²) case requirement for compositional tasks and a PoisonedRAG result where 5 adversarial texts reached a 90% attack success rate.
Why it matters: HKR-H/K/R all pass: the hook is concrete, the summary gives Ω(k²) and 90% attack success, and the issue matters to agent-memory and RAG-security builders. Strong research signal, not a same-day model-release event.
Jiqizhixin translated Sebastian Raschka’s blog on recent LLM architecture changes, covering long-context cost reductions in Gemma 4, Laguna XS.2, and ZAYA1-8B; the article states that Gemma 4 E2B saves about 2.7GB of KV cache at 128K context with bfloat16 precision.
Why it matters: HKR-H/K/R pass: notable model names, a concrete 128K bf16 KV-cache saving, and inference-cost relevance. As a translated survey rather than a release, it stays in the 72–77 featured band.
Odyssey Labs released Agora-1, described as the first real-time multi-agent world model, using a GoldenEye deathmatch demo where multiple humans and AI agents interact in the same simulated world; the post says a playable research preview is available now, but does not disclose model architecture or latency figures.
Why it matters: HKR-H/K/R all pass: Agora-1 combines multi-agent world modeling with a live human-AI preview. Sparse details on architecture, latency, cost, and benchmarks keep it in the 78–84 band.
EvolveR lets agents distill reusable experience from successful and failed trajectories, maintain a scored experience library, and train retrieval behavior with GRPO; the paper reports the best average performance on seven complex QA benchmarks using Qwen2.5-3B and 7B.
Why it matters: HKR-H/K/R all pass: the agent self-growing-skill angle is clickable, with mechanism and benchmark specifics. Since only a media summary is available and no repo, absolute scores, or reproduction details are disclosed, it stays in the 78–84 research band.