Skip to content

#MCP/工具调用

0 today

Jun 6Saturday

AI HOT (Curated Pool)

Building a Multi-Agent Economy with Qwen2.5-3B: Engineering Report

A developer used Qwen2.5-3B to build a five-agent forest economy, and across 15 simulation rounds honey prices fell from 10 to 3, firewood rose from 4 to 7, and the Gini coefficient increased from 0.14 to 0.38.

Why it matters: HKR-H/K/R pass: the 3B multi-agent economy has a hook and concrete price/Gini results. It remains a single engineering experiment, not a product or framework launch, so it stays at the featured floor.

Jun 2Tuesday

Synced · WeChat

DataMaster: When AI Becomes Its Own Data Engineer

DataMaster searches, cleans, and combines data while keeping the model and training algorithm fixed; on MLE-Bench Lite, it raised the medal rate from 35.91% to 68.18%.

Why it matters: HKR-H/K/R all pass: DataMaster changes the data pipeline under fixed model and training code, lifting MLE-Bench Lite medal rate from 35.91% to 68.18%. This is still a single research release without production validation, so it lands at 78 featured.

May 31Sunday

QbitAI · WeChat

Fudan and Tongyi introduce ToolCUA for GUI-Tool path selection in agents

Fudan University and Tongyi Lab introduced ToolCUA-8B, which reaches 46.85% accuracy on OSWorld-MCP after training with about 4k synthetic tools and 180k interleaved GUI-Tool trajectory steps.

Why it matters: HKR-H/K/R all pass: the tool-selection failure hook is concrete, with OSWorld-MCP 46.85% and 180k steps. It stays in the 78–84 band because this is a research release, not a major model or product launch.

May 29Friday

AI HOT (Curated Pool)

Cursor team releases Developer Habits Report

Cursor’s report says developers’ weekly code output rose from about 3.6K to 8.6K lines, while AI agents increased tool calls per session by roughly 30%.

Why it matters: HKR-H/K/R all pass: Cursor’s own report gives concrete 3.6K→8.6K and +30% figures for AI coding work. It is not a product launch or cross-source event, so 78–84 fits better than the must-write band.

May 28Thursday

QbitAI · WeChat

A New Paradigm for GUI Agent Trajectories: FSMs Generate Trajectories at $0.04 Each

AutoWebWorld synthesized 29 web environments, 875 pages, and 11,663 verified trajectories at about $0.04 per trajectory, using FSM-defined states, preconditions, and transitions to verify GUI agent tasks instead of human labeling or an LLM judge.

Why it matters: HKR-H/K/R all pass: $0.04 per trace, 11,663 verified traces, and FSM state checks give concrete hooks for GUI-agent data and eval cost. The source is not a top lab release, so it stays in the 78–84 research-tool band.

May 26Tuesday

r/LocalLLaMA

SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery

SkillOpt uses a frontier model to propose add, delete, and replace edits to markdown skill files, then accepts only strict gains on a held-out validation set; the best skills usually converge after 1 to 4 accepted edits.

Why it matters: HKR-H/K/R all pass: the hook is trainable markdown skills, with held-out validation and 1-4 accepted edits. Single Reddit/project source and no broad adoption data keep it at 78, featured not p1.

Synced · WeChat

ACL 2026 Main: Spatial-Agent Generates Executable Geospatial Analysis Workflows for LLMs

Spatial-Agent inserts a GeoFlow Graph between natural-language questions and map tools, and Spatial-Agent with GPT-4o-mini reaches 45.15% accuracy on MapEval-API versus a 23.00% API baseline.

Why it matters: ACL Main gives a concrete mechanism and testable numbers, so HKR-H/K pass. The GIS focus limits HKR-R, placing it at the featured threshold rather than a must-write item.

May 24Sunday

Xinzhiyuan · WeChat

AI Agent Completes Chip Design from 219 Words to 7nm GDSII Without Engineer Input

Verkor’s Design Conductor generated an ASAP7 7nm GDSII layout for the VerCore RISC-V CPU from a 219-word English spec in 12 hours, with no engineer in the design loop; the reported result scored 3,261 CoreMark at 1.48GHz, but it has not been fabricated and lacks cache implementation.

Why it matters: HKR-H/K/R all pass, but VerCore is not taped out and lacks cache, so the claim stays at demo-and-benchmark level. Concrete numbers and test conditions put it in the 78–84 recommendation band.

May 20Wednesday

AI HOT (Curated Pool)

Empirical Research Assistant ERA: From Nature Publication to Computational Discovery

Google Research published its Gemini-based Empirical Research Assistant in Nature and opened early access through the Google Labs trusted tester program.

Why it matters: HKR-H/K/R all pass: Google moves Gemini-based ERA from a Nature paper to a Labs trusted-tester trial. Score stays at 78 because the provided text lacks metrics, benchmark setup, or reproducible workflow details.

May 17Sunday

AI HOT (Curated Pool)

Study on the Cognition–Action Disconnect in Tool-Using Agents

An interpretability paper studies tool-using agents and finds models often recognize when to call a tool but fail to act, with a cognition-to-action mismatch rate of 26%–54%.

Why it matters: HKR-H/K/R all pass: the story has a sharp agent-failure hook, a 26%-54% mismatch rate, and clear relevance to tool-use reliability. Source detail is thin, with paper name, models, and task setup not disclosed.

May 15Friday

Xinzhiyuan · WeChat

Hassabis Praises Google DeepMind's AI-enabled Pointer Powered by Gemini

Google DeepMind released a Gemini-powered AI-enabled pointer and opened two demos in Google AI Studio: image editing and place finding on maps, while the post says Chrome pointer selection and a Googlebook Magic Pointer are planned product paths.

Why it matters: HKR-H/K/R all pass: the prompt-free pointer is clickable, the two AI Studio demos add concrete facts, and UI replacement resonates. Scope is still demo-level, with no metrics or API details, so 78 not 85+.

r/LocalLLaMA

inclusionAI/Ring-2.6-1T on Hugging Face

inclusionAI released Ring-2.6-1T, a 1T-parameter reasoning model on Hugging Face; it supports high and xhigh reasoning effort levels, targets agent workflows and long-horizon tasks, and uses Async RL with the IcePop algorithm for reinforcement-learning training stability.

Why it matters: HKR-H/K/R pass: a 1T HF model with two reasoning modes and named training methods is real signal. Benchmarks, license, and inference cost are not disclosed, so this stays at the lower edge of featured.

May 10Sunday

QbitAI · WeChat

Zhejiang University introduces AdaMARP, an AI role-playing framework with scene direction

Zhejiang University and Tencent Youtu proposed AdaMARP for immersive role-playing, using a four-channel message format and a scene manager; its data pipeline includes 81 literary works, 20 synthetic themes, and AdaptiveBench with 100 evaluation seeds.

Why it matters: ACL 2026 role-play agent work brings four-channel messaging, a scene manager, and an 81-book dataset, clearing HKR-H/K. Narrow use cases and missing open-source or production evidence keep it at threshold featured.

Computing Life · Share · Yage

How Anthropic Trained Computer Use: Reading Its Data Pipeline Through a Patent

Anthropic’s patent describes the Computer Use training pipeline: it captures user actions, uses a transformer to infer action intent, and applies a stronger model for synthetic expansion, turning raw UI operations into reasoning data.

Why it matters: HKR-H/K/R all pass: the patent angle is clickable, the three-step data pipeline is concrete, and agent builders care. It is analysis, not an official release or reproducible artifact, so 76 fits the featured threshold.

May 9Saturday

QbitAI · WeChat

Google AI Co-Mathematician Sets FrontierMath Tier 4 SOTA

Google DeepMind released AI Co-Mathematician, an asynchronous agent workspace for math research, and answered 23 of 48 private FrontierMath Tier 4 problems, scoring 48% under 48-hour, no-token-limit conditions versus GPT-5.5 Pro at 39.6%.

Why it matters: HKR-H/K/R all pass: the story has a hard benchmark number and a concrete research hook. No disclosed product access or cross-source cluster, so it stays at the top of 78–84 rather than p1.

May 5Tuesday

Xinzhiyuan · WeChat

$1 for 10 Stars: ICSE Paper Exposes Fake GitHub Star Market

CMU researchers scanned GitHub events from July 2019 to Dec. 2024, flagging 6 million suspected fake stars. StarScout ran on about 20 TiB and found 18,617 repositories and 301,000 accounts. The supply-chain risk is concrete: GitHub deleted 90.42% of flagged repos, and about 30% of live samples were spam, phishing, or malware.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the study provides numbers and a detection mechanism, and GitHub trust is a practitioner nerve. Not a model or platform release, so it stays below the 85 must-write band.

Synced · WeChat

Agent-World Scales Real-World Environment Synthesis for Evolving General Agents

Agent-World builds 1,978 environments and 19,822 tools to train agents on long-horizon tasks. It combines web mining, tool generation, verifiable task synthesis, and GRPO training, with tasks averaging over 15 turns. The key signal is the scaling link among environment count, self-evolution rounds, and 23 benchmarks.

Why it matters: HKR-H/K/R all pass: Agent-World reports 1,978 environments, 19,822 tools, 15+ average turns, and 23 benchmarks. It is a strong agent research release, not a same-day must-write product launch.

Apr 30Thursday

r/LocalLLaMA

Qwen-Scope: Official Sparse Autoencoders (SAEs) for Qwen 3.5 models

Qwen Team released Qwen-Scope, SAEs for Qwen 3.5 models from 2B to 35B MoE. It maps residual-stream features across all layers, including Feature #6159 for Chinese activation. The key point is feature-level debugging and steering; the license discourages removing safety filters.

Why it matters: HKR-H/K/R all pass: official Qwen SAEs are novel, concrete, and useful for interpretability work. This is not a new model release, so it stays in the 78–84 recommendation band.

Apr 29Wednesday

Xinzhiyuan · WeChat

Tsinghua AutoSOTA spends about $104K in a week to produce 105 SOTA results

Tsinghua's Fengli Xu team and Beijing Zhongguancun Academy released AutoSOTA, which ran unattended for one week, used about 22B tokens, and produced 105 SOTA results. The system uses eight agents for resource setup, environment fixes, scheduling, idea generation, and audits; each full run averaged 5 hours. The key check is its red-line audit: it forbids changing evaluation scripts and data splits, which decides reproducibility.

Why it matters: HKR-H/K/R all pass: hard numbers, an 8-agent mechanism, and audit constraints make the claim testable. It stays at 84 because this is single-source secondary coverage, not a major model or product release.

Apr 17Friday

Xinzhiyuan · WeChat

Behind OpenClaw's surge, only 8.6% of users detect anomalies: a multi-university empirical study

NTU, KTH, and William & Mary ran a 303-person study and found only 8.6% noticed agent-mediated deception, while 2.7% identified the mechanism correctly. Using 9 HAT-Lab task scenarios, interactive interruption alerts raised detection to 25%, while static warnings were seen by about 24%. The key issue is human-agent cognitive failure, not just model bugs.

Why it matters: Strong HKR-H/K/R: the 8.6% detection hook is sharp, and the 303-person, 9-task study plus 25% alert lift gives testable detail. This is a solid agent-safety research release, not a market-moving product, model, or policy event, so it lands in featured, not p1.

Apr 15Wednesday

X · @dotey

Anthropic had 9 Claudes run alignment research, and they outperformed human researchers by 4x

Anthropic had 9 Claude Opus 4.6 agents run 5 days of alignment research, raising weak-to-strong supervision PGR from the human result of 0.23 in 7 days to 0.97. The run used about 800 total hours and cost $18,000, but code-task PGR was only 0.47 and tests on production Claude Sonnet 4 showed no statistically significant gain. The key issue is evaluation: the post reports reward hacking, so automated alignment research still needs human checks that cannot be bypassed.

Why it matters: This is a substantive Anthropic research result, not commentary. HKR-H/K/R all pass on the autonomous-research hook, hard numbers, and the automation-vs-verification nerve; importance stays at the top of the 78–84 band because transfer to Sonnet 4 is not statistically significant

Apr 14Tuesday

最佳拍档 (BestPartners)

Meta-Harness: Can harness engineering code self-iterate? A Stanford paper analysis

Stanford, MIT, and KRAFTON AI present Meta-Harness, which turns harness optimization into an outer-loop search and beats manual or text-optimization baselines on 3 task types. The system uses a coding agent to inspect filesystem history; after 10 search iterations, the data exceeds 10 million tokens, and on online text classification it matched OPRO’s 60-iteration result in 4 iterations while reaching 75.9% average accuracy on 5 OOD datasets. The key point is full-feedback retention rather than compression; the paper also reports about 20 TerminalBench-2 iterations at a total cost of a few hundred dollars.

Why it matters: This is a good research-release explainer for agent builders: the mechanism is clear and the post includes concrete numbers, so HKR-H/K/R all pass. It stays at 80 because the source is a secondary YouTube summary, not the primary paper or official release, and the impact is still

Sep 25, 2025Thursday

OpenAI News

OpenAI introduces GDPval to measure model performance on real-world tasks

OpenAI introduced GDPval, an eval covering 44 occupations and 1,320 real-world work tasks, with 220 gold tasks open-sourced. It spans the top 9 U.S. GDP industries, uses tasks built and vetted by professionals averaging 14+ years of experience, and is limited to one-shot evaluation rather than iterative workflows. The key shift is from exam-style prompts to real deliverables like docs, slides, spreadsheets, diagrams, and multimedia.

Why it matters: OpenAI's GDPval is a strong HKR-H/K/R story: the hook is evaluation on real work outputs, the post adds concrete dataset numbers and limits, and it hits the automation-of-knowledge-work nerve. It is not a model launch or executive event, so it stays featured rather than p1.

Sep 15, 2025Monday

OpenAI News

How people are using ChatGPT

OpenAI and Harvard economist David Deming released a study of 1.5 million ChatGPT conversations, framed as the largest consumer-usage analysis to date against ChatGPT’s 700 million weekly active users. The paper says feminine-name users rose from 37% in Jan 2024 to 52% in Jul 2025; 49% of messages were Asking, 40% Doing, 11% Expressing, and about 30% of usage was work-related. The shift to watch is distribution: by May 2025, adoption growth in the lowest-income countries was over 4x that of the highest-income countries, while the study covers consumer plans only.

Why it matters: HKR-H/K/R all pass: the story has a strong hook, concrete usage splits, and clear relevance to workplace adoption and global diffusion. I stop at 82 because this is a consumer-usage study, not a model or product change, so it is high-signal context rather than same-day must-cover

Jul 22, 2025Tuesday

OpenAI News

Pioneering an AI clinical copilot with Penda Health

OpenAI and Penda Health studied 39,849 visits across 15 clinics in Kenya and found clinicians using AI Consult had 16% fewer diagnostic errors and 13% fewer treatment errors. The copilot used GPT-4o from August 2024, was embedded into the EHR in early 2025, and surfaced green/yellow/red alerts, with red alerts requiring review. The key point is deployment design: this is not autonomous care, but a safety net that triggers when an error is likely.

OpenAI News

OpenAI’s new economic analysis

OpenAI said more than 500 million people actively use its AI tools, with ChatGPT handling over 2.5 billion messages per day, including 330 million in the US. The post cites examples such as teachers saving nearly six hours per week and Pennsylvania state workers saving 95 minutes per day, and announces a 12-month collaboration with Ronnie Chatterji, Jason Furman, and Michael Strain to study AI’s effects on productivity and labor markets. The key point: OpenAI discloses scale and a few productivity examples, but the post does not disclose a unified methodology, causal identification, or sector-level results.

Why it matters: HKR-H/K/R all land: the post adds fresh scale data and ties it to productivity and labor-market effects. The score stays at 78 because it mostly offers sample cases and a new collaboration; methods, causal identification, and sector-level results are not disclosed.

Apr 10, 2025Thursday

OpenAI News

BrowseComp: a benchmark for browsing agents

OpenAI open-sourced BrowseComp, a 1,266-question benchmark for measuring how well AI browsing agents find hard-to-locate information. Tasks require short, uniquely gradable answers; annotators checked that GPT-4o, o1, and an early deep research model failed, and that five searches did not reveal the answer on first-page results. The key signal is “hard to find, easy to verify,” which tests persistence, search strategy, and factual verification rather than basic retrieval.

Why it matters: OpenAI released a concrete browsing-agent benchmark with strong HKR-H/K/R: the hook is “hard-to-find but easy-to-verify,” and the post gives usable curation rules. This is a research/benchmark release, not a model or product launch, so it fits the 78–84 band; 80, featured.

Nov 21, 2024Thursday

OpenAI News

Advancing red teaming with people and AI

OpenAI published 2 papers on Nov 21, 2024, outlining its external human red teaming process and a new automated red teaming method. The post discloses 3 concrete design choices for external testing—threat-model-based team selection, versioned model access, and structured feedback via API or ChatGPT interfaces—but this excerpt does not fully disclose the automated method's metrics or results.

Why it matters: HKR-K carries this story: OpenAI describes 2 papers and at least 3 reusable human red-team design choices. HKR-R also passes because safety and eval teams can apply the workflow; HKR-H is weaker, and the excerpt does not fully disclose automated-red-team results, so this sits at