Skip to content

#Agent

39 today

May 18Monday

Bloomberg Technology

Baidu AI Sales Eclipse Waning Legacy Ads for the First Time

Baidu reported a 1% revenue decline as growth in nascent AI businesses offset shrinking traditional internet revenue; the post does not disclose AI sales, advertising revenue, or details of the agentic AI pivot.

Why it matters: Baidu revenue fell 1% while AI sales topped legacy ads for the first time, so HKR-H/K/R pass. Missing AI/ad dollar splits and agentic-AI mechanics keep it in the low featured band, not p1.

r/LocalLLaMA

I built a coding agent that gets 87% on benchmarks with a 4B parameter model

SmallCode passes 87 of 100 benchmark tasks with Gemma 4 activating 4B parameters per token. The author attributes the result to compound tools, compile and lint feedback, task decomposition after two repeated failures, and optional escalation to Claude or OpenAI for one task.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the benchmark identity plus replication details are incomplete. It fits a concrete first-person experiment above the featured bar, not the 78+ band.

Synced · WeChat

openJiuwen releases JiuwenSwarm, an open-source multi-agent swarm framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Swarm Skills Hub, and self-evolving Swarm Skills, and reports a 94.2% PinchBench score versus 91.6% for OpenClaw.

Why it matters: HKR-H/K/R all pass: an open-source agent-swarm framework with named components and a PinchBench 94.2% claim. It stays at 78 because openJiuwen is not a top lab and the summary lacks license, reproduction setup, and baselines.

AI HOT (Curated Pool)

Tencent AI Design Agent Ardot Enters Public Beta: Generates Editable Designs and Converts Them to Code

Tencent Cloud opened public beta for Ardot, an AI design agent that generates editable app pages, websites, and posters from one-sentence prompts, then converts designs to code.

Why it matters: HKR-H/K/R pass on a concrete Tencent product beta for editable design-to-code workflows. Missing pricing, model details, benchmarks, and field results keep it at the lower featured threshold.

AI HOT (Curated Pool)

Grok launches Skills feature

xAI launched Grok Skills on May 18, 2026, letting users set preferences, formatting rules, or workflows once and keep them active across all conversations on web, iOS, and Android.

Why it matters: HKR-H/K/R all pass: Grok Skills adds persistent preferences and workflows across web, iOS, and Android. This is a mid-weight xAI product update; rollout scope, limits, and pricing are not disclosed.

AI HOT (Curated Pool)

Composer 2.5 release and technical analysis

Cursor released Composer 2.5, built on a Moonshot open-source checkpoint, trained with synthetic data from real codebases at 25 times the previous scale, and updated with text-feedback reinforcement learning and a sharded Muon optimizer.

Why it matters: HKR-H/K/R all pass: Cursor is a core coding-agent surface, and the post gives concrete training details around Moonshot, 25x data, RL, and Muon. It lacks benchmarks, pricing, or user-facing capability limits, so it stays in the 78–84 band.

Google DeepMind

Google DeepMind adds Street View grounding to Project Genie

Google DeepMind has added Street View real-scene grounding to its experimental prototype Project Genie. Users can pick a US location, then pair it with a style and characters to generate a world.

Why it matters: With Street View imagery wired in, agents and robots can train and navigate in virtual environments that track real places.

Google DeepMind

Introducing Google Antigravity 2.0

Google 发布智能体开发平台 Google Antigravity 2.0。该平台在 Google DeepMind 官网被列为面向开发者的 agentic development platform,与 Gemini 应用、Google AI Studio 并列。原文未披露版本功能、参数或可用性细节。

May 17Sunday

Hacker News front page

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

MinishLab open-sourced Semble, a code-search tool for agents that combines Model2Vec embeddings, BM25, RRF fusion, and reranking; on a 63-repo benchmark, it used 98% fewer tokens than grep+read, reached 0.854 NDCG@10, and ran CPU queries in about 1.5 ms.

Why it matters: HKR-H/K/R all pass: the 98% token claim is clickworthy, the 63-repo benchmark adds substance, and coding-agent context cost is a real practitioner nerve. Impact is still toolchain-level, so it stays below must-write.

r/LocalLLaMA

MiroThinker-1.7 Open-Weight Deep Research Agent Based on Qwen3 MoE

MiroMindAI released the MiroThinker-1.7-deepresearch and mini APIs, with the mini version using 30B total parameters and 3B active parameters, weights on HuggingFace, and context management based on sliding window K=5 plus episode restarts.

Why it matters: HKR-H/K/R all pass, but the source is a Reddit thread and the lab is not top-tier. Open weights, MoE sizing, and context-management details clear featured, not same-day must-write.

Bloomberg Technology

Apple’s New ChatGPT-Like Siri App Will Have Auto-Deleting Chats

The title says Apple’s ChatGPT-like Siri app will support auto-deleting chats; the RSS snippet only adds that iOS 27 will include a Genmoji upgrade, and the post does not disclose retention periods, release timing, or feature details.

Why it matters: HKR-H and HKR-R pass because Bloomberg frames a specific Apple Siri privacy angle; HKR-K fails since retention and feature mechanics are missing, so this stays at the low featured threshold.

Google DeepMind

Google DeepMind launches Gemini for Science toolset

Google DeepMind released Gemini for Science, which includes three experimental tools on Google Labs: Hypothesis Generation, built on Co-Scientist.

Why it matters: Google is packaging research prototypes like Co-Scientist and AlphaEvolve into apply-to-use science tools, showing what agentic research looks like in practice.

AI HOT (Curated Pool)

Microsoft AI CEO predicts AI will automate all white-collar jobs within 18 months

Mustafa Suleyman predicts AI will reach human-level performance within 18 months and automate most professional tasks, including accounting, law, marketing, and project management.

Why it matters: HKR-H and HKR-R are strong, and HKR-K passes on the testable 18-month timeline. The score stays in the low 78–84 band because this is a CEO forecast, not evidence, benchmarks, or a shipped capability.

r/LocalLLaMA

DeepSeek V4's 1M Context Window: The Breaking Point

A Reddit user tested DeepSeek V4 on 45k, 180k, and 520k-token codebases and found 150k-250k tokens best for coding work. Past 300k tokens, line-number precision degraded; at 520k, outputs shifted toward architecture summaries and skipped implementation details.

Why it matters: A single Reddit post limits authority, but HKR-H/K/R all pass: it is a numbered first-person test with a concrete long-context failure pattern. The right band is featured, not 78+, because replication and model details are thin.

Synced · WeChat

AI agents may spend 1,000x more tokens without better results: the hidden bill

Researchers used OpenHands to analyze traces from 8 frontier models on 500 swe-bench-verified tasks, finding that agentic coding reached a 154:1 input-output token ratio and that human difficulty labels correlated weakly with token use at Kendall tau 0.32.

Why it matters: All HKR axes pass: strong cost-performance hook, concrete benchmark setup and correlation numbers, and direct resonance with coding-agent economics. It is not a model or platform launch, so it fits the 78–84 quality-recommendation band.

Synced · WeChat

What Are World Models? Their History and the $10 Billion Bet

Jiqizhixin translated a MoE Capital blog tracing two world-model lineages. The article says more than $10 billion entered the category over 18 months, and cites DreamDojo as using 44,711 hours of first-person video pretraining to reach r=0.995 correlation with real-world robot policy outcomes.

Why it matters: HKR-H/K/R all pass: the hook is strong and the article gives concrete figures, but it is a compiled explainer rather than a new release. It fits the featured-threshold band for a strong commentary/tutorial.

Synced · WeChat

Peter Steinberger Says His Monthly Token Bill Hit $1.3M, Covered by OpenAI

Peter Steinberger used 603 billion tokens across 7.6 million requests in 30 days, with the bill exceeding $1.3 million; he said disabling fast mode cut the price by 70%, and OpenAI does not charge him for the tokens.

Why it matters: HKR-H/K/R all pass: the story has a sharp cost hook, concrete usage numbers, and strong practitioner resonance. It is a first-person bill disclosure, not an OpenAI pricing or product launch, so it sits just above the featured threshold.

AI HOT (Curated Pool)

MagicPath Integrates with Codex to Combine Design and Development

MagicPath AI CEO @skirano demonstrated MagicPath running inside Codex as a native canvas, with users configuring it through one command, dragging UI elements, and letting Codex generate and edit code in real time.

Why it matters: HKR-H/K/R pass: MagicPath puts a draggable design canvas inside Codex with one-command setup and live code edits. Single-demo sourcing and missing framework support, permissions, and reproducible cases keep it at the lower featured band.

AI HOT (Curated Pool)

Study on the Cognition–Action Disconnect in Tool-Using Agents

An interpretability paper studies tool-using agents and finds models often recognize when to call a tool but fail to act, with a cognition-to-action mismatch rate of 26%–54%.

Why it matters: HKR-H/K/R all pass: the story has a sharp agent-failure hook, a 26%-54% mismatch rate, and clear relevance to tool-use reliability. Source detail is thin, with paper name, models, and task setup not disclosed.

AI HOT (Curated Pool)

Ring-2.6-1T Open-Sourced and Listed on OpenRouter for Agent Workflows

AntLingAGI open-sourced Ring-2.6-1T and listed it on OpenRouter with a 75% discount through the end of May; the trillion-scale reasoning model targets agent workflows, including planning, tool use, context maintenance, and complex task execution, using Async RL and IcePop training methods.

Why it matters: HKR-H/K/R all pass: a 1T open agent model is clickable, with OpenRouter access, discount, and training methods disclosed. Score stays at 74 because benchmarks, license, and context window are not given.

May 16Saturday

TechCrunch · AI

OpenAI co-founder Greg Brockman takes charge of product strategy

Greg Brockman has officially taken charge of OpenAI’s product strategy, and Wired reports that he described a plan in a staff memo to combine ChatGPT and Codex into one unified experience.

Why it matters: HKR-H/K/R all pass: OpenAI co-founder product control plus a reported ChatGPT-Codex unification matters. No launch date, feature boundary, or rollout plan is disclosed, so this stays below a major product release.

AI HOT (Curated Pool)

Anthropic Founder’s Playbook warns AI can raise startup failure rates

Anthropic published Founder’s Playbook, arguing that AI tools such as Claude Code reduce prototyping cost but increase startup failure risk across the Idea, MVP, Launch, and Scale stages through false validation, confirmation bias, agentic technical debt, and founder decision bottlenecks.

Why it matters: HKR-H/K/R pass: the Anthropic founder playbook has a sharp counterintuitive angle, a four-stage mechanism, and clear founder resonance. It stays near the featured floor because no dataset or reproducible test is disclosed.

Google DeepMind

Strengthening Singapore’s AI Future: A New National Partnership

Google DeepMind 宣布与新加坡政府达成国家 AI 合作,在新加坡推出多项计划,聚焦医疗健康、科学发现与教育。合作内容包括探索 AI 辅助临床医生、用 AlphaFold 和 Google Earth 推进东南亚传染病研究、为盲人及低视力跑者开发基于 Gemma 的跑步助手,并向中小学至初级学院教育者提供 Gemini for Education。

AI HOT (Curated Pool)

Researchers use Anthropic Mythos to build a macOS kernel exploit bypassing Apple M5 MIE

Three researchers used Anthropic Mythos to develop a macOS kernel exploit in six days, moving from discovery on April 25 to completion on May 1, bypassing Apple’s MIE memory-integrity system for M5 and A19 chips and gaining root via standard unprivileged system calls; the full technical report will follow Apple’s patch.

Why it matters: HKR-H/K/R all pass: Anthropic Mythos, a 6-day macOS kernel exploit, and M5/A19 MIE bypass create real dual-use signal. Kernel-exploit depth and single X-source sourcing keep it below the 85 must-write band.

AI HOT (Curated Pool)

Codex adds multi-device remote control and shared context

Codex controls multiple devices through ChatGPT, switches by project to access each device’s context and files, and supports remote SSH setup for other VMs.

Why it matters: HKR-H/K/R all pass, but the item is a thin X-post summary with no official release note, pricing, permission model, or reproducible demo. Treat it as a mid-weight coding-agent product update at the featured threshold.

r/LocalLLaMA

Qwen3.6-35B-A3B and 9B land on the public Terminal-Bench 2.0 leaderboard

little-coder × Qwen3.6-35B-A3B scored 24.6% ±3.2 on Terminal-Bench 2.0, above Gemini 2.5 Pro on Gemini CLI at 19.6% and Qwen3-Coder-480B on Terminus 2 at 23.9%.

Why it matters: HKR-H/K/R all pass, but this is a Reddit post with leaderboard numbers only; test setup and reproducibility details are not disclosed. Strong code-agent benchmark signal, not a 78+ release story.

AI HOT (Curated Pool)

OpenAI Restructures as Brockman Takes Over Product Strategy

OpenAI merged ChatGPT, Codex, and API into one product organization, with Greg Brockman taking over product strategy; the post says Anthropic’s valuation reached $900 billion, but it does not disclose the restructuring timeline.

Why it matters: HKR-H/K/R all pass: this is an OpenAI top-level product reorg covering ChatGPT, Codex, and API. Single-source summary keeps it below the highest band, but it is same-day must-write news.

QbitAI · WeChat

Codex Integrates HeyGen for Prompt-Based Video Generation and Editing

Codex integrates the HeyGen plugin to run image generation, talking-avatar video, subtitles, and edits from natural-language prompts; the article tests roughly one-minute avatar generation, trimming content after 10 seconds, and deleting a blink at the eighth second.

Why it matters: HKR-H/K/R all pass, backed by a numbered hands-on test. The scope is still one Codex-to-HeyGen plugin workflow, not a model or platform release, so it lands in the 72-77 featured band.

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

AI HOT (Curated Pool)

Ignoring Token Costs, Using 100 AI Instances to Automate an Open Source Project

The OpenClaw team runs about 100 Codex instances to handle code review, security analysis, issue deduplication, test reproduction, task creation from meetings, spam filtering, and performance regression monitoring.

Why it matters: HKR-H/K/R all pass: 100 Codex instances running open-source maintenance is a strong operational anecdote with concrete task types. Single X post, no cost, outcome metrics, or reproducible setup, so it stays in the lower featured band.

AI HOT (Curated Pool)

Runway Agent Generates Complete Ads in One Session

Runway Agent turns product photos and ideas into fully produced ads in one session; the post does not disclose the model, pricing, generation length, or regional availability.

Why it matters: Runway’s ad-generation Agent clears HKR-H/K/R as a mid-weight product update. Missing model, pricing, duration, and region details keep it at the featured threshold, not a must-write release.

The Verge · AI

OpenAI keeps shuffling executives to win the AI agent battle

OpenAI announced a reorganization Friday that makes Greg Brockman the official lead for product, and its memo says the company will combine ChatGPT and Codex into one unified agentic experience.

Why it matters: HKR-H/K/R all pass: the power shuffle is clickable, the product merge is new, and OpenAI's agent roadmap matters. It stays at 83 because no capability shipped and timing, pricing, and technical details are absent.

The Verge · AI

AI radio hosts demonstrate why AI can’t be trusted alone

Andon Labs had Claude, ChatGPT, Gemini, and Grok run separate radio stations with $20 in seed money each; the RSS snippet says all failed, but the post does not disclose the full experimental results.

Why it matters: HKR-H/R are strong because the agent-failure setup is memorable and relevant. HKR-K is present but thin: it gives four models and $20 budgets, while full experimental results are not disclosed.

May 15Friday

AI HOT (Curated Pool)

Feishu Open-Source CLI Tool Gets 10,000 Stars in 45 Days with Visible AI Operations

Feishu’s open-source lark-cli gained over 10,000 GitHub stars in 45 days, letting AI create groups and documents through the command line with each operation previewable and reviewable.

Why it matters: HKR-H/K/R all pass, but the source is a single social post and lacks usage, contributor, or adoption data. This fits the lower featured band for an open-source agent tool update.

Alibaba Technology · WeChat

Qoder 1.0 launches as an agentic development workspace beyond AI IDE

Alibaba released Qoder 1.0 with downloads for Windows, macOS, and Linux, adding a standalone Quest workspace, cross-project parallel agent tasks, a team knowledge engine, and Experts mode with five roles for planning, research, coding, review, and testing.

Why it matters: Alibaba’s Qoder 1.0 is a mid-weight AI coding product release with concrete agent-workflow features and developer resonance. No pricing, benchmark, or task-success data is disclosed, so it stays near the featured threshold.

Synced · WeChat

Amazon employees reportedly tokenmaxx to meet AI usage KPIs

Amazon required more than 80% of developers to use AI tools each week and created an internal token-consumption leaderboard. Employees reportedly used the internal MeshClaw agent to inflate usage, while Amazon has limited visibility of the statistics to each employee and their direct manager.

Why it matters: HKR-H/K/R all pass: Amazon’s AI-use KPI became token-gaming, with >80% target, leaderboard, MeshClaw, and visibility changes. Impact is workplace-significant, not major-release level, so featured not p1.

Synced · WeChat

MemPrivacy Shows a Privacy Layer for AI Memory

MemTensor and HONOR open-sourced MemPrivacy for edge-cloud agent memory protection using local reversible pseudonymization; MemPrivacy-4B-RL reached 85.97% composite F1 on MemPrivacy-Bench, 50.47 percentage points above OpenAI privacy-filter, while the benchmark covers 200 users and more than 155,000 privacy items.

Why it matters: HKR-H/K/R all pass: the story has a sharp memory-privacy hook, a concrete reversible pseudonymization mechanism, and benchmark numbers. Single-source release from non-frontier labs keeps it at 78.

Xinzhiyuan · WeChat

Hassabis Praises Google DeepMind's AI-enabled Pointer Powered by Gemini

Google DeepMind released a Gemini-powered AI-enabled pointer and opened two demos in Google AI Studio: image editing and place finding on maps, while the post says Chrome pointer selection and a Googlebook Magic Pointer are planned product paths.

Why it matters: HKR-H/K/R all pass: the prompt-free pointer is clickable, the two AI Studio demos add concrete facts, and UI replacement resonates. Scope is still demo-level, with no metrics or API details, so 78 not 85+.

AI HOT (Curated Pool)

Databricks brings GPT-5.5 to enterprise agent workflows

Databricks made GPT-5.5 available through AI Unity Gateway for AgentBricks and Agent Supervisor API workflows; on OfficeQA Pro, it became the first model above 50% accuracy and reduced errors by 46% versus GPT-5.4.

Why it matters: HKR-H/K/R all pass: GPT-5.5 enters Databricks workflows with 50% OfficeQA Pro accuracy and 46% fewer errors than GPT-5.4. It stays below a full model-release score because the page is a sales-led OpenAI customer story using Databricks’ own benchmark.

AI HOT (Curated Pool)

Connect Grok to the Hermes Agent

xAI connects Grok subscription accounts to Nous Research’s open-source Hermes Agent across all subscription tiers, letting users run Grok 4.3 text chat and reasoning, generate spoken replies with text-to-speech, create images and videos with Grok Imagine, and connect the agent to WhatsApp or Discord.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight xAI product integration with an open-source agent, not a flagship model release. Featured fits; it does not clear the 85+ same-day bar.