Skip to content

#Agent

39 today

May 11Monday

AI HOT (Curated Pool)

Cog House Opens for the First Time: Scott Wu and the Rise of Cognition AI

Cognition AI disclosed internal footage of Cog House, while Devin reached $445 million in annualized revenue within 18 months of launch and the company is valued at about $25 billion.

Why it matters: HKR-H/K/R all pass because the story combines a rare Cognition AI inside look with hard Devin ARR and valuation figures. It stops below P1 because this is a profile-style reveal, not a funding, product, or model release.

AI HOT (Curated Pool)

AntLingAGI Releases Trillion-Parameter Ring-2.6-1T Model

AntLingAGI released Ring-2.6-1T, a trillion-parameter thinking model available for free on OpenRouter until May 15, with adjustable thinking intensity, agent-oriented multi-step execution, tool calling, and tasks covering math logic and scientific research.

Why it matters: HKR-H/K/R all pass, but the post is thin: no benchmarks, pricing, architecture, or training details. Treat as a mid-weight model launch on OpenRouter, not a same-day must-write.

Import AI (Jack Clark)

Import AI 456: RSI and Economic Growth; Radical Optionality for AI Regulation; and a Neural Computer

Import AI 456 covers radical optionality for AI regulation and a Neural Computer paper, listing seven proposed governance tool categories, including transparency, reporting, audits, whistleblower protections, evaluations, model-weight security, and talent, while also noting Meta and KAIST prototypes using Wan 2.1 for CLI and GUI neural-computer experiments; the RSS snippet is truncated before full prototype results.

Why it matters: HKR-H/K/R all pass: this is a high-signal Import AI roundup, not a hard launch. The concrete value is the 7 regulatory tools plus Wan 2.1 prototypes, so it clears featured but stays below major-release bands.

AI HOT (Curated Pool)

Tencent Hunyuan Hy3 Preview Released for Complex Agent Tasks

Tencent Hunyuan opened early access to the Hy3 preview, which uses a 256K context window and a mixture-of-experts architecture with fast and slow thinking for complex agent tasks.

Why it matters: HKR-H/K/R all pass: Tencent Hunyuan Hy3 preview names 256K context and a fast/slow-thinking MoE for complex agents. Benchmarks, pricing, and access scope are not disclosed, keeping it in the 78–84 band.

r/LocalLLaMA

ExLlamaV3 Major Updates

ExLlamaV3 added DFlash in v0.0.31, raising Coding throughput from 59.21 t/s to 177.67 t/s; v0.0.32 optimized five models, with Trinity-Nano gaining 72.4% on 6000 Pro², while v0.0.33 adds DFlash model quantization plus bug fixes and efficiency work.

Why it matters: HKR-H/K/R all pass, but the blast radius is mostly LocalLLaMA and ExLlama users. This fits a mid-weight open-source inference update, not a same-day industry-wide story.

AI HOT (Curated Pool)

OpenAI Launches DeployCo to Help Enterprises Build Businesses Around Intelligence

OpenAI launched DeployCo, an enterprise deployment company focused on moving AI systems into production, while the RSS snippet does not disclose pricing, customer names, deployment scope, or launch timeline.

Why it matters: OpenAI launching DeployCo is a real enterprise strategy signal: HKR-H has a separate-company hook and HKR-R hits deployment competition. HKR-K is weak because pricing, customers, and timing are absent, so it sits at the featured floor.

Xinzhiyuan · WeChat

The Second Half of Agent Evaluation: Why a Live Benchmark Is Needed

Claw-Eval-Live evaluates 13 frontier models on 105 tasks, and the top model stays below a 70% pass rate, while HR tasks average only 6.8% pass rate.

Why it matters: HKR-H/K/R all pass: the live benchmark hook is specific, and the post gives 105 tasks, 13 models, HR at 6.8%. Claw-Eval-Live still lacks proven field impact, so this sits in the lower featured band.

Xinzhiyuan · WeChat

Largest IPO Nears, Topping SpaceX; 2028 AI Self-Iteration Countdown

Xinzhiyuan says Anthropic is considering a near-$1 trillion valuation, with ARR rising to $45 billion in five months; Jack Clark predicts a greater than 50% chance that AI systems can autonomously build better versions of themselves by the end of 2028, while the article cites a 72% Kalshi probability of an IPO announcement before November 1.

Why it matters: HKR-H/K/R all pass: the hook is sharp and the post gives valuation, ARR, and 2028 odds. Source is secondary and IPO/ARR claims lack official confirmation, so it stays in 78-84.

Xinzhiyuan · WeChat

Claude Mythos Hits 50% Success on 16-Hour Tasks in METR Time Horizons

Claude Mythos Preview reached a 50% success rate on METR Time Horizons tasks that take humans 16 hours, while only 5 of 228 tasks exceeded the 16-hour range, so the article says METR lacks enough samples to quantify longer-horizon performance.

Why it matters: HKR-H/K/R all pass: the 16-hour task result is a strong hook, and the METR sample caveat adds substance. Capped at 82 because only 5 tasks exceed 16 hours, so the 2027 extrapolation is not same-day P1 material.

Computing Life · Share · Yage

DeployCo Arrives: OpenAI and Anthropic Form AI Deployment JVs with PE on the Same Day

OpenAI and Anthropic announced AI deployment joint ventures with private equity on May 4, and the snippet cites divergent terms, including a 17.5% guaranteed return versus no guaranteed return.

Why it matters: HKR-H/K/R all pass: the angle has tension, the facts include PE JVs and a 17.5% floor, and the nerve is model-lab commercialization. Single-source commentary keeps it in the 78–84 band, not must-write.

Computing Life · Share · Yage

Google shuts down Project Mariner; Anthropic and OpenAI also hit limits

Google quietly shut down Project Mariner on May 4, and the post says Google, Anthropic, and OpenAI reached the same conclusion: standalone browser agents do not work, while GUI automation still has room outside headless dedicated environments.

Why it matters: HKR-H/K/R all pass: the shutdown date, route-level claim, and Google/OpenAI/Anthropic contrast carry signal. Single-source summary lacks an official notice or failure metrics, so this stays in the low featured band.

AI HOT (Curated Pool)

Local models handle half of daily tasks and respond faster than cloud models

A five-week experiment tested about 1,400 daily work tasks, where local 35B models such as Qwen 3.6 35B handled about 50% and averaged 2.8-second responses, 2.1 times faster than Claude Opus 4.5, while the cloud model still led complex reasoning by about 20%.

Why it matters: HKR-H/K/R all pass: Tom Tunguz’s experiment reports ~1,400 tasks, ~50% success, 2.8s latency, and a speed comparison to Claude Opus 4.5. Strong practitioner signal, but not a model launch or platform-level update.

AI HOT (Curated Pool)

Codex autonomously completes a security audit and earns a bounty

A user instructed Codex to earn $5; Codex spent about 22 hours finding an open-source security audit bounty, submitting a valid PR, communicating with maintainers, passing GitHub verification, and ultimately receiving a $16.88 payment.

Why it matters: HKR-H/K/R all pass: a Codex agent allegedly closed a bounty loop in 22 hours with concrete money and workflow details. Single social-post evidence lacks reproducible logs, so it stays below P1.

AI HOT (Curated Pool)

MachinaCheck: Multi-agent CNC manufacturability analysis system built on AMD MI300X

MachinaCheck runs Qwen 2.5 7B locally on AMD MI300X to analyze STEP files for CNC manufacturability, reducing drawing review for quote analysis from 30–60 minutes to 30 seconds while using 192GB HBM3 to keep customer design data on-premises.

Why it matters: HKR-H/K/R all pass, but this is an AMD hackathon project on Hugging Face, not a broad model or platform launch. Concrete numbers carry it to the featured threshold.

May 10Sunday

r/LocalLLaMA

We tried vectors, ASTs, and brute-force context stuffing for code retrieval; LLM semantic graphs worked best

ByteBell open-sourced a code indexing system that stores per-file LLM-generated purpose, summary, business context, entities, classes, functions, keywords, and imports in a Neo4j graph, then uses full-text search instead of vector similarity, with SHA-256 diffing to reindex only changed files and keep LLM calls proportional to churn.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, and the post gives a concrete Neo4j semantic-graph mechanism with SHA-256 incremental rebuilds. Reddit sourcing and missing metrics keep it at the 72–77 featured threshold.

Synced · WeChat

A Framework for Mechanic-Aware Iteration in AI Game Generation

CreativeGame makes an agent write a mechanic contract before four code-generation stages, then evaluates iterations with CreativeProxyReward, two hard gates for runtime and static errors, and lineage-aware memory shared within each game evolution tree.

Why it matters: HKR-H/K/R pass, but this is a game-generation research framework without disclosed open-source status, metrics, or production adoption. It fits the 72–77 band rather than a must-write item.

Xinzhiyuan · WeChat

Harsh Claim: Top Silicon Valley AI Is One Year Ahead of the World

Elad Gil claims top AI lab employees are 3-4 months ahead of Silicon Valley, while Silicon Valley is 3-6 months ahead of New York; the post cites Mythos’ 73% success rate in expert cyberattack simulations as evidence in a disputed “geographic time gap” argument.

Why it matters: HKR-H/K/R all pass: the lab-to-user lag hook is clickable, and the post cites 3–4 months, 3–6 months, and a 73% Mythos figure. It is secondhand commentary, not a model or product release, so it stays in the 72–77 threshold band.

QbitAI · WeChat

Zhejiang University introduces AdaMARP, an AI role-playing framework with scene direction

Zhejiang University and Tencent Youtu proposed AdaMARP for immersive role-playing, using a four-channel message format and a scene manager; its data pipeline includes 81 literary works, 20 synthetic themes, and AdaptiveBench with 100 evaluation seeds.

Why it matters: ACL 2026 role-play agent work brings four-channel messaging, a scene manager, and an 81-book dataset, clearing HKR-H/K. Narrow use cases and missing open-source or production evidence keep it at threshold featured.

Computing Life · Share · Yage

How Anthropic Trained Computer Use: Reading Its Data Pipeline Through a Patent

Anthropic’s patent describes the Computer Use training pipeline: it captures user actions, uses a transformer to infer action intent, and applies a stronger model for synthetic expansion, turning raw UI operations into reasoning data.

Why it matters: HKR-H/K/R all pass: the patent angle is clickable, the three-step data pipeline is concrete, and agent builders care. It is analysis, not an official release or reproducible artifact, so 76 fits the featured threshold.

May 9Saturday

AI HOT (Curated Pool)

YC CEO Open-Sources Personal AI OS GBrain for a Compounding Second Brain

Y Combinator CEO Garry Tan open-sourced GBrain, a personal AI operating system that processed more than 20 books in five months and manages over 100,000 pages of structured knowledge.

Why it matters: HKR-H/K/R pass: Garry Tan’s open-source personal knowledge system has a notable-user hook and three concrete usage numbers. Missing repo activity, architecture detail, and tests keep it at the featured threshold.

AI HOT (Curated Pool)

Peekaboo 3.0 Launches With Action-First macOS Control and UI Detection

Peekaboo 3.0 is now live with action-first macOS control, unified screenshots and UI detection, cleaner JSON exchange between CLI and MCP, and improved snapshots; the post does not disclose pricing, model choices, or release timeline beyond the 3.0 launch.

Why it matters: HKR-H/K/R all pass for a concrete desktop-agent tooling update. Score stays at the featured floor because pricing, model details, and adoption data are not disclosed.

AI HOT (Curated Pool)

Baidu releases ERNIE 5.1 with compressed parameters and training cost

Baidu released ERNIE 5.1 with total parameters reduced to about one third of the original scale, active parameters to about one half, and pretraining cost to about 6% of same-scale models; the model is available on the ERNIE platform and Baidu AI Studio.

Why it matters: HKR-H/K/R all pass: Baidu ERNIE 5.1 is a domestic flagship-model release with concrete compression and 6% pretraining-cost claims. That puts it in the must-write band.

AI HOT (Curated Pool)

Using Codex to debug and verify fixes in parallel

The author uses Codex in temporary crabbox environments to recreate bug states, verify failures, apply fixes, and re-verify them, while running 10 sessions in parallel to avoid local state pollution and speed loss.

Why it matters: HKR-H/K/R all pass, but this is a single first-person workflow note, not a product release or benchmark. The 10-session Codex/crabbox setup earns featured-level practical signal, near the lower band.

AI HOT (Curated Pool)

ERNIE 5.1 Released With Pretraining Cost at 6% of Comparable Models

Baidu released ERNIE 5.1, saying it builds on ERNIE 5.0 pretraining and improves search, reasoning, knowledge QA, creative writing, and agent capabilities, with pretraining cost at about 6% of comparable models.

Why it matters: Baidu released ERNIE 5.1 with a concrete “6% of reference pretraining cost” claim. HKR-H/K/R all pass, with a domestic flagship-model bump, but sparse technical detail keeps it below the 90s.

Xinzhiyuan · WeChat

CUHK Open-Sources ArbiterOS Agent Governance Kernel With 92.95% High-Risk Interception

CUHK CURE Lab open-sourced ArbiterOS, an agent runtime governance kernel that intercepts, parses, governs, and observes actions before execution, raising high-risk step interception on OpenClaw tasks from 6.17% to 92.95%.

Why it matters: HKR-H/K/R all pass: the story has a sharp execution-control hook, a concrete 6.17%→92.95% result, and clear agent-safety resonance. It is a strong open-source research tool, not a top-lab model release, so it stays in the 78–84 band.

QbitAI · WeChat

Why Perfect AI Agents Do Not Exist: Five Design Philosophies and Trade-offs Behind Claude Code

MBZUAI VILA Lab and UCL analyze Claude Code v2.1.88 source code and identify 5 design philosophies, 13 design principles, 7 permission layers, and 5 context-compaction layers behind its production-agent architecture.

Why it matters: All HKR axes pass: the contrarian Claude Code angle is clickable, the v2.1.88 permission/context mechanisms add substance, and agent tradeoffs resonate with builders. It is third-party analysis, not an Anthropic release, so it stays below must-write.

QbitAI · WeChat

Google AI Co-Mathematician Sets FrontierMath Tier 4 SOTA

Google DeepMind released AI Co-Mathematician, an asynchronous agent workspace for math research, and answered 23 of 48 private FrontierMath Tier 4 problems, scoring 48% under 48-hour, no-token-limit conditions versus GPT-5.5 Pro at 39.6%.

Why it matters: HKR-H/K/R all pass: the story has a hard benchmark number and a concrete research hook. No disclosed product access or cross-source cluster, so it stays at the top of 78–84 rather than p1.

QbitAI · WeChat

Qwen AI Glasses S1 Adds Spatial 3D Display, Proactive Reminders, and Daily AI Features

Qwen AI Glasses S1 added spatial 3D display and proactive services, with ride-hailing, instant shopping, and photo-based homework help scheduled for this month; Wellsenn XR says Qwen AI Glasses hold 53% of China’s online AI glasses sales since March 8.

Why it matters: HKR-H/K/R all pass, but this is an AI-glasses feature update rather than a model or platform release. The 53% online-sales share and this-month feature list justify low featured range.

Synced · WeChat

DeepSeek Reportedly Raises RMB 50B, with Liang Wenfeng Funding 40%, Valuation Reaching RMB 350B

DeepSeek is negotiating a $7.3 billion funding round at an estimated $51.5 billion valuation; Liang Wenfeng reportedly plans to contribute 40%, while Tencent and China’s RMB 60 billion national AI fund are also in talks.

Why it matters: HKR-H/K/R all pass: the DeepSeek funding rumor has large numbers, a founder contribution ratio, and named backers. Because it is still reported as talks with no official confirmation, it stays at 84 and featured, not p1.

Synced · WeChat

OpenAI's Jiayi Weng: Is the Next AI Training Paradigm Beyond Gradients?

OpenAI researcher Jiayi Weng proposes Heuristic Learning: codex gpt-5.4 reached a perfect 864 score on Breakout and generated 342 search trajectories across Atari 57, with updates applied to code, tests, replays, and memory rather than neural-network weights.

Why it matters: HKR-H/K/R all pass: an OpenAI researcher proposes Heuristic Learning with concrete hooks like Breakout 864 and 342 Atari 57 trajectories. This is strong research/commentary signal, not an official model or product release, so it stays in the 78–84 band.

Latent Space

Anthropic growing 10x/year while others lay off over 10% of staff

Anthropic is described as growing 10x annually and being valued at $1T-$1.2T, while the post cites layoffs of 40% at Block, 14% at Coinbase, and 20% at Cloudflare under AI-readiness framing.

Why it matters: HKR-H/K/R all pass: the title has contrast, the post gives growth, valuation, and layoff figures, and it hits jobs plus AI-capital concentration. It is high-signal industry commentary, not an official funding or product event, so 78-84 fits.

May 8Friday

AI HOT (Curated Pool)

Robotics Endgame: A Physical AGI Roadmap and LLM Analogy

The speaker presented a physical AGI roadmap with six named components: video world models, WAM, EgoScale, dexterity scaling laws, physical reinforcement learning, and DreamDojo; the snippet also mentions a 2016 OpenAI DGX-1 signing story with Jensen and Elon.

Why it matters: HKR-H/K/R all pass: the physical-AGI endgame hook is strong, the post gives a 6-part roadmap, and robotics practitioners will debate the path. It is still a personal roadmap, not a release or benchmark, so it sits in 78–84.

Hacker News front page

Show HN: Git for AI Agents

regent-vcs released the open-source re_gent project for AI-agent version control, currently supporting Claude Code, with workflows for tracking why an agent changed files, rewinding sessions, and bisecting agent actions; the post does not disclose the license, storage format, or installation details.

Why it matters: HKR-H/K/R all pass: the Git analogy is clicky, the mechanism is concrete, and Claude Code rollback pain is real. The post lacks license, storage format, and install details, so it stays at the featured threshold.

AI HOT (Curated Pool)

Running Codex Safely at OpenAI

OpenAI runs Codex with four safeguards: sandbox isolation, human approval, strict network policies, and native agent telemetry; the post does not disclose evaluation metrics, incident rates, or enterprise deployment requirements.

Why it matters: HKR-H/K/R all pass: the OpenAI Codex post gives concrete safety mechanisms for code agents. I keep it at 74 because it lacks eval data, incident rates, or enterprise rollout details.

Alibaba Technology · WeChat

The AI-Native Era: Where R&D Organizations Go Next

Xu Xiaobin cites internal interviews showing that engineers who use AI heavily cut coding time from 30% to 5%, raised Agent conversation time from 5% to 60%, and increased end-to-end delivery efficiency by 2 to 3 times, while pure coding efficiency rose 10 times.

Why it matters: Alibaba Tech’s internal-interview numbers make HKR-H/K/R pass, but this is org-methodology commentary rather than a product or model release, so it sits just above the featured threshold.

Synced · WeChat

ICLR 2026: NVIDIA and Purdue Use an Agentic Loop for Text-to-3D Scene Generation

NVIDIA Cosmos Lab and Purdue University proposed Scenethesis, a language-and-vision agentic framework for text-to-3D scene generation that uses visual grounding, SDF-based physical constraints, and a judge module; experiments report about 72% first-pass success, 91% after self-checking, and collision rate reduction from 6.1% to 0.8%.

Why it matters: HKR-H/K/R all pass: NVIDIA/Purdue plus an agent loop is clickable, and the post gives SDF constraints, a judge module, and 72%→91% results. Strong research signal, but not a product release, so it stays in 78–84.

QbitAI · WeChat

OpenAI releases three realtime voice models for reasoning, translation, and transcription

OpenAI launched GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper as API models, covering 128K-context voice reasoning, streaming translation from more than 70 input languages into 13 output languages, and realtime transcription priced at $0.017 per minute.

Why it matters: OpenAI shipped three realtime voice APIs across reasoning, translation, and transcription, hitting HKR-H/K/R. The 128K context, 70+ languages, and $0.017/min price make this a same-day must-write item.

QbitAI · WeChat

All Labs Watch ByteDance, Everyone Praises DeepSeek: A U.S. Researcher’s 36-Hour China AI Trip

Ai2 researcher Nathan Lambert visited Zhipu, Moonshot AI, Tsinghua, Meituan, Xiaomi, and 01.AI within 36 hours, and said Chinese labs closely watch ByteDance and respect DeepSeek, while student participation in core work, open source habits, and in-house control of the technical stack mark key differences.

Why it matters: HKR-H/K/R all pass: the piece has a named US researcher’s dense China-lab tour plus concrete claims on ByteDance, DeepSeek, open source, and in-house stacks. It is strong industry field reporting, not a model launch or major deal, so it sits at featured rather than p1.

AI HOT (Curated Pool)

China releases L1-L4 AI terminal intelligence standards covering 7 device categories

MIIT and other agencies released AI terminal intelligence standards covering 7 device categories. The framework uses a “2+N” structure with L1 response, L2 tool, L3 assistance, and L4 collaboration; L4 details come later. The post does not disclose concrete test metrics.

Why it matters: HKR-H/K/R pass: the story has a clear L1-L4 standards hook, concrete “2+N” and 7-category details, and compliance impact for device AI teams. Missing L4 rules and test metrics keep it near the featured floor.

Ruan YiFeng's Weblog

Technology Enthusiast Weekly Issue 395: The Third Way of Software Development

Ruanyifeng Weekly issue 395 frames AI-assisted coding as a “mystery house” style of software development and cites HN SOTA, which ranks model popularity by scanning 200 top Hacker News topics each day and their programming or AI discussions.

Why it matters: HKR-H/K/R pass: the “third way/mystery house” framing, HN SOTA’s 200 daily HN topics, and developer workflow anxiety all land. It is commentary, not a model or product release, so it stays at 72.