Skip to content

#Agent

39 today

May 8Friday

AI HOT (Curated Pool)

OpenAI launches official openai-cli for terminal API calls

OpenAI open-sourced openai-cli for direct API calls from the terminal. The Apache 2.0 tool installs via Homebrew or Go and covers Responses API, structured output, image editing, transcription, and key config. The key detail is Agent workflows using cloud tools like web search and code interpreter.

Why it matters: HKR-H/K/R all pass: official OpenAI terminal tooling is clickable, with concrete install/license/API details and workflow resonance. It is still a developer tooling update, not a model or major capability release, so 76 fits the featured threshold.

Bloomberg Technology

Cloudflare to Cut One-Fifth of Workers in Move to AI-First Model

Cloudflare plans to cut over 1,100 jobs globally, about one-fifth of its workforce. The cuts are tied to an agentic AI-first operating model; the post does not disclose roles, timing, or cost targets.

Why it matters: HKR-H/K/R all pass: Bloomberg reports a 20% Cloudflare cut tied to an agentic AI-first operating model. Role mix, timing, and cost targets are not disclosed, keeping it below P1.

AI HOT (Curated Pool)

Codex Plugin Now Supports Parallel Runs Across Chrome Tabs

OpenAI says Codex now runs in Chrome on macOS and Windows. The plugin works across tabs in the background without taking browser control; the post does not disclose version, concurrency limits, or enterprise policy.

Why it matters: HKR-H/K/R all pass, but the post gives platform and execution mechanics only; version, concurrency limits, and enterprise controls are not disclosed. Score: 76 as a practical OpenAI Codex product update.

NVIDIA Blog

Powering the Next American Century: Chris Wright and NVIDIA’s Ian Buck on Genesis Mission

The U.S. DOE and NVIDIA are building two AI supercomputers at Argonne; Equinox uses 10,000 Grace Blackwell GPUs. Solstice will use 100,000 Vera Rubin GPUs, which Buck said reach 5,000 exaflops. The key bottleneck is grid work: Wright said AI can cut interconnection studies from years to weeks or hours.

Why it matters: HKR-H/K/R all pass: the GPU counts, DOE-NVIDIA role, and grid bottleneck are concrete. NVIDIA-blog sourcing keeps it below must-write; this fits the 78–84 band.

AI HOT (Curated Pool)

Agent Pull Requests Are Everywhere: How to Review Them

GitHub published a guide for reviewing pull requests generated by AI agents. The snippet lists 3 focus areas: code changes, logic or security bugs, and pre-merge technical debt. The key issue is a review process before automated commits reach production.

Why it matters: HKR-H/K/R all pass: GitHub gives a practical checklist for agent-generated PRs with 3 review areas. It is guidance, not a product or model release, so it stays at the featured threshold.

AI HOT (Curated Pool)

Work with Claude across Excel, PowerPoint, Word, and Outlook

Claude now connects to four Microsoft apps: Excel, PowerPoint, Word, and Outlook. Excel, PowerPoint, and Word are generally available; Outlook is in public beta. Admins can deploy via Microsoft admin center and monitor with OpenTelemetry.

Why it matters: HKR-H/K/R all pass: Claude enters 4 Microsoft 365 apps with rollout status and OpenTelemetry details. This is a strong Anthropic product update, but not a model release or core capability jump, so it stays in the 78–84 band.

AI HOT (Curated Pool)

Perplexity launches Personal Computer app for Mac

Perplexity opened its Personal Computer Mac app to all users. It runs on any Mac and works across local files, native Mac apps, the web, and Perplexity secure servers. The post does not disclose pricing, permission boundaries, or task success rates.

Why it matters: HKR-H/K/R all pass, but the source is a single product post with no pricing, permission model, or task success rate. Score stays in the mid-weight product-update band.

May 7Thursday

AI HOT (Curated Pool)

Trillion-parameter instruction model Ling-2.6-1T released

inclusionAI says Ling-2.6-1T is now live on OpenRouter. The trillion-parameter instruction model uses “fast thinking” and claims top AIME26 and SWE-bench Verified results with about 75% lower cost. The post does not disclose pricing, context length, or full benchmark scores.

Why it matters: HKR-H/K/R all pass: a 1T instruction model on OpenRouter with fast thinking, AIME26/SWE-bench claims, and ~75% cost reduction. Missing price, context window, and full scores keep it in the 78–84 band.

Hacker News front page

AlphaEvolve: Gemini-powered coding agent scaling impact across fields

Google DeepMind describes AlphaEvolve as a Gemini-powered coding agent; the body is only an RSS snippet. The title discloses coding-agent scope and cross-field impact, but the post does not disclose model version, benchmarks, or deployments.

Why it matters: HKR-H and HKR-R pass on a DeepMind Gemini coding-agent announcement, but HKR-K fails: only title-level facts are disclosed. This reaches featured threshold, not 78+, because evals, model version, and deployments are absent.

AI HOT (Curated Pool)

Apify mcpc and x402 Give AI Agents an Auto-Payment Wallet

Apify mcpc integrates the x402 payment protocol, letting AI agents auto-sign payments on HTTP 402. x402 compresses paid API settlement into one HTTP round trip plus a signature; mcpc supports Claude Code and USDC-funded wallets. The key point is machine settlement for paid tool calls, not the wallet label.

Why it matters: HKR-H/K/R all pass: the hook is fresh, the mechanism is concrete, and agent payments hit a real practitioner nerve. It is still a mid-weight integration with no usage scale, pricing, or production case disclosed.

r/LocalLLaMA

Qwen/WebWorld 32B/14B/8B (Qwen3 finetune)

Qwen released WebWorld 32B/14B/8B, Qwen3 finetunes for training and evaluating web agents. It uses 1M+ real web trajectories and supports 30+ step simulation plus A11y Tree, HTML, XML, Markdown, and natural-language states. Agents trained on its synthetic trajectories gain 9.9% on MiniWob++ and 10.9% on WebArena.

Why it matters: HKR-H/K/R all pass: WebWorld has an agent hook, concrete scale, and benchmark gains. It is a useful Qwen research release for agent builders, but limited source detail keeps it below the 85 must-write band.

QbitAI · WeChat

Vidu Claw Generates Ad Videos From One Prompt and a Hundred-Yuan Budget

Shengshu Technology opened Vidu Claw, which generates ad scripts, voiceover, music, editing, and final videos from one prompt; its Video Plan includes up to 40 minutes of daily generation across video, image, and audio.

Why it matters: HKR-H has a concrete ad-test hook, HKR-K adds the 40-minute daily quota and one-prompt workflow, and HKR-R hits production-cost pressure. No benchmark or pricing detail, so this stays at the featured threshold.

Ben's Bites

Elon Doubled Limits

Ben’s Bites says Anthropic doubled Claude usage on paid plans via SpaceX’s Colossus 1. The issue also lists GPT-5.5 Instant, ChatGPT spreadsheet integration, and three Claude Managed Agents features. The title names Elon, but the post does not disclose exact limits.

Why it matters: HKR-H/K/R pass: the SpaceX Colossus 1 angle, 2x Claude usage, and quota pressure are all concrete. Missing exact caps, pricing, and rollout scope keep it in the low featured band.

AI HOT (Curated Pool)

Consistent web search and scraping for all models

OpenRouter released tools for tool-calling models to run web search and page scraping. The post says multiple search and scraping engines are supported, but does not disclose names, pricing, or limits. The key item is cross-model tool interface consistency.

Why it matters: HKR-H/K/R all pass, but engines, pricing, and limits are not disclosed. This is a mid-weight Product update: useful for model-agnostic agent stacks, not a major model or capability release.

AI HOT (Curated Pool)

Anthropic Institute Outlines Four Core Research Areas

Anthropic Institute named four research areas: economic diffusion, threats and resilience, real-world AI systems, and AI-driven R&D. The post says it will publish a more granular Anthropic Economic Index and study how AI tools speed AI research. The results will inform Anthropic’s Long-Term Benefit Trust.

Why it matters: HKR-K comes from 4 named research tracks and the Economic Index plan; HKR-R is strong on labor and governance. It is an agenda, not a model, product, or finished result, so it stays in the 72–77 band.

Latent Space

Anthropic-SpaceXAI's 300MW/$5B/yr Deal for Colossus I, ARR Growth Is 8000% Annualized

Anthropic announced a SpaceX compute partnership, doubled Claude Code’s 5-hour limits for Pro, Max, Team, and seat-based Enterprise, raised Opus API limits, and said Claude inference would ramp on Colossus within days; the post treats the 300MW and $5B-per-year figures as widely circulated but not canonized in Anthropic’s own announcement.

Why it matters: HKR-H/K/R all pass: the compute-deal numbers and Claude Code limit changes are concrete and practitioner-relevant. The 300MW/$5B/year claim is unofficial, so it stays below P1.

Xinzhiyuan · WeChat

Claude Managed Agents Add Dreaming, With Reported Task Completion Up to 6x

Anthropic added Dreaming, Outcomes, and multi-agent orchestration to Claude managed agents; Harvey reports about 6x higher task completion. Dreaming reads up to 100 sessions; one demo distilled 5.3M tokens into 98 rules, while Outcomes raised success by up to 10 points. Opus 4.7 and Sonnet 4.6 require access, with $0.08 per session-hour runtime fees.

Why it matters: HKR-H/K/R all pass: Anthropic adds Dreaming, Outcomes, and multi-agent orchestration with 100-session memory, $0.08/session-hour runtime, and Harvey’s ~6x completion claim. This is a same-day Claude agent update.

Bloomberg Technology

Kimi Chatbot Maker Moonshot AI Valued at $20 Billion in Meituan-Led Round

Moonshot AI raised about $2 billion, reaching a $20 billion valuation. The title says Meituan led the round; the post does not disclose investors, stake size, or use of funds. It signals strong demand for Chinese AI startups.

Why it matters: Bloomberg reports Moonshot AI raised about $2B at a $20B valuation, a major capital event for a Chinese model lab. HKR-H/K/R all pass; investor details and use of funds are not disclosed, so this sits in the lower 85–94 band.

AI HOT (Curated Pool)

Amp releases Neo CLI as coding agents shift toward long-horizon workflows

Amp released Neo, a CLI tool covering remote orchestration, automatic context compression, and a Plugin API. Neo lets local threads be controlled remotely, allows all operations by default, and moves safety control to plugins; the post does not disclose version, pricing, or performance gains.

Why it matters: HKR-H/K/R all pass: Neo adds remote orchestration, context compression, Plugin API, and default-allow permissions. Amp’s reach and missing price/version/perf data keep it in the 72–77 band.

Synced · WeChat

TACO Lets CLI Agents Drop Useless Context Through Self-Evolving Compression

TACO proposes a training-free terminal-observation compression framework, improving success rate and token efficiency on TerminalBench 1.0/2.0 and related benchmarks. It evolves rules within tasks, writes validated rules to a global pool, and finds 24.6%–44.1% low-value redundancy in TerminalBench 2.0 raw prompts. The key signal is stability: Top-30 rule retention exceeds 90% after multiple evolution rounds.

Why it matters: HKR-H/K/R all pass: the paper targets CLI-agent context bloat with a no-training rule-pool mechanism and concrete TerminalBench numbers. It is strong agent research, not a major model or product launch, so it sits in the 78–84 featured band.

Synced · WeChat

Claude, GPT and Gemini score 0% completion on ProgramBench

ProgramBench tested Claude Opus 4.7, GPT-5.4 and Gemini 3.1 Pro, with 0% full completion on rebuilding software projects. It gives only executables and usage docs, removes source/tests, and grades behavioral equivalence via agent-driven fuzzing. The key signal is system-level engineering, not function-level code generation.

Why it matters: HKR-H/K/R all pass: the 0% result is clickable, the setup is concrete, and the coding-agent gap matters to practitioners. Still, it is a single benchmark report, below a major model or product release.

AI HOT (Curated Pool)

Open Slide lets AI write PPT code

Open Slide builds PPTs with React, using a workflow designed for AI agents. It integrates SVGL with 1,500+ brand logos, supports manual edits, and lets AI read user comments for revisions.

Why it matters: HKR-H/K/R pass: the programmable-slide angle is clickable, with concrete React and 1500+ logo details, and deck work is a real practitioner pain. No usage metrics or hands-on test keeps it at the featured threshold.

Computing Life · Share · Yage

Agent Filesystems: From Feeding Models Memory to Letting Models Browse Files

The article frames agent filesystems as a three-stage shift from raw context to memory systems to filesystem-as-context, covering design choices from Turso, Anthropic, Vercel, and Manus, and listing four overlooked blind spots.

Why it matters: HKR-H/K/R all pass, but this is design commentary rather than a product or research release. Named comparisons across Turso, Anthropic, Vercel, and Manus justify featured, not the 78+ band.

The Verge · AI

Google shuts down Project Mariner

Google shut down Project Mariner on May 4, 2026. The experimental web-task agent once supported up to 10 concurrent tasks. Its technology moved into Google products, including Gemini Agent.

Why it matters: HKR-H/K/R all pass, but the disclosed facts are limited to shutdown timing, a 10-task limit, and migration into Gemini Agent. Strong source authority supports low featured, not a major launch.

Bloomberg Technology

Meta-Backed Scale AI Wins $500 Million Defense Department Deal

The Pentagon awarded Scale AI a $500 million contract to sift data and support decisions. The post says Meta Platforms backs Scale AI, but does not disclose term, deployment scope, or model details. The key signal is US military AI spending moving into data workflows.

Why it matters: HKR-H/K/R pass on the $500M Pentagon contract and defense-AI procurement angle. Missing term, deployment scope, and model details keep it in the 72–77 featured band.

r/LocalLLaMA

Analysis of 922 Agentic Task Traces Finds DeepSeek v4’s Cost Edge in Caching

A Reddit user analyzed 922 agentic task traces and reported $0.01 per task for DeepSeek v4 Flash versus $1.52 for Opus 4.7. Both used about 960K tokens per task, but DeepSeek showed a 97% cache hit rate versus 87%, with a 0.02 cache read/write price ratio versus 0.08. The key issue is caching, not headline pricing.

Why it matters: HKR-H/K/R all pass: 922 agent traces tie a large cost gap to cache hit rate and cache read/write pricing. Reddit single-source data and incomplete method detail keep it in the 78–84 band.

May 6Wednesday

r/LocalLLaMA

CopilotKit (MIT): Open-source building blocks for agent apps and generative UI

CopilotKit offers MIT-licensed React components and claims 30k GitHub stars. It covers chat, streaming, tool calls, HITL, and generative UI, with AG-UI support for LangGraph, CrewAI, LlamaIndex, and other backends. The key point is decoupling the UI layer from agent frameworks.

Why it matters: HKR-H/K/R all pass: MIT open source, 30k stars, and AG-UI links to major agent backends. Kept in 72–77 because the post lacks a new version, benchmark, or named production adopter.

TechCrunch · AI

Apple to Pay $250M to Settle Lawsuit Over Siri's Delayed AI Features

Apple agreed to pay $250 million to settle a class action over delayed Siri AI features. The snippet cites overpromised rollout claims; the post does not disclose user count, payout rules, or feature timing.

Why it matters: HKR-H/K/R all pass: Apple’s $250M Siri AI settlement has a strong hook, a concrete number, and product-liability resonance. Missing payout rules, affected-user count, and feature timeline keep it in the 78–84 band.

r/LocalLLaMA

An Open Benchmark for Testing RAG on Realistic Company-Internal Data

EnterpriseRAG-Bench released a 500k-document corpus for testing RAG on company-internal data. It simulates Redwood Inference across 9 sources and includes 500 questions over 10 retrieval failure modes. Baselines show BM25 beats vector search overall, while agentic/bash retrieval has the best completeness at higher cost and latency.

Why it matters: HKR-H/K/R all pass: the benchmark targets a real enterprise RAG pain point, with 500k docs and testable BM25-vs-vector results. Single Reddit-source benchmark release keeps it below same-day must-write.

Latent Space

AINews: Silicon Valley Gets Serious About Services

Anthropic and OpenAI announced enterprise services companies: Anthropic’s unnamed JV is funded with $1.5 billion, while OpenAI’s The Deployment Company has raised about $4 billion at a $10 billion pre-money valuation.

Why it matters: HKR-H/K/R all pass: the hook is labs turning into services operators, with $1.5B and ~$4B figures. The scale and OpenAI/Anthropic names put it in must-write territory.

Synced · WeChat

DeepSeek Version of Claude Code Tops Trending Chart With 8,700 Stars

DeepSeek TUI topped GitHub trending with over 8,700 stars. Hunter Bown built it in Rust for local terminal use with DeepSeek V4, supporting chat, file edits, shell commands, and task management. The key detail is RLM mode: up to 16 V4 Flash subtasks, plus a 1M-token context window and approval gates.

Why it matters: HKR-H/K/R all pass: the 8,700-star hook is strong, RLM adds concrete mechanisms, and coding-agent competition resonates. It is a third-party open-source tool, not an official DeepSeek model release, so it stays in the 78–84 band.

Synced · WeChat

Two Chinese open-source projects turn Mac into a private AI workstation

Mininglamp open-sourced Cider and Mano-P 1.0 for Apple Silicon local inference and GUI agents. Cider speeds Qwen3-VL-2B prefill by 57%–61% on M5 Pro; Mano-P 1.0-72B scores 58.2% on OSWorld. The key constraint is W8A8 memory: on 16GB devices accuracy falls from 58.0% to 54.0%, so 32GB+ is recommended.

Why it matters: HKR-H/K/R all pass: the Mac-local workstation angle is clickable, and Cider/Mano-P include testable numbers. Score stays at 80 because the source entity is not a top-tier model lab.

Xinzhiyuan · WeChat

Salesforce plans to hire 1,000 graduates as agent roles expand

Salesforce CEO Marc Benioff said the company will hire 1,000 graduates or interns for Agentforce growth. The post cites Agentforce ARR up 169% to $800 million, with roles covering prompts, evals, agent supervision, and delivery. The key shift is entry roles moving from execution to agent orchestration and output checks.

Why it matters: HKR-H/K/R all pass: 1,000 junior hires, $800M Agentforce ARR, and 169% growth give concrete signal, with a strong jobs angle. This is Salesforce hiring plus Agentforce expansion, not a major model or product release.

Xinzhiyuan · WeChat

Coding at 12, Building a $2B Google Business at 28: He Tells Young People to Stop Chasing Coding

Xinzhiyuan says Alon Chen coded at 12 and managed a $2B Google business at 28. He argues Gen Z should stop chasing coding, citing 30% AI-written Microsoft code and 25%+ at Google. The sharper signal is execution, problem framing, and communication, not coding as a sole moat.

Why it matters: HKR-H/K/R all pass, but this is a career commentary piece, not a model or product release. The two AI-code-share numbers lift it above generic advice, placing it at the featured threshold.

Xinzhiyuan · WeChat

GPT-5.5 Instant becomes ChatGPT’s free default model

OpenAI made GPT-5.5 Instant the default ChatGPT model, rolling it out free to all users. AIME 2025 rose from 65.4% to 81.2%, responses are 30.2% shorter, and hallucinations fell 52.5% versus GPT-5.3 Instant on high-risk prompts. Plus and Pro web users get chat, file, and Gmail personalization first; the API model ID is chat-latest.

Why it matters: HKR-H/K/R all pass: a free default ChatGPT model switch, concrete benchmark and behavior deltas, and direct impact on daily OpenAI workflows. This fits the 85–94 must-write band.

TechCrunch · AI

SAP Bets $1.16B on 18-Month-Old German AI Lab and Says Yes to NemoClaw

SAP plans to buy 18-month-old German AI startup Prior Labs in a $1.16B bet. The RSS snippet says SAP restricts customer agent use to a few options such as Nvidia NemoClaw; the post does not disclose deal structure, closing date, or technical details.

Why it matters: HKR-H/K/R all pass: $1.16B for an 18-month-old AI lab is a strong enterprise-AI hook. Kept at 76 because deal structure, closing timeline, and technical details are not disclosed.

The Verge · AI

Apple agrees to pay iPhone owners $250 million for not delivering AI Siri

Apple agreed to pay $250 million to settle a class action over Apple Intelligence availability claims. It covers US buyers of iPhone 16 models and iPhone 15 Pro from June 10, 2024 to March 29, 2025. Eligible claims pay $25 per device, with an upper range of $95 depending on claim volume.

Why it matters: HKR-H/K/R all pass: the $250M Apple settlement turns an AI Siri delay into a legal-cost story. It lands in 78–84: concrete facts and major platform impact, but no new model or capability ships.

Financial Times · Technology

Apple reaches $250mn settlement over delayed ‘AI Siri’

Apple reached a $250mn settlement over delayed “AI Siri” features. iPhone buyers sued over 2024 marketing for features not yet launched; the post does not disclose payout scope, court filings, or launch timing.

Why it matters: FT reports Apple reached a $250mn settlement over delayed “AI Siri.” HKR-H is the legal twist, HKR-K has the amount and 2024 ad claim, HKR-R hits AI feature delivery risk; missing payout scope keeps it below 85.

The Verge · AI

Apple could let you pick a favorite AI model in iOS 27

Apple plans to let third-party chatbots run system-wide Apple Intelligence in iOS 27, iPadOS 27, and macOS 27. Mark Gurman says Extensions can handle Siri, Writing Tools, and Image Playground this fall. The post does not disclose supported models, pricing, or developer APIs.

Why it matters: HKR-H/K/R all pass: the Apple system-level model picker is a strong hook, with named Extension targets. Scored 80 because model list, pricing, and developer APIs are not disclosed, and this remains a roadmap report.

Financial Times · Technology

Meta plans advanced agentic AI assistant for consumers

Meta plans a consumer agentic AI assistant; the RSS body has one sentence. It says Meta is funding an OpenClaw counterpart for everyday task execution. The post does not disclose model size, launch timing, pricing, regions, or permission controls.

Why it matters: FT reports Meta plans a consumer agentic assistant, with HKR-H/K/R present. Details on launch, pricing, model, and permission design are missing, so this sits at the lower featured band.