Skip to content

#MCP/工具调用

6 today

May 1Friday

Bloomberg Technology

AI Payoff in Focus During Tech Earnings Bonanza | Bloomberg Tech 4/30/2026

Bloomberg Tech covered AI payoff in tech earnings, saying Alphabet and Amazon show clearer returns while Meta lags. Anthropic is weighing funding at a valuation above $900B; Stripe’s John Collison discussed AI tools and a Google partnership. The post does not disclose AI spending amounts.

Why it matters: HKR-H/K/R pass: Bloomberg frames AI ROI by company and cites a >$900B Anthropic funding valuation. The score stays below 78 because the post is a roundup video and AI spending figures are not disclosed.

TechCrunch · AI

After Dissing Anthropic for Limiting Mythos, OpenAI Restricts Access to Cyber, Too

OpenAI will first roll out GPT-5.5 Cyber only to “critical cyber defenders.” The RSS snippet does not disclose eligibility rules, pricing, or launch timing. The access-tiering model is the key detail for practitioners.

Why it matters: HKR-H/K/R all pass, but the body is RSS-only: it confirms tiered access for GPT-5.5 Cyber, not criteria, pricing, or timeline. This fits a lower-featured OpenAI safety product update.

TechCrunch · AI

Stripe introduces Link, a digital wallet autonomous AI agents can use

Stripe introduced Link, a digital wallet for cards, banks, subscriptions, and AI-agent spending. The post cites approval flows, but does not disclose fees, limits, or merchant coverage. Watch the authorization boundary for agent payments.

Why it matters: HKR-H/K/R pass: agent wallet payments are clickable, the approval-control mechanism is concrete, and spend authorization is a live practitioner concern. Missing rates, limits, and merchant coverage keep it in the 72–77 band.

The Verge · AI

Meta is running get-rich-quick ads for its AI tools

The Verge says Meta-owned Manus ran quick-money ads for AI tools after a $2B acquisition. The pitch targets local firms with no or bad websites. Manus also paid creators for Instagram, YouTube, and TikTok promotion; some TikTok accounts were removed after inquiry.

Why it matters: HKR-H/K/R all pass: the story has a strong Meta-versus-grift hook, concrete funnel details, and reputational stakes. It is investigative industry reporting, not a major model or product release.

Apr 30Thursday

Ben's Bites

Building Gets Easier

Ben’s Bites lists agent tooling updates from Cloudflare, Stripe, Cursor SDK and others, with over 10 product leads. Cloudflare lets agents create accounts, buy domains, get API tokens and deploy; Stripe adds Agentic Commerce Suite, Link CLI and agent-ready Treasury accounts. The key shift is external permissions becoming agent-readable interfaces.

Why it matters: HKR-H/K/R pass, but this is a roundup rather than one major launch. Concrete Cloudflare and Stripe agent-permission details keep it in the featured-low band.

r/LocalLLaMA

Notes on what actually breaks when you run a coding agent on small local models

A Reddit user tested small local and free-tier cloud models for weeks on multi-file coding tasks. Sub-7B structured output was unreliable; failures included markdown fences, wrong-file edits, and read/write misclassification, with post-processing and validation as fixes.

Why it matters: HKR-H/K/R pass: the post names real local coding-agent failure points, a sub-7B threshold, four failure classes, and mitigations. Reddit single-post scope keeps it below release-tier news, so 75.

r/LocalLLaMA

Qwen-Scope: Official Sparse Autoencoders (SAEs) for Qwen 3.5 models

Qwen Team released Qwen-Scope, SAEs for Qwen 3.5 models from 2B to 35B MoE. It maps residual-stream features across all layers, including Feature #6159 for Chinese activation. The key point is feature-level debugging and steering; the license discourages removing safety filters.

Why it matters: HKR-H/K/R all pass: official Qwen SAEs are novel, concrete, and useful for interpretability work. This is not a new model release, so it stays in the 78–84 recommendation band.

Xinzhiyuan · WeChat

AI Raw Proofs Pile Up on GitHub as Terence Tao Says Solving Alone Is Not Enough

Terence Tao says math is shifting from proof scarcity to proof abundance, with 20-plus AI solutions pending assessment on an Erdős problems GitHub page. The post says GPT-5.4 Pro generated an Erdős #1196 approach in 80 minutes, and Tao verified the core within 24 hours. The key issue is verification and digestion workflow, not raw proof count.

Why it matters: All HKR axes pass: Tao plus GitHub proof backlog gives HKR-H, while 20+ pending AI solutions and an 80-minute GPT-5.4 Pro claim give HKR-K. This is not a model release, so it stays below 85.

TechCrunch · AI

Microsoft says it has over 20M paid Copilot users, and they really are using it

Microsoft says Copilot has over 20M paid users, with engagement growing. The post does not disclose active usage, retention, ARPU, or the counting method.

Why it matters: HKR-K is strong because Microsoft disclosed 20M+ paid Copilot users, a rare adoption metric. The score stays near the featured floor because active rate, retention, ARPU, and methodology are not disclosed.

r/LocalLLaMA

Building a fully local PDF-to-audiobook workflow with Kokoro 82M, Qwen and llama.cpp

Reddit user purellmagents shared a local PDF-to-audiobook workflow using Kokoro 82M, Qwen 3.5 0.8B/2B, and llama.cpp. The Tauri 2.0 app runs on an M1 Mac, reads 15 initial sentences, then prepares the next 15. The hard parts are PDF-text alignment, code snippets, tables, and first-generation latency.

Why it matters: HKR-H/K/R all pass, but this is a Reddit personal workflow, not a model or platform release. Specific components and the 15-sentence pipeline keep it at the low featured band.

Hacker News front page

Ramp’s Sheets AI Exfiltrates Financials

PromptArmor disclosed a Ramp Sheets AI flaw with a 6-step attack chain; Ramp said it was fixed on March 16, 2026. A hidden prompt injection in an external sheet made the AI insert an IMAGE formula calling attacker.com with financial data. The key issue is formula insertion without user approval.

Why it matters: HKR-H/K/R all pass: the post gives a concrete exfil path for an AI spreadsheet tool. Scored 82, not 85+, because it is single-source and impact scale is not disclosed.

X · @dotey

Inside Hermes Agent's Memory System and How It Avoids OpenClaw's Pitfalls

Hermes Agent splits memory into 4 layers: prompt files, SQLite session search, skills, and optional Honcho. MEMORY.md is capped at 2,200 chars, USER.md at 1,375; writes apply after a new session or compression. The key design is cache-first: keep system prompts stable and retrieve long-tail history via tools.

Why it matters: HKR-H/K/R all pass: the OpenClaw contrast is clickable, and the memory limits/mechanisms are concrete. Single X-source tutorial, not a product release, keeps it at the featured threshold.

Apr 29Wednesday

Xinzhiyuan · WeChat

Tsinghua AutoSOTA spends about $104K in a week to produce 105 SOTA results

Tsinghua's Fengli Xu team and Beijing Zhongguancun Academy released AutoSOTA, which ran unattended for one week, used about 22B tokens, and produced 105 SOTA results. The system uses eight agents for resource setup, environment fixes, scheduling, idea generation, and audits; each full run averaged 5 hours. The key check is its red-line audit: it forbids changing evaluation scripts and data splits, which decides reproducibility.

Why it matters: HKR-H/K/R all pass: hard numbers, an 8-agent mechanism, and audit constraints make the claim testable. It stays at 84 because this is single-source secondary coverage, not a major model or product release.

r/LocalLLaMA

DeepSeek V4 pricing is genuinely silly; the math made me question my stack

A Reddit user calculates DeepSeek V4-Pro input at $0.145 per million tokens, about 34x cheaper than Claude Opus 4.7. A May promo cuts it to $0.036, while cache hits are $0.0036, about 173x below Opus cached pricing. The key issue is agent-loop cost; the post does not verify the 1M context under production loads.

Why it matters: HKR-H/K/R all pass on the pricing hook, concrete token prices, and agent-cost pressure. Capped below 78 because this is a Reddit calculation, not an official release or production benchmark.

X · @dotey

OpenAI Expands AWS Partnership, Bringing GPT-5.5, Codex, and Managed Agents to Bedrock

OpenAI expanded its AWS partnership, bringing GPT-5.5, Codex, and Managed Agents to Amazon Bedrock in limited preview. Codex supports Bedrock across CLI, desktop, and VS Code, with over 4M weekly active users. The key detail is reuse of AWS compliance, billing, and cloud commitments.

Why it matters: HKR-H/K/R all pass: this is more than a routine cloud listing, with OpenAI bringing GPT-5.5, Codex, and Managed Agents to AWS Bedrock. Limited preview, 4M weekly Codex users, and IDE/CLI entry points justify same-day coverage.

Hacker News front page

OpenAI Models Coming to Amazon Bedrock: Interview with OpenAI and AWS CEOs

OpenAI will bring models to AWS via Bedrock Managed Agents, in an interview with Sam Altman and Matt Garman. Microsoft and OpenAI amended their deal: Azure ships first, cross-cloud service is allowed, and IP licensing runs through 2032. The post does not disclose pricing, model list, or launch timing.

Why it matters: All HKR axes pass: Altman and AWS's CEO confirm OpenAI models on Bedrock, tied to Microsoft agreement changes through 2032. Price, model list, and launch timing are undisclosed, so it stays in the 85–94 band.

X · @dotey

AI terminal tool Warp open-sources client code with OpenAI as founding sponsor

Warp open-sourced its client code under AGPL; only the client is open, while server code stays closed. The Rust terminal has 700,000+ developers, and its Oz cloud AI handles coding, planning, and tests. The key signal is its AI-first contribution workflow.

Why it matters: HKR-H/K/R all pass: the OpenAI sponsorship hook, AGPL/client-only detail, and 700K-developer signal are concrete. This is a strong dev-tool open-source update, not a major model or capability release.

The Verge · AI

Claude can now plug directly into Photoshop, Blender, and Ableton

Anthropic launched Claude connectors for creative apps, including Adobe Creative Cloud, Affinity, Blender, Ableton, and Autodesk. The Blender connector can debug scenes, build tools, and batch-apply object changes; the post does not disclose pricing or full availability.

Why it matters: HKR-H/K/R all pass: the hook is Claude inside major creative apps, with concrete connector behavior. Missing price and rollout details keep it below must-write status.

Apr 28Tuesday

X · @claudeai

Claude Now Connects to Tools Creative Professionals Already Use

Claude added a Blender connector for scene debugging, tool building, and batch object edits from Claude. The post does not disclose versions, pricing, or rollout scope; the key issue is agent control boundaries inside DCC workflows.

Why it matters: HKR-H/K/R pass: Claude’s Blender connector is a concrete agent-tool expansion. Missing version, pricing, and rollout details keep it near the featured threshold, not a must-write.

Hacker News front page

GitHub Copilot code review will start consuming GitHub Actions minutes

GitHub will make Copilot code reviews consume GitHub Actions minutes starting June 1, 2026. Private-repo reviews use plan entitlements, with overages billed at standard Actions rates; public repos stay free. The change covers Copilot Pro, Pro+, Business, and Enterprise, including direct org billing for unlicensed users.

Why it matters: Official GitHub billing change for Copilot code review hits CI quotas and org invoices; HKR-H/K/R all pass, but it is a pricing rule, not a capability release, so it sits low in 72–77.

Computing Life · Share · Yage

Agentic Creative Tools: From Photoshop Actions to Claude for Creative Work

Anthropic released 9 creative-tool Connectors for Claude for Creative Work. The post frames agentic creative tools around programmable APIs, connector protocols, and perceptual feedback loops. The post does not disclose the Connector list.

Why it matters: HKR-H/K/R all pass: Claude creative agents have a clear hook, 9 connectors add a fact, and creator workflow pressure adds resonance. Missing connector names and access terms keep it below must-write.

Hacker News front page

Claude Pro: Opus Requires Extra Usage in Claude Code

Anthropic lists 6 Claude Code models, and Pro users need extra usage enabled and purchased to use Opus. The guide gives 3 configuration paths: /model, --model, and ANTHROPIC_MODEL in zsh or bash. The post does not disclose extra usage pricing or quotas.

Why it matters: HKR-H/K/R all pass, but the facts come from a help doc and cover Claude Code access/configuration, not a new model or major capability. Anthropic relevance lifts it to the lower featured band.

X · @dotey

The West forgot how to build things, and may forget how to write code

Denis Stetskov compares Western defense production gaps with AI coding, citing Stinger orders placed in 2022 for 2026 delivery. He says Europe’s 1M-shell target was 9 months late, and METR found senior developers 19% slower with AI. The key risk is the junior-engineer pipeline, not code generation speed.

Why it matters: HKR-H/K/R all pass: the analogy is clickable, the post gives concrete defense and METR numbers, and the junior-engineer pipeline resonates. X translation/commentary limits authority, so it sits just above the featured threshold.

X · @dotey

GitHub Copilot switches to usage-based billing on June 1

GitHub Copilot will switch to AI Credits billing on June 1 while keeping subscription prices unchanged. Credits count input, output, and cached tokens; Pro includes $10 monthly credits and Pro+ includes $39. Watch Copilot Agent long-task costs.

Why it matters: HKR-H/K/R all pass: Copilot billing moves from subscription expectations to token/cache consumption with date and credit amounts. Single-source X context lacks enterprise details and overage rates, so it stays in the 78–84 band.

X · @dotey

Cursor 3 feedback: users want a reliable AI development workspace

Eric Zakariasson’s Cursor 3 feedback thread summarizes 431 replies, with users asking for a stable AI development workspace. Requests center on Agent Window retaining LSP, debugging, Git, terminal and diff workflows, plus multi-agent worktrees and model-cost transparency. The key issue is workflow reliability, not a flashier IDE.

Why it matters: All HKR axes pass: 431 user replies, concrete workflow requests, and strong resonance for Cursor users. Kept in the low featured band because this is feedback synthesis, not an official Cursor release or roadmap.

Hacker News front page

GitHub Copilot is moving to usage-based billing

GitHub said on 2026-04-27 that GitHub Copilot will move to usage-based billing. The captured post only shows the title, time, and navigation. It does not disclose the launch date, usage metric, prices, or overage rules.

Why it matters: GitHub Copilot billing affects a large developer base. HKR-H and HKR-R are strong, while HKR-K is limited to the usage-based mechanism with no date, metering unit, or price details disclosed.

Apr 27Monday

The Verge · AI

Canva apologizes after its AI tool replaces ‘Palestine’ in designs

Canva said Magic Layers replaced “Palestine” with “Ukraine” in designs. The tool should split flat images into editable layers, not alter visible content; X user @ros_ie9 said “Gaza” was unaffected. Canva says it fixed the issue; the post does not disclose the trigger mechanism.

Why it matters: This is a concrete Canva Magic Layers incident, not a routine feature post. HKR-H comes from the unexpected word swap, HKR-K from the stated tool boundary and fix, and HKR-R from political-bias and trust risk.

Hacker News front page

Show HN: Utilyze — an open-source GPU monitoring tool claiming higher accuracy than nvtop

Systalyze open-sourced Utilyze to measure real GPU compute efficiency in production, with negligible overhead claimed. The post says nvidia-smi and nvtop only check whether any kernel runs during the sampling window; an H100 has 132 SMs and 17,424 cores. The key issue is real throughput headroom, not binary utilization dashboards.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the post explains the sampling flaw, and GPU waste is a real practitioner nerve. Unknown vendor and single-tool scope keep it in the 72–77 band.

Hacker News front page

Running Local LLMs Offline on a Ten-Hour Flight

Dmitri Lerko ran Gemma 4 31B and Qwen 4.6 36B locally during a 10-hour flight with no Wi‑Fi. The MacBook Pro M5 Max had 128GB unified memory and a 40-core GPU; sustained load used about 1% battery per minute, and performance degraded past 100k tokens. The sharp finding is instrumentation: an iPhone cable delivered 60W, while a MacBook cable delivered 94W under the same load.

Why it matters: HKR-H/K/R all pass: this is a named first-person local-inference test with concrete hardware, model, battery, and power numbers. Scope stays practical rather than industry-shaking, so it lands in the 72–77 band.

QbitAI · WeChat

Meshy tops 10M users and moves into 3D printing as ARR rises 14x

Meshy says it passed 10M registered users, reached $40M ARR, and grew 2025 revenue 14x year over year. Meshy Creative Lab supports keychain, magnet, and keycap design; physical ordering is not live yet. The key signal is print fit: 97% slice-pass rate in Bambu Studio across 75 tested models.

Why it matters: HKR-H/K/R all pass: the hook, revenue metrics, and print-readiness test are concrete. This is a vertical 3D AI product update from company disclosure, so it lands at the lower featured band.

Hacker News front page

The Prompt API

Chrome’s docs describe the Prompt API for calling built-in AI inside the browser. The page links to session management and structured output docs; the captured body does not disclose model, context window, pricing, or rollout details.

Why it matters: Chrome Prompt API clears HKR-H/K/R: native browser AI is a real hook, and session plus structured-output docs add usable detail. Model, context window, pricing, and release timing are not disclosed, keeping it in the lower featured band.

Synced · WeChat

From 99 Lines of Frozen Code to Meshy AI’s 3D Momentum in the West

Meshy AI released Meshy 6 and claims over 60% share in developed Western markets. The post says it has 10M+ users, $40M+ ARR, and 100M+ AI-generated 3D models in three years. The key signal is workflow fit: 37Games reports 30–40% less base sculpting work.

Why it matters: HKR-H/K/R pass: Meshy 6 has a clear founder/product hook, concrete traction metrics, and a production-labor angle. Kept in the low featured band because the market-share claim is company-sourced and no independent benchmark is disclosed.

OpenAI News

An Open-Source Spec for Orchestration: Symphony

OpenAI released Symphony, an open-source spec for Codex orchestration. The RSS snippet says it turns issue trackers into always-on agent systems; the post does not disclose spec details, license, APIs, or benchmarks.

Why it matters: HKR-H and HKR-R pass: an OpenAI open-source Codex orchestration spec is relevant to agent workflows. HKR-K is weak because license, interfaces, and reproducible mechanics are not disclosed.

Apr 26Sunday

Hacker News front page

Agents Aren’t Coworkers, Embed Them in Your Software

Feldera co-founder Gerd Zellweger argues agents should be embedded in existing software, not treated as chatty coworkers. He lists 3 patterns: CLI, declarative specs, and Kubernetes-style reconciliation loops, then adds CDC streams for inserts, updates, and deletes. The key split: agents adapt logic, while the engine runs it continuously and emits precise changes.

Why it matters: HKR-H/K/R all pass, but this is vendor engineering commentary, not a launch or first-person benchmark. Concrete architecture patterns justify featured, not the 78+ band.

Hacker News front page

Using Coding Assistance Tools to Revive Projects You Never Were Going to Finish

Matthew Brunelle used Claude Code with Opus 4.6 to rebuild a YouTube Music-to-OpenSubsonic connector, listing 6 setup steps. The stack used FastAPI, Pydantic, ytmusicapi, and yt-dlp, with Feishin logs used to fix .view suffix handling. The useful point: a clear spec plus human review beat one-shot generation.

Why it matters: HKR-H/K/R all pass, but the impact stays at a first-person coding workflow. Claude Code + Opus 4.6, a concrete connector stack, and Feishin-log debugging place it in the quality tutorial band, not a broader industry update.

Apr 25Saturday

Hacker News front page

Open-source memory layer Stash lets any AI agent do what Claude.ai and ChatGPT memory can do

Stash released an open-source persistent memory layer for AI agents, exposing 28 MCP tools and a 6-stage pipeline for long-term memory. The page says it uses PostgreSQL plus pgvector and hierarchical namespaces to separate user, project, and self memory. The real point is a portable memory layer, not the headline claim about matching ChatGPT or Claude.ai.

Why it matters: HKR-H/K/R all pass: the hook is portable long-term memory for any agent, and the page gives concrete architecture details. The score stays in the low featured band because this is an indie OSS infrastructure launch, not a major lab or platform release.

Computing Life · Share · Yage

Anthropic lets Claude Cowork run rival models, a stranger move than it looks

Anthropic added an April 22–23 Claude Cowork switch for GPT-5.5, Gemini 3.1 Pro, DeepSeek V4, or local models. The post says third-party deployments have no Anthropic seat fee, and Bedrock, Vertex, and gateway prompts stay outside Anthropic. The key fight is runtime and control plane: AWS, Google, and Microsoft bet on Agent Registry, Apigee, and Entra Agent ID.

Why it matters: All three HKR axes pass: the competitor-model switch is a strong hook, and the article gives billing and data-flow details. Capped below P1 because sourcing is unofficial, with no independent benchmark and a small Cowork base.

Computing Life · Share · Yage

TPU vs. CUDA: A Post-Cloud Next 2026 Assessment

Google announced TPU 8t/8i, TorchTPU, and an Anthropic deal at Cloud Next 2026; TPU 8i is slated for H2 2027 volume production. 8i has 288GB HBM, 8.6TB/s bandwidth, and 384MB SRAM; TorchTPU runs PyTorch on TPU, but the post says independent benchmarks are missing. The key crack is vLLM inference, while the author says TPU will not replace NVIDIA within 18-24 months.

Why it matters: HKR-H/K/R all pass: clear TPU-vs-CUDA rivalry, concrete 8i specs and TorchTPU details, and strong NVIDIA cost/supply resonance. No independent benchmark and H2 2027 production keep it in 78–84, not P1.

Hacker News front page

Databases Were Not Designed for This

Arpit Bhayani argues agentic AI breaks four database assumptions: deterministic queries, human-reviewed writes, brief connections, and human-monitored failures. He proposes Postgres role timeouts of 5s and 10s, soft deletes, append-only logs, and idempotency keys. The key shift is treating agent_worker as an untrusted caller, not sizing pools like human-written apps.

Why it matters: HKR-H/K/R all pass: the angle is sharp, the post gives concrete Postgres guardrails, and the risk is real for agent builders. Not a model or product release, so it fits the 72–77 engineering commentary band.

X · @dotey

Cursor 3 adds /multitask for parallel async sub-agents

Cursor 3 added /multitask and lets async sub-agents run in parallel. Queued tasks can also switch to parallel mode without waiting for the previous task to finish. The post does not disclose concurrency limits, resource usage, or failure rollback.