Skip to content

#Agent

36 today

Apr 28Tuesday

X · @dotey

GitHub Copilot switches to usage-based billing on June 1

GitHub Copilot will switch to AI Credits billing on June 1 while keeping subscription prices unchanged. Credits count input, output, and cached tokens; Pro includes $10 monthly credits and Pro+ includes $39. Watch Copilot Agent long-task costs.

Why it matters: HKR-H/K/R all pass: Copilot billing moves from subscription expectations to token/cache consumption with date and credit amounts. Single-source X context lacks enterprise details and overage rates, so it stays in the 78–84 band.

X · @dotey

Cursor 3 feedback: users want a reliable AI development workspace

Eric Zakariasson’s Cursor 3 feedback thread summarizes 431 replies, with users asking for a stable AI development workspace. Requests center on Agent Window retaining LSP, debugging, Git, terminal and diff workflows, plus multi-agent worktrees and model-cost transparency. The key issue is workflow reliability, not a flashier IDE.

Why it matters: All HKR axes pass: 431 user replies, concrete workflow requests, and strong resonance for Cursor users. Kept in the low featured band because this is feedback synthesis, not an official Cursor release or roadmap.

Apr 27Monday

Dwarkesh Patel podcast

What I've been Thinking About This Weekend: Open Questions, Intelligence vs Power, Verification in Science

Dwarkesh lists open AI questions, including that five hyperscalers own over 70% of global AI compute. He asks about coding agents, KV cache costs, merging training with inference, and online learning; the post gives questions, not experimental answers.

Why it matters: HKR-H/K/R all pass: Dwarkesh adds a concrete compute-concentration claim and practitioner-relevant questions. No experiment, release, or policy change, so it stays in the 72–77 commentary band.

TechCrunch · AI

China blocks Meta’s $2B Manus deal after months-long probe

China ordered Meta to unwind its $2B Manus acquisition. The title says the probe lasted months; the post does not disclose the legal mechanism, timeline, or Meta’s next steps. The deal risk directly hits Meta’s AI agents push.

Why it matters: HKR-H/K/R all pass: a $2B Meta-Manus AI-agent deal was blocked by China after a months-long probe. The article lacks the legal mechanism and Meta’s next step, so it lands at 88, not 90+.

Hacker News front page

Show HN: OSS Agent Dirac topped TerminalBench on Gemini-3-flash-preview

Dirac-run released Dirac and says it topped TerminalBench using Gemini-3-flash-preview. The repo claims 50-80% lower API costs via Hash Anchored edits, parallel operations, and AST manipulation; the post does not disclose full scores.

Why it matters: HKR-H/K/R all pass: an OSS coding agent claims a TerminalBench lead with cost and mechanism details. Held to 78 because the post relies on repo claims and lacks full leaderboard scores or reproduction logs.

Mistral AI

Mistral AI opens public preview of Workflows

Mistral AI has put Workflows, its enterprise AI orchestration layer, into public preview. It offers durable execution, observability and human-in-the-loop approvals. ASML, ABANCA and CMA-CGM are already using it to automate critical processes.

Why it matters: It lays out Workflows' orchestration features, deployment model and customer cases, showing the engineering bar for enterprise AI processes.

Xinzhiyuan · WeChat

First Spatio-Temporal Time-Series Reasoning Framework for LLMs | ACL'26

Emory University, Microsoft, and partners introduced STReasoner for spatio-temporal time-series reasoning, with ST-Bench covering four task types. It uses Network SDE plus Multi-Agent data generation, then Align, SFT+CoT, and S-GRPO training. The article claims inference cost is 0.004× closed models, with code on GitHub.

Why it matters: HKR-H and HKR-K pass: the story has a “first framework” hook plus ST-Bench, S-GRPO, 0.004× cost, and code release. HKR-R is weak because spatiotemporal reasoning is a narrower research lane.

QbitAI · WeChat

Stanford-led LLM-as-a-Verifier claims SOTA on Terminal-Bench 2.0

Stanford, Berkeley and Nvidia introduced LLM-as-a-Verifier, claiming SOTA on Terminal-Bench 2.0 and SWE-Bench Verified. It selects trajectories via score-token granularity, repeated checks and criteria decomposition; ForgeCode accuracy reached 86.4%.

Why it matters: HKR-H/K/R all pass: Stanford, Berkeley, and NVIDIA offer a concrete verifier mechanism and benchmark numbers. It is still a benchmark research release, not a major model or product launch, so it fits the 78–84 band.

QbitAI · WeChat

DeepSeek V4 Cuts Prices Permanently; Cached Inputs Get 90% Off, Coding Test Costs Drop 83%

DeepSeek V4 cut prices twice in two days: input/output pricing is 75% lower, with cached inputs getting another 90% off. QbitAI’s coding test fell from 31.73 yuan for 35M tokens to 5.34 yuan under new pricing, an 83% drop. The key case is high cache-hit workloads, with V4-Pro at about 95–96% cache hits.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 pricing has a sharp cost hook, concrete test numbers, and strong cost resonance. It is still a pricing update, not a new model release, so it stays below the 85 P1 band.

QbitAI · WeChat

Meshy tops 10M users and moves into 3D printing as ARR rises 14x

Meshy says it passed 10M registered users, reached $40M ARR, and grew 2025 revenue 14x year over year. Meshy Creative Lab supports keychain, magnet, and keycap design; physical ordering is not live yet. The key signal is print fit: 97% slice-pass rate in Bambu Studio across 75 tested models.

Why it matters: HKR-H/K/R all pass: the hook, revenue metrics, and print-readiness test are concrete. This is a vertical 3D AI product update from company disclosure, so it lands at the lower featured band.

Hacker News front page

The Prompt API

Chrome’s docs describe the Prompt API for calling built-in AI inside the browser. The page links to session management and structured output docs; the captured body does not disclose model, context window, pricing, or rollout details.

Why it matters: Chrome Prompt API clears HKR-H/K/R: native browser AI is a real hook, and session plus structured-output docs add usable detail. Model, context window, pricing, and release timing are not disclosed, keeping it in the lower featured band.

Synced · WeChat

ACL 2026: Sending AI “~” May Cause It to Delete Your Home Directory

ACL 2026 accepted an LLM safety paper on emoticon semantic confusion. The team tested 6 models with 3,757 cases; average confusion was 38.6%, with over 90% silent failures. The key risk is agent execution, where “ignore emoticons” prompts had limited effect.

Why it matters: ACL 2026 safety research clears HKR-H/K/R: a sharp file-deletion hook, concrete test numbers, and direct agent-execution risk. It is strong research, not a model launch or platform incident, so it stays in the 78–84 band.

OpenAI News

An Open-Source Spec for Orchestration: Symphony

OpenAI released Symphony, an open-source spec for Codex orchestration. The RSS snippet says it turns issue trackers into always-on agent systems; the post does not disclose spec details, license, APIs, or benchmarks.

Why it matters: HKR-H and HKR-R pass: an OpenAI open-source Codex orchestration spec is relevant to agent workflows. HKR-K is weak because license, interfaces, and reproducible mechanics are not disclosed.

Hacker News front page

If You Stop Hiring Juniors, Your Senior Engineers Own You

Justin Smestad argues that firms stopping junior hiring in 2026 risk costly senior-heavy teams by 2030. The mechanism: a senior can demand a 40% raise; without a two-year bench, replacement may take six months. The key issue is pipeline leverage, not quarterly headcount savings.

Why it matters: HKR-H/K/R all pass, but this is an individual commentary, not a model, product, or research release. The 40% raise and 6-month replacement claims give it enough signal for low featured.

Apr 26Sunday

Hacker News front page

Agents Aren’t Coworkers, Embed Them in Your Software

Feldera co-founder Gerd Zellweger argues agents should be embedded in existing software, not treated as chatty coworkers. He lists 3 patterns: CLI, declarative specs, and Kubernetes-style reconciliation loops, then adds CDC streams for inserts, updates, and deletes. The key split: agents adapt logic, while the engine runs it continuously and emits precise changes.

Why it matters: HKR-H/K/R all pass, but this is vendor engineering commentary, not a launch or first-person benchmark. Concrete architecture patterns justify featured, not the 78+ band.

TechCrunch · AI

Anthropic created a test marketplace for agent-on-agent commerce

Anthropic tested Project Deal, an agent marketplace with 69 employees given $100 budgets. The pilot produced 186 deals worth over $4,000 and ran four model setups. Advanced models got better outcomes, but users did not notice the gap.

Why it matters: HKR-H/K/R all pass: Anthropic tested agent commerce with concrete counts, budgets, trades, and model-market splits. Score stays at 82 because this is an internal test market, not a public product or model release.

Apr 25Saturday

Hacker News front page

What's Missing in the 'Agentic' Story

Mark Nottingham critiques the “AI agent works for you” story and lists 8 trust-misalignment cases online. One example says Microsoft’s new Outlook sends third-party email passwords to its cloud and 700+ data partners. The key issue is delegation boundaries, not model capability alone.

Why it matters: HKR-H/K/R all pass, but this is sourced commentary rather than a model or product release. Mark Nottingham’s Web-protocol authority and HN traction put it at the featured threshold, not P1.

Hacker News front page

Open-source memory layer Stash lets any AI agent do what Claude.ai and ChatGPT memory can do

Stash released an open-source persistent memory layer for AI agents, exposing 28 MCP tools and a 6-stage pipeline for long-term memory. The page says it uses PostgreSQL plus pgvector and hierarchical namespaces to separate user, project, and self memory. The real point is a portable memory layer, not the headline claim about matching ChatGPT or Claude.ai.

Why it matters: HKR-H/K/R all pass: the hook is portable long-term memory for any agent, and the page gives concrete architecture details. The score stays in the low featured band because this is an indie OSS infrastructure launch, not a major lab or platform release.

Computing Life · Share · Yage

Anthropic lets Claude Cowork run rival models, a stranger move than it looks

Anthropic added an April 22–23 Claude Cowork switch for GPT-5.5, Gemini 3.1 Pro, DeepSeek V4, or local models. The post says third-party deployments have no Anthropic seat fee, and Bedrock, Vertex, and gateway prompts stay outside Anthropic. The key fight is runtime and control plane: AWS, Google, and Microsoft bet on Agent Registry, Apigee, and Entra Agent ID.

Why it matters: All three HKR axes pass: the competitor-model switch is a strong hook, and the article gives billing and data-flow details. Capped below P1 because sourcing is unofficial, with no independent benchmark and a small Cowork base.

Computing Life · Share · Yage

Anthropic’s Three Experiments in Claude-Run Commerce: From a Fridge to a Market

Anthropic ran 3 Claude commerce experiments in 12 months, spanning a mini-fridge, a multi-agent store, and a 69-person Slack market. Project Deal closed 186 trades; Opus sellers earned $2.68 more than Haiku, while Opus buyers paid $2.45 less. The key signal: weaker-model users did not perceive the loss.

Why it matters: HKR-H/K/R all pass: Anthropic’s real-commerce agent tests include transaction counts, model deltas, and failure cases. It is a strong research analysis, not a new model launch, so it stays in the 78–84 band.

Hacker News front page

Databases Were Not Designed for This

Arpit Bhayani argues agentic AI breaks four database assumptions: deterministic queries, human-reviewed writes, brief connections, and human-monitored failures. He proposes Postgres role timeouts of 5s and 10s, soft deletes, append-only logs, and idempotency keys. The key shift is treating agent_worker as an untrusted caller, not sizing pools like human-written apps.

Why it matters: HKR-H/K/R all pass: the angle is sharp, the post gives concrete Postgres guardrails, and the risk is real for agent builders. Not a model or product release, so it fits the 72–77 engineering commentary band.

MIT Technology Review · AI

Three reasons why DeepSeek’s new model matters

DeepSeek released a V4 preview with two versions: V4-Pro and V4-Flash. V4-Pro costs $1.74/M input tokens and $3.48/M output tokens; V4-Flash is about $0.14/$0.28, and both support 1M-token context. The key point is attention efficiency and open weights pressuring agentic coding costs.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 is a domestic flagship release with 1M context, two price tiers, and open-weight cost pressure. The preview status keeps it below a full GPT/Claude major release, but it is same-day material.

X · @dotey

Cursor 3 adds /multitask for parallel async sub-agents

Cursor 3 added /multitask and lets async sub-agents run in parallel. Queued tasks can also switch to parallel mode without waiting for the previous task to finish. The post does not disclose concurrency limits, resource usage, or failure rollback.

Hacker News front page

Could a Claude Code routine watch my finances?

Matt May used Claude Code routines with his Driggsby MCP server and Plaid to automate a daily finance email; he says the project took 2 months and about 75k lines of Rust. The post says the Gmail connector can only create drafts, so he added a restricted `email_me()` MCP tool that sends Markdown-only mail to a verified owner address. The practical angle is operability: routine behavior changes via prompt edits, and he already runs alerts on 7-day card anomalies and daily checking outflows over $500.

Why it matters: This is a strong first-person implementation write-up: Claude Code routines + Plaid, Gmail draft-only limits, a constrained email tool, and concrete anomaly rules. HKR-H/K/R all pass, but it is still a single product blog post rather than a lab or platform release, so it lands in

X · @AnthropicAI

New Anthropic research: Project Deal

Anthropic announced Project Deal and had Claude buy, sell, and negotiate for employees in a San Francisco office marketplace. The setup is confirmed as an internal marketplace; the post does not disclose scale, model version, or outcome metrics.

Why it matters: This clears featured on HKR-H and HKR-R: Anthropic has attention weight, and an agent negotiating office deals is inherently discussable. It stays mid-band because HKR-K is weak; the post gives the setup, but not sample size, model version, success metrics, or controls.

Apr 24Friday

Hacker News front page

Affirm Retooled Its Engineering Organization for Agentic Software Development in One Week

In February 2026, Affirm paused normal engineering work for one week and asked 800+ engineers to complete a full agentic workflow from ideation to submitted PR; it says over 60% of PRs are now agent-assisted. The post adds that 80%+ of engineers were weekly active users of AI dev tools by December 2025, and a nine-engineer group spent two weeks defining a default workflow around Claude Code, local-first development, and human checkpoints; the captured body does not fully disclose later implementation details or measured outcomes.

TechCrunch · AI

In another wild turn for AI chips, Meta signs deal for millions of Amazon AI CPUs

Meta signed a deal for millions of Amazon-built AI CPUs for agentic AI workloads. The snippet confirms CPUs, not GPUs, and a scale of “millions”; the post does not disclose chip model, price, delivery timeline, or deployment details. The signal to watch is agent workloads pulling demand beyond GPUs.

Why it matters: Meta buying millions of Amazon AI CPUs is an unusual infra move, so HKR-H and HKR-R are strong. HKR-K clears because the story gives scale, chip class, and agentic-workload use, but model, price, delivery, and deployment details are undisclosed, so it stays in the 78–84 band.

The Verge · AI

China’s DeepSeek previews new AI model a year after jolling US rivals

DeepSeek released a preview of its open-source V4 model on Friday and said it can compete with closed systems from Anthropic, Google, and OpenAI. The RSS snippet says V4 improves coding and highlights compatibility with Huawei tech; parameter count, benchmark scores, and rollout details are not disclosed. The part to watch is the pairing of agent-focused coding gains with tighter alignment to China’s domestic chip stack.

Why it matters: This is a flagship Chinese model update with HKR-H/K/R: a new open-source V4 preview, coding gains, and Huawei compatibility. It stays below the 85 band because the story withholds params, benchmark scores, and launch timing.

Synced · WeChat

Anthropic confirms three bugs caused Claude Code's apparent quality drop

Anthropic said Claude Code's quality drop over the past month came from 3 harness and prompt issues, while model capability itself and the Claude API were unchanged. The issues were a Mar. 4 default reasoning shift from high to medium, a Mar. 26 session-cache bug, and an Apr. 16 25/100-word prompt limit; fixes or rollbacks landed on Apr. 7, Apr. 10, and Apr. 20.

Why it matters: Anthropic published a concrete postmortem for Claude Code regressions with three dated causes and fixes, so HKR-H/K/R all pass. It matters to a Claude-heavy developer audience and affects multiple Sonnet/Opus versions, but it remains an incident report, not a market-wide model or

Latent Space

GPT 5.5 and OpenAI Codex Superapp

OpenAI launched GPT-5.5 for ChatGPT and Codex, while API access is delayed for safeguards. The post cites 82.7% Terminal-Bench 2.0, 58.6% SWE-Bench Pro, and a 1M API context window. The sharper signal is Codex: browser control and Prism integration point to a desktop superapp strategy.

Why it matters: All HKR axes pass: GPT-5.5 is a major OpenAI model update with benchmark numbers and API conditions. Codex plus browser control and Prism raises the coding-agent stakes; this fits the Claude 4.7-level 85–94 band.

X · @dotey

DeepSeek releases and open-sources V4 preview; 1M context is standard across all services

DeepSeek released and open-sourced the V4 preview, making 1M context standard across all official services with no tier or price split. The post says V4-Pro and V4-Flash use token compression plus DSA sparse attention to cut compute and memory costs for 1M context; legacy APIs remain for 3 months and stop after July 24.

Why it matters: DeepSeek is a flagship Chinese model vendor, and this V4 preview is a substantive release with open source and 1M context made standard across official services. HKR-H/K/R all pass: the post includes mechanisms and a migration deadline, and the tier reset makes it a same-day P1.

X · @op7418

DeepSeek V4 detailed official announcement is out

DeepSeek says V4 Pro has 1.6T total parameters with 49B active, while Flash has 284B total and 13B active; both were pretrained on 32T tokens. Web and app Expert mode map to Pro, and Fast mode maps to Flash. The post also says several benchmarks are on par with Opus 4.6, with stronger agent ability and world knowledge, plus a new attention mechanism that reduces compute and memory demand.

Why it matters: This is a flagship DeepSeek release, scored on par with peer US lab model launches. HKR-H/K/R all pass on concrete scale numbers, 32T data, and an inference-efficiency mechanism; benchmark setup, pricing, and API availability are not disclosed in the summary.

Computing Life · Share · Yage

Skills Are Products With Built-in Suicide Genes

The author argues Anthropic Skills cannot stand alone as paid products, citing direct sales, hosting, and API funneling as 3 dead ends. The post cites PromptBase at about $5M annual revenue, Stripe’s 2.9% plus 30 cents fee, and Snyk finding 13.4% of skills with critical issues. The sharper point is charging for relationships, time-sensitive access, physical accountability, and judgment.

Why it matters: HKR-H/K/R all pass: the hook is sharp, and the post tests three business paths with named examples. It is strong commentary, not a new Anthropic release, so it lands at the featured threshold rather than 78+.

Hugging Face Blog

DeepSeek-V4: a million-token context that agents can actually use

DeepSeek released V4 with two MoE checkpoints, Pro and Flash, both supporting a 1M-token context. Pro has 1.6T total and 49B active parameters; Flash has 284B total and 13B active. The key detail is KV cost: Pro uses 27% of V3.2 single-token FLOPs and 10% of its KV cache; Flash uses 10% and 7%.

Why it matters: DeepSeek-V4 is a flagship Chinese model release with 1M-token context and KV cache at 7%–10% of V3.2. HKR-H/K/R all pass, placing it in the 85–94 same-day band.

Ruan YiFeng's Weblog

Tech Weekly Issue 394: The Second Wave of API Opening

Ruanyifeng’s Weekly Issue 394 argues that production-ready LLMs in H2 2025 triggered a second API-opening wave. The post says agents need platform APIs to act, citing Tencent opening WeChat interfaces after OpenClaw and adoption of MCP and Skills. The key shift is consumer services exposing actions, not only cloud APIs.

Why it matters: HKR-H/K/R all pass: the historical API-wave frame is clickable, and the post gives mechanisms around agent action APIs, MCP/Skills, and WeChat access. This is strong commentary, not a model or major product release, so it stays in the 72–77 band.

The Verge · AI

Claude is connecting directly to personal apps like Spotify, Uber Eats, and TurboTax

Anthropic added personal app connectors to Claude, covering services such as Spotify, Uber, AllTrails, Instacart, and TurboTax. After connection, Claude can suggest relevant apps inside chats, such as using AllTrails for hike recommendations; the post does not disclose launch count, regions, or plan access. The key shift is Claude moving from work apps into personal consumer workflows.

Why it matters: This gets Anthropic’s positive signal: a substantive product update, but not a model release. HKR-H/K/R all pass because personal-app connectors are a strong hook, the story confirms in-chat app invocation, and it hits the fight for assistant entry points; missing pricing, region

X · @dotey

Anthropic launches memory for Claude Managed Agents in public beta

Anthropic has launched memory for Claude Managed Agents in public beta, letting agents retain and reuse experience across sessions. Memory is stored as files on a filesystem, with shared permissions, concurrent access, audit logs, and rollback; Rakuten reports a 97% drop in first-time errors, and Wisedocs reports 30% faster document validation. The key detail is the implementation path: it uses a filesystem, not a dedicated vector database.

Why it matters: Anthropic adds cross-session memory to Claude Managed Agents beta and discloses the implementation plus two user numbers: Rakuten 97% and Wisedocs 30%. HKR-H/K/R all pass, but the scope is still limited to the managed-agent beta, so this lands at 83 and featured.

X · @claudeai

Memory on Claude Managed Agents is now in public beta

Claude has put Memory for Managed Agents into public beta, and agents can now learn from every session. The post only says it uses an intelligence-optimized memory layer balancing performance and flexibility; it does not disclose capacity, retention, pricing, or access conditions. What matters for practitioners is when persistent memory becomes default and how it changes agent evals and state management.

Why it matters: Memory on Claude Managed Agents is a substantive Anthropic product update with clear practitioner resonance, so HKR-H and HKR-R pass. HKR-K is weak because the post omits capacity, retention, pricing, and default-on conditions, keeping it in low featured rather than p1.

Bloomberg Technology

An AI Agent Takes Over a Store and Orders Too Many Candles

Andon Market in San Francisco’s Cow Hollow put store operations under an AI agent named Luna, which handles assortment and pricing, and the headline says it over-ordered candles. The RSS snippet only confirms Luna acts like a CEO; the post does not disclose the candle quantity, failure mechanism, financial impact, or remediation. The real signal is that a retail operating loop was delegated to an agent.

Why it matters: Bloomberg reports a real store delegating assortment and pricing to an AI agent, turning agent risk into a concrete incident. HKR-H and HKR-R pass, but HKR-K is limited because quantity, loss, trigger, and rollback are undisclosed, so this sits at the low end of featured.

X · @dotey

Codex now supports GPT-5.5 and adds five capability upgrades

Codex now supports GPT-5.5 and adds 5 upgrades aimed at moving it from a coding tool to an agent that can execute longer tasks. The RSS snippet says it can control browsers and computers, create files in Microsoft Office and Google Drive, and use gpt-image-2; an auto-review mode invokes a separate review agent for high-risk actions. What matters is longer task chains, but the post does not disclose pricing, rollout scope, or safety thresholds.

Why it matters: This is a substantive Codex product update: the main signal is the shift toward an agent that can execute chained tasks, not just a new model toggle. HKR-H/K/R all pass, but the item is second-hand and omits pricing, rollout scope, and safety thresholds, so it lands as featured,