Skip to content

#Agent

36 today

Apr 15Wednesday

OpenAI News

The next evolution of the Agents SDK

OpenAI published a post about the next evolution of the Agents SDK. Only the title is available, with no body text or details, so specific features, numbers, and timing cannot be confirmed. For AI developers, it signals continued updates to the Agents SDK, but the scope is unclear from the source provided.

Why it matters: This is a substantive OpenAI developer-platform update: the post confirms native sandbox execution, a stronger agent-loop harness, and harness/compute separation, so HKR-H/K/R all pass. It stays below P1 because pricing, rollout scope, and performance numbers are not disclosed in

X · @dotey

pi maintainer Mario Zechner sets a new rule: unapproved issues and PRs will be auto-closed immediately

pi maintainer Mario Zechner says any issue or PR submitted without prior approval will be auto-closed, after he started receiving 30 to 50 issues per day and most were AI-agent spam. He will still review closed submissions daily; strong issues can earn an “lgtmi” tag, and strong issue-plus-fix PRs can earn “lgtm,” exempting future submissions from auto-close. The shift to watch is simple: open source projects are raising contribution gates to filter zero-cost AI-generated noise.

Why it matters: Featured on strong HKR-H/K/R: a maintainer-level policy change with concrete spam numbers and a review mechanism. Importance stays in the mid-70s because the blast radius is mainly the OSS agent/dev community, not a major model or platform release.

最佳拍档 (BestPartners)

Will OpenClaw Go Closed Source? Peter Steinberger on OpenClaw at AI Engineer

Peter Steinberger said at the April 9, 2026 AI Engineer event that OpenClaw will not go closed source; the project reached nearly 30,000 commits and almost 2,000 contributors in 5 months. The talk says OpenClaw logged 1,142 security reports, 99 marked critical, 469 public with a 60% closure rate, and Fast Mode cut his parallel sessions from nearly 10 to 5-6. The key signal is the operating model: local-first, model-neutral, and a foundation for security maintenance; the post does not disclose a release date or implementation details for Dreaming.

Why it matters: HKR-H/K/R all pass: the close-source question is a strong hook, and the talk adds concrete stats on contributors, advisories, and Fast Mode. The score stays near the featured floor because this is a YouTube recap, and several teased items lack mechanism or release details.

X · @dotey

Anthropic's Anthony Morris says Claude Code desktop has been rebuilt from the ground up

Anthropic's Anthony Morris said Claude Code desktop was rebuilt from the ground up to make it easier to run multiple Claude coding tasks in parallel within one repository. The post cites Git worktree isolation as the mechanism: each session gets an independent code copy, with changes kept separate until merge, plus visual diff review, app preview, and a plugin marketplace. The workflow shift matters more than the headline, but the post does not disclose release timing, performance data, or supported platforms.

Why it matters: This is a substantive Claude Code product update aimed at a real workflow pain point: parallel coding sessions in the same repo. HKR-H/K/R all pass through the strong hook, concrete worktree-based mechanism, and developer resonance, but missing launch date, performance data, and

X · @op7418

Claude Code's newly released routines feature looks strong

Claude Code released routines, which package prompts, repos, environments, and connectors into cloud automation triggered by schedules, HTTP API, or GitHub events. Each trigger starts a full Claude Code cloud session that can run shell, use repo skills, and access external services, then hand work back to local follow-up. The post does not disclose pricing, quotas, or supported platforms.

Why it matters: This is a substantive Claude Code workflow update: routines package prompts, repo, environment, and connectors into cloud jobs triggered by schedule, HTTP API, or GitHub events. HKR-H/K/R all pass, but price, quota, and supported platforms are not disclosed, so it stays featured,

X · @dotey

Claude Code adds Routines for trigger-based automated tasks

Anthropic added Routines to Claude Code in research preview, letting preset tasks run in the cloud via 3 triggers: schedules, GitHub events, and API calls. The post cites auto doc sync on release-branch merges and code review on PRs; Pro, Max, Team, and Enterprise users can access it, but the daily run cap is not disclosed. The key detail is permissions: every Routine acts as the user, including GitHub commits and Slack messages.

Why it matters: This is more than a minor feature tweak. Routines moves Claude Code toward an event-driven cloud agent, with concrete details on 3 triggers, plan availability, and a user-identity permission model, so HKR-H/K/R all pass. It stays below p1 because this is still a research preview,

X · @claudeai

Now in research preview: routines in Claude Code

Anthropic launched routines in research preview for Claude Code: configure a prompt, repo, and connectors once, then run it on a schedule, via API, or from an event. Routines run on Anthropic web infrastructure, so a laptop does not need to stay open; the post does not disclose pricing, quotas, or rollout scope. The key point is hosted execution, not one-off code completion.

Why it matters: This is a substantive Claude Code expansion from local interactive coding to hosted, scheduled, and event-driven execution. HKR-H/K/R all pass, and the Anthropic update gets a policy bump, but price, quotas, and rollout scope are not disclosed, so it stays featured rather than P1

Apr 14Tuesday

最佳拍档 (BestPartners)

Global GPU shortage worsens: H100 rental prices rose nearly 40% in five months

SemiAnalysis says Nvidia H100 one-year rental pricing rose from $1.70 to $2.35 per GPU-hour between Oct 2025 and Mar 2026, up nearly 40% in five months. The post attributes this to Anthropic-driven demand, multi-agent and media generation workloads, and memory cost spikes, with LPDDR5 and DDR5 contract prices up about 4x and 5x year over year; much new capacity is already prebooked. The key variable is the supply gap, not Blackwell refreshes alone.

Why it matters: Strong HKR-H/K/R: the story has a sharp price-shock hook, concrete market data, and clear resonance with compute-cost anxiety. It stays below P1 because this is a secondary video synthesis of a SemiAnalysis report, not a primary company or product announcement.

X · @dotey

Rather than AI First, this is really Software Engineering First

The post argues “AI First” is an engineering problem: if AI writes code in 2 hours, review, testing, deploy, monitoring, and rollback must also run automatically, with humans kept at key decision points. Its concrete prerequisites are automated tests, CI/CD, A/B testing, production monitoring, task management, and a clear architecture; without them, a 25-person team just shifts bottlenecks from coding to QA and ops. The real boundary is use case fit: API services, data platforms, and internal tools fit better than complex UI, core products, or high-security systems.

Why it matters: This is a strong practitioner commentary rather than a news event. HKR-H lands on the contrarian framing, HKR-K on concrete prerequisites and scope limits, and HKR-R on the bottleneck-shift argument; it stays in the mid-70s because there are no named cases, first-person tests, or

X · @dotey

Vercel open-sources Open Agents, a reference implementation for enterprise coding agent platforms

Vercel open-sourced Open Agents as a forkable reference for enterprise coding-agent platforms, with a three-layer architecture and features like voice input and PR creation. Its key design keeps the agent outside the sandbox and uses tools such as file I/O, shell, and search to control execution; the post also cites Anthropic Managed Agents pricing at $0.08 runtime per hour and $10 per 1,000 web searches. The part to watch is the agent-sandbox split, not the packaging choice.

Why it matters: This fits the 78–84 band: a notable open-source coding-agent framework with concrete architecture, remote sandbox operation, and Anthropic pricing, so HKR-H/K/R all land. It stops short of must-write status because this is strong infra reference material, not a model or industry-

最佳拍档 (BestPartners)

Meta-Harness: Can harness engineering code self-iterate? A Stanford paper analysis

Stanford, MIT, and KRAFTON AI present Meta-Harness, which turns harness optimization into an outer-loop search and beats manual or text-optimization baselines on 3 task types. The system uses a coding agent to inspect filesystem history; after 10 search iterations, the data exceeds 10 million tokens, and on online text classification it matched OPRO’s 60-iteration result in 4 iterations while reaching 75.9% average accuracy on 5 OOD datasets. The key point is full-feedback retention rather than compression; the paper also reports about 20 TerminalBench-2 iterations at a total cost of a few hundred dollars.

Why it matters: This is a good research-release explainer for agent builders: the mechanism is clear and the post includes concrete numbers, so HKR-H/K/R all pass. It stays at 80 because the source is a secondary YouTube summary, not the primary paper or official release, and the impact is still

Apr 13Monday

最佳拍档 (BestPartners)

2027 Is the Enterprise AI Singularity Year: Sundar Pichai on 10 Years as Google CEO, Transformer and Search

Sundar Pichai said in a Stripe interview that Alphabet plans $175B-$185B in 2026 capex and that 2027 will be the breakout year for enterprise AI agent workflows. He said Google cut Search latency by 30% over five years while adding AI features, manages teams with 10 ms or 30 ms latency budgets, and sees 2026-2027 constrained by wafers, memory, power, and permitting. The point to watch is not search replacement but search evolving into an agentic manager, while TPU allocation has become Google's scarcest internal resource.

Why it matters: High-signal executive commentary rather than a product launch. HKR-H/K/R all pass on the 2027 agent call, concrete capex and latency details, and the search-plus-compute nerve hit; score stays below P1 because this is a second-hand recap, not the primary interview.

Apr 12Sunday

X · @dotey

UC Berkeley team used a cheating AI to break 8 major agent benchmarks and score near perfect without solving tasks

A UC Berkeley team used a cheating AI with no LLM calls to break 8 major agent benchmarks, scoring 73% to 100% without solving tasks. The post cites three cases: a 10-line Python hook bypassed SWE-bench tests across 500 tasks, WebArena exposed answers via file://, and FieldWorkArena gave full credit to an empty {} reply. The real issue is benchmark isolation failure; the team is turning its scanner into the open-source BenchJack project.

Why it matters: HKR-H/K/R all pass: the claim is clicky, concrete, and directly threatens trust in agent evals. I stop at 84, not 85+, because the current input is a social summary; paper status, full methods, and outside replication are not disclosed here.

X · @Yuchenj_UW

MiniMax M2.7 is open-source!

MiniMax open-sourced M2.7 and said its research agent now handles 30%–50% of the R&D workflow. The post says the agent covers literature review, experiment orchestration, log debugging, code fixes, and merge requests; M2.7 also rewrote its own harness for 100+ automated rounds, with a 30% gain on internal coding evals.

Why it matters: HKR-H/K/R all pass: open-sourcing plus a research agent doing 30%-50% of R&D is a strong hook, and the post includes 100+ self-rewrite loops with +30% internal coding eval. It stays at 78 because license, repo, benchmark context, and external reproduction are not disclosed.

Apr 11Saturday

X · @dotey

Anthropic launches Claude Managed Agents beta; Michael Cohen explains secure third-party key management for agents

Anthropic added Vaults to the Claude Managed Agents beta to manage each end user's third-party credentials with a per-user vault_id and automatic injection at session runtime. The post shows a three-step flow—create a Vault, bind credentials to an MCP server address, and pass vault_id when creating a session—and prices CMA at token usage plus $0.08 per session-hour. The key design is isolation: credentials never enter Claude's context window, code runs in a sandbox, auth goes through a dedicated proxy, and the harness cannot access secrets.

Why it matters: This adds the missing implementation detail for Claude Managed Agents: third-party credential isolation. HKR-H/K/R all pass via a concrete security hook, reproducible vault_id flow, pricing, and a real operator pain point; impact stays at the developer integration layer, so it is

X · @dotey

OpenAI Codex team's Nick Baumann: build dedicated CLI tools for AI instead of feeding messy data repeatedly

OpenAI Codex engineer Nick Baumann says teams should wrap repeated data access into parameterized CLI tools with JSON output instead of repeatedly dumping logs, docs, and API responses into Codex. The post lists 3 examples in daily use: codex-threads for past sessions, slack-cli for threaded Slack search, and typefully-cli for posting workflows; access still goes through the existing auth gateway. The point for practitioners is narrower interfaces: models handle focused commands more reliably than raw, noisy source data.

Why it matters: This is a practical workflow note from an OpenAI Codex team member, not a formal launch, but it offers a reusable mechanism: wrap noisy context behind parameterized JSON-returning CLIs and shows 3 live examples. HKR-H/K/R all land; no benchmark, scale, or major product release,so

QbitAI · WeChat

OpenClaw-style methods reach multimodal generation, with a 6B model beating Nano Banana 2 on some tasks

A team led by Shanghai AI Laboratory introduced GEMS, adding Agent Loop, Memory, and Skills to multimodal generation, and reports that 6B Z-Image-Turbo beats Nano Banana 2 on some tasks. The post reports +14.22 average gains on 5 mainstream tasks and +8.92 over the best baseline on 4 downstream tasks; the paper and code are public, but the post does not disclose Nano Banana 2's full setup.

Why it matters: Strong HKR-H/K/R: the hook is a 6B multimodal model beating Nano Banana 2, and the post includes mechanism plus testable deltas (+14.22 / +8.92) with paper and code. It stays below P1 because the article does not disclose the full Nano Banana 2 comparison setup.

X · @dotey

Anthropic launches Claude for Word beta add-in

Anthropic released a beta Claude for Word add-in for paid Claude Team and Enterprise users, with direct sidebar editing for .docx and .docm files. Edits appear in Word’s native track changes flow, the add-in can reuse conversation context from Excel and PowerPoint, and it supports reference uploads plus reusable team Skills. The key point is shared context across Office apps; the post does not disclose pricing, regions, or a wider rollout timeline.

Why it matters: This is a substantive Anthropic product update for Team and Enterprise, not a generic integration post. HKR-H/K/R all pass on novelty, concrete mechanics, and workflow resonance, but the beta scope is limited and price, regions, and GA timing are undisclosed, so it lands in mid-"

X · @dotey

Claude Code adds ultraplan: start planning in terminal, review in browser, then run in cloud or locally

Claude Code opened a preview of ultraplan to users with the web app enabled, requiring v2.1.91+, and planning starts from /ultraplan in the terminal. Claude drafts a plan in the cloud after reading the repo, users review and annotate it in the browser, then choose cloud execution with a PR or local terminal execution. The key change is splitting planning from execution: planning moves to the cloud without blocking the terminal, and the post says token use is close to local plan mode.

Why it matters: This is more than a routine feature add: Claude Code splits planning from execution, with /ultraplan in terminal, cloud-side repo reading, browser review, and cloud PR or local execution. HKR-H/K/R all pass, with a Claude-specific bump, but it is still a preview and sourced froma

Apr 10Friday

最佳拍档 (BestPartners)

LLM self-evolution: Shinka Evolve, AlphaEvolve, and sample efficiency

Sakana AI open-sourced Shinka Evolve and uses a UCB bandit to switch among GPT-5, Claude Sonnet 4.5, Gemini, and others, aiming to cut the thousands of program evaluations common in AlphaEvolve-style search. The post says it beat AlphaEvolve’s classic circle-packing result with fewer evaluations and adds full-file rewrites, crossover, editable-region guards, and a meta-notebook; the post does not disclose exact metrics, cost, or the repo link. The part to watch is surrogate-task design and hard verification: the system still needs humans to define problems.

Why it matters: Featured, not P1: HKR-H/K/R all pass. The piece has a strong hook, concrete mechanisms like UCB model routing and program crossover, and a real nerve around eval cost and hard verification. It stays at 80 because key metrics, cost, and the primary release link are not disclosed.

QbitAI · WeChat

Claude bug mixes up speaker roles, issues self-instructions, and blames the user

A developer said Claude 3.5 and Claude 4 can confuse user, assistant, and system roles under complex or malicious context, and the Hacker News post drew heavy discussion. The post cites inputs like <stop> and <end prompt> as a repro clue; Anthropic's fix status and scope are not disclosed. The real issue is control-data separation, not a single prompt failure.

Why it matters: This clears all HKR axes: the angle is clickworthy, the post includes a concrete repro clue, and the failure mode matters to anyone shipping agents. I kept it below P1 because scope, affected versions, and Anthropic’s fix status are not disclosed.

X · @dotey

Anthropic launches Advisor Tool API: cheaper models execute while pricier models advise on hard decisions

Anthropic launched the advisor tool API, letting Sonnet or Haiku execute tasks and consult Opus on hard decisions; it is in beta and requires the anthropic-beta: advisor-tool-2026-03-01 header. The RSS snippet says Sonnet+Opus gains 2.7 points on multilingual SWE-bench while cutting per-task cost by 11.9%; Haiku+Opus rises from 19.7% to 41.2% on BrowseComp at 15% of Sonnet's cost. The key detail is the call path: model switching happens inside one Messages API request, advisor and executor tokens are billed separately, and max_uses caps consultations.

Why it matters: This is a substantive Anthropic API update with concrete mechanics: in-request model routing, separate token billing, max_uses, and two benchmark/cost deltas. HKR-H/K/R all pass, so it merits featured, but it is still below a model-release tier event.

X · @claudeai

We're bringing the advisor strategy to the Claude Platform.

Claude is adding the advisor strategy to Claude Platform, with Opus as the advisor and Sonnet or Haiku as the executor. The RSS snippet says this yields near-Opus-level agent intelligence at lower cost; the post does not disclose pricing, benchmark scores, or rollout timing.

Why it matters: Anthropic ships a substantive Claude Platform update, and HKR-H/K/R all pass: the Opus-advisor plus Sonnet/Haiku-executor setup is novel, concrete, and directly relevant to agent builders. The score stays below P1 because price, benchmarks, and rollout timing are not disclosed.

Apr 9Thursday

QbitAI · WeChat

Beyond MoE, Tencent introduces MoT: a 2B embodied model ranks first in 16 of 22 evaluations

Tencent Hunyuan and Robotics X released HY-Embodied-0.5; its MoT-2B uses 4B total params with 2B active and ranks first in 16 of 22 embodied evaluations. The post says it uses 100M+ embodied data, 600B+ pretraining tokens, 30M+ mid-training samples, plus visual latent tokens, bidirectional attention, RFT, RL, and online distillation. The key point is a rebuilt edge-oriented embodied stack, not a simple VLM fine-tune.

Why it matters: Strong on HKR-H/K/R: the headline has a real hook, the body includes concrete numbers and training mechanisms, and the edge-robotics angle lands with practitioners. I keep it at 83, not 85+, because this is a high-quality embodied-model release, not a broad same-day industry-def

QbitAI · WeChat

Claude launches managed agent service, then faces an open-source alternative from Multica

Anthropic has opened Claude Managed Agents and charges $0.08 per session-hour plus token usage. The service supports hours-long runs, sandboxing, checkpoint recovery, and multi-agent orchestration; web search costs $10 per 1,000 searches, while some memory and orchestration features remain in research preview. The title mentions a “lobster ban,” but the post does not disclose that context; the real shift is Anthropic selling agent infrastructure to enterprises.

Why it matters: This is a substantive Anthropic product update: a managed-agent service with explicit pricing, sandboxing, checkpoint restore, and search costs, so HKR-K is strong. HKR-R is real for Claude-heavy teams weighing build vs. buy, but the scope is smaller than a model launch, so it is

X · @op7418

Meta releases Muse Spark model

Meta released the Muse Spark model with native multimodal reasoning, tool use, visual chain-of-thought, and multi-agent orchestration, but it is only available in the Meta AI app and is not open source for now. The snippet says its Contemplating mode coordinates multiple parallel agents for reasoning, and its Artificial Analysis score is below Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. The post does not disclose model size, pricing, or rollout timing.

Why it matters: A major-lab model launch plus the “poached team’s first output” angle lands HKR-H/K/R. The score stays near the featured floor because the post offers capability claims and relative benchmark placement only; params, pricing, rollout timing, and access scope are not disclosed.

X · @dotey

Anthropic launches Claude Managed Agents, a managed API for building and deploying agents, now in public beta

Anthropic launched Claude Managed Agents, a managed API for building and deploying agents, in public beta. It offers a production sandbox, long-running sessions, and multi-agent coordination; Anthropic says internal tests showed up to a 10-point success-rate gain on structured file-generation tasks versus standard prompt loops. Pricing uses standard Claude token fees plus $0.08 per active session-hour; the real signal is Anthropic moving agent infrastructure into its platform layer.

Why it matters: Anthropic packaged managed agents, sandboxing, and long-running sessions into a public-beta API, which is a real workflow update for developers. HKR-H/K/R all pass: strong platform hook, concrete facts like a 10-point gain and $0.08 per hour, and clear resonance around developer-

X · @claudeai

Introducing Claude Managed Agents: everything you need to build and deploy agents at scale.

Claude has launched Claude Managed Agents in public beta on Claude Platform, claiming to compress the path from agent prototype to launch into days. The post discloses only a performance-tuned agent harness plus production infrastructure; pricing, toolchain support, model scope, and quotas are not disclosed.

Why it matters: Anthropic gets a positive bump, and HKR-H/HKR-R pass because managed agent deployment is a strong hook for Claude-heavy builders. HKR-K is limited: the post discloses a harness and prod infra, but not pricing, toolchain support, model scope, or quotas.

Apr 8Wednesday

MIT Technology Review · AI

Mustafa Suleyman: AI development won’t hit a wall anytime soon—here’s why

Mustafa Suleyman argues frontier AI training compute rose from about 10^14 to over 10^26 FLOPs since 2010, a 1 trillion-fold increase, so AI development is not near a wall. He cites a 7x Nvidia chip gain in six years, 3x more HBM3 bandwidth, and Epoch AI estimates that compute needed for fixed performance halves every eight months. The piece is commentary from Microsoft AI’s CEO, not an independent study; the post does not disclose a reproducible basis for the 200GW-by-2030 claim.

Why it matters: HKR-H/K/R all pass: Suleyman takes a hard line in the scaling-wall debate and cites 10^26 flops, 7x chip gains, 3x bandwidth, and 8-month efficiency halving. Held at 82 because this is executive commentary, not independent research, and the 2030 200GW math is not disclosed.

X · @dotey

Hermes Agent is gaining traction; I installed it and the experience was decent

Nous Research open-sourced Hermes Agent in late February, and the post says it reached nearly 30,000 GitHub stars in under two months. The post describes a closed learning loop: after complex tasks with 5+ tool calls, Hermes writes Markdown skills, with one Reddit report claiming 3 skills in 2 hours and a 40% speedup on repeated research work. The key angle is its self-hosted agent engine that combines skill generation, SQLite-based memory retrieval, and five-layer safety controls.

Why it matters: HKR-H/K/R all pass: the piece combines strong OSS momentum, concrete mechanics, and a real builder nerve around self-hosted learning agents. It stays at 78 because the evidence is mostly social commentary and light user feedback, not a primary release or broad independent eval.

Latent Space

Extreme Harness Engineering for Token Billionaires: 1M LOC, 1B toks/day, 0% human code, 0% human review

OpenAI Frontier says it built an internal beta over five months with a repo above 1M LOC, over 1B tokens per day, and 0% human-written or human-reviewed code before merge. The post says the team treated failures as missing capability, context, or structure, then used Symphony orchestration, specs, tests, observability, and sub-1-minute build loops to constrain Codex. The shift to watch is from humans reviewing code to humans designing the harness; the $2k-$3k/day cost is cited secondhand in the post.

Why it matters: HKR-H/K/R all pass: the headline is clickworthy, and the piece includes concrete workflow details plus scale numbers. It stays below p1 because this is an interview-style report, not an official launch, and key claims like 1B tokens/day and cost lack independent verification.

Apr 7Tuesday

X · @dotey

Milla Jovovich and Ben Sigman release open-source AI memory system MemPalace, claim perfect LongMemEval score

Milla Jovovich and Ben Sigman released the open-source memory system MemPalace and claimed a perfect LongMemEval score. The project runs fully local with no cloud or API key, says AAAK compresses context 30x, and uses 19 MCP tools for retrieval. The key issue is evaluation: Penfield Labs says the “perfect” result measured retrieval only, not end-to-end QA, and AAAK dropped retrieval accuracy from 96.6% to 84.2%.

Why it matters: HKR-H lands on the celebrity/open-source hook and the 'perfect score' dispute. HKR-K/R land on concrete metrics and the familiar nerve of eval gaming vs real memory utility; source authority is still just an X post, so this stays featured, not higher.

Latent Space

[AINews] Gemma 4 crosses 2 million downloads

Google’s Gemma 4 reached about 2 million downloads in its first week. The post compares that with Gemma 3 at 6.7 million over the past year, Gemma 2 at 1.4 million since June 2024, and Qwen 3.5 at about 27 million in roughly 1.5 months. The signal for practitioners is local deployment: one iPhone 17 Pro demo ran Gemma 4 E2B at about 40 tok/s via MLX, with support across Hugging Face, vLLM, llama.cpp, Ollama, and NVIDIA.

Why it matters: HKR-H/K/R all pass: the story has a clean hook, concrete comparative download data, and a real open-model adoption nerve. It stays low-featured because this is a secondary-source uptake snapshot, not a primary Google release or a substantive capability update.

MIT Technology Review · AI

The one piece of data that could actually shed light on your job and AI

University of Chicago economist Alex Imas argues that AI job displacement depends less on task exposure and more on industry-level price elasticity data; the piece cites OpenAI estimating real estate agents as 28% exposed. It adds that the US task catalog started in 1998, and Anthropic compared it with millions of Claude chats in February. The key variable is whether lower prices raise demand enough, and the post does not disclose any economy-wide dataset yet.

Why it matters: Strong HKR-K: it reframes job impact around price elasticity, with concrete anchors like OpenAI's 28% exposure for real-estate agents and Anthropic's O*NET-to-Claude mapping. HKR-R is clear because it hits job displacement anxiety, but this is commentary, not a fresh dataset or a

Apr 6Monday

X · @dotey

Xiaomi MiMo lead Luo Fuli on token costs in the Agent era

Luo Fuli said Agent workloads can resend 100k+ tokens across repeated tool calls, and global compute cannot keep up with that burn. She said OpenClaw makes several times more requests than Claude Code and can push real API cost to tens of times the subscription price; the post does not disclose a pricing formula.

Why it matters: A named Xiaomi MiMo lead makes a concrete, testable critique of agent cost: 100k+ token context replay, multi-tool-call overhead, and several-times request inflation vs Claude Code. HKR-H/K/R all pass, but missing public benchmark setup and pricing keeps it at the low end of the

Apr 4Saturday

X · @dotey

Mintlify uses ChromaFs to make AI document retrieval look like a file system

Mintlify routes its AI doc assistant’s grep, cat, and ls calls through ChromaFs into database queries, cutting session startup from 46s to 100ms and pushing marginal compute cost per chat near zero. Built on Vercel Labs’ just-bash, it maps pages to files and sections to directories; at 850,000 chats per month, replacing real sandboxes saves over $70,000 a year in compute. The real shift is retrieval design: not faster vector RAG, but model-led exploration of structured docs, and the post says this may not fit messy knowledge bases.

Why it matters: This is a substantive engineering write-up, not a routine product note. HKR-H/K/R all pass: the fake-filesystem angle is novel, the post includes hard numbers (46s→100ms, 850k chats/month, >$70k/yr), and it hits operator concerns around latency, cost, and retrieval design; strong

Latent Space

Marc Andreessen introspects on The Death of the Browser, Pi + OpenClaw, and Why “This Time Is Different”

Marc Andreessen argues in a 76-minute interview that this AI cycle differs from 2016 because of reasoning, coding, agents, and recursive self-improvement. The post gives one concrete mechanism: Pi/OpenClaw as LLM + shell + filesystem + markdown + cron loop; it mentions “death of the browser,” but does not disclose a verifiable timeline or product plan. The sharper point is his Unix-like framing of file-backed agent state and portability.

Why it matters: This is a strong commentary piece, not a market-moving event. HKR-H comes from the browser-death hook, HKR-K from the Pi+OpenClaw mechanism, and HKR-R from the interface/distribution nerve; lack of roadmap, metrics, or launch details keeps it at the low end of featured.

Apr 3Friday

X · @op7418

Alibaba released the Qwen 3.6 Plus model

Alibaba released Qwen 3.6 Plus with a 1M context window, 64K input, and nearly 991K max output. The RSS snippet says it improves over Qwen 3.5 on agents, coding, image, and document understanding, priced at RMB 2 per 1M input tokens and RMB 12 per 1M output tokens; benchmark scores and test conditions are not disclosed.

Why it matters: Alibaba shipping Qwen 3.6 Plus is a substantive domestic model update. HKR-H/K/R all pass on the 1M-context plus pricing combo, but it stays below P1 because benchmark scores, baselines, and test conditions are not disclosed in the body.

X · @op7418

Karpathy shared how he builds a local AI knowledge base

Karpathy uses Obsidian and local Markdown to build a personal wiki, stores source material in a RAW folder, then has an LLM generate summaries, indexes, concept pages, links, and visualizations. The setup can answer questions over the wiki and write reports or new files, but the post also says AI-generated content can pollute the corpus and should be separated from trusted sources; the post does not disclose the model, scale, or automation details.

Why it matters: HKR-H and HKR-R land because Karpathy’s local-first wiki workflow is inherently clickable and discussable for AI practitioners. HKR-K lands on the RAW→LLM→summary/index/link mechanism, but missing model, corpus size, and automation details keep it in the mid-70s.

X · @op7418

Google releases Gemma 4 for on-device use under Apache 2.0

Google released Gemma 4 in four variants—E2B, E4B, 26B MoE, and 31B Dense—targeting phones, edge devices, and up to single-H100 workstations. The RSS snippet says the 26B MoE activates 3.8B parameters and adds native function calling, JSON output, multimodal I/O, speech-to-text, and Apache 2.0 licensing; the post does not disclose benchmarks, context length, or rollout details.

Why it matters: Google releasing Gemma 4 is a substantive open-model update. HKR-H/K/R all pass on the size spread, 3.8B-active MoE detail, and deployment-cost relevance; it stays at 81 because benchmarks, context window, and test conditions are not disclosed here.