Skip to content

#Agent

36 today

Apr 20Monday

Import AI (Jack Clark)

Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4

Import AI 454 covers HiFloat4, Anthropic automated alignment R&D, and a Chinese model safety study. HiFloat4 reached about 1.0% relative BF16 loss on Ascend NPUs, versus MXFP4's about 1.5%. Anthropic's Claude Opus 4.6 AARs used 800 hours and about $18,000 to raise PGR from a 0.23 human baseline to 0.97.

Why it matters: HKR-H/K/R all pass: Jack Clark links Anthropic AAR, HiFloat4, and Chinese model safety with hard numbers on cost, PGR, and loss. It is strong research commentary, not the original release, so it fits 78–84.

Synced · WeChat

How to Do Vibe Coding Correctly? A Masterclass from Anthropic's Coding Agent Lead

Anthropic researcher Erik Schluntz said his team merged a 22,000-line production change, mostly written by Claude, cutting work from two weeks to one day. His workflow spends 15-20 minutes on repo exploration and planning, limits edits to leaf nodes, keeps humans on core logic, and validates with long stress tests plus a few E2E tests. The key issue is boundary control, not handing AI the system core; he also said task length AI can handle doubles about every seven months.

Why it matters: HKR-H/K/R all pass: this is an Anthropic field report with concrete numbers and reproducible workflow rules for production coding agents. It stays at featured, not p1, because it is a strong practitioner lesson rather than a major model or product launch.

Synced · WeChat

In the first year of “deployment mode,” AgiBot expanded its rollout plans to seven solutions

AgiBot said at its April 17 Shanghai event that it released 4 robots, 6 AI models, and 7 standardized deployment solutions, and framed 2026 as the first year of embodied AI “deployment mode.” The post cites concrete metrics: Expedition A3 runs 8-10 hours, WITA Omni 1.0 targets sub-500ms interaction latency, and BFM was trained on 100 million-plus frames and 700 hours of motion-capture data; it also claims 5,100-plus shipments and 39% share in 2025, with the 10,000th robot rolling off in March 2026. The real point for practitioners is repeatable delivery rather than launch volume: the post lists 7 scenarios from 3C line loading to patrol, but independent validation details are not disclosed.

Why it matters: HKR-H/K/R all pass: the story leads with seven deployment playbooks and backs it with shipment, share, latency, and training figures. It stays at 76 because key outcome claims are company-sourced; customer impact and independent validation are not disclosed.

Xinzhiyuan · WeChat

Agent isn’t the key: RUC's AiScientist shows 23 hours and 74 rounds of long-horizon memory

A Renmin University of China team released AiScientist, which ran 23 hours and 74 experiment loops on MLE-Bench Lite Detecting Insults, raising validation AUC from 0.903 to 0.982 with 18 best-so-far updates. The paper says its core is File-as-Bus, which persists analysis, code, logs, and results in the workspace; removing it drops PaperBench by 6.41 points and MLE-Bench Lite Any Medal by 31.82 points. The real lever here is state continuity, not simply adding more agents.

Why it matters: HKR-H lands because the title flips a live assumption: memory continuity, not more agents. HKR-K lands on the 23h/74-run setup, AUC 0.903→0.982, and ablations; HKR-R lands because builders are debating multi-agent stacks vs durable state.

r/LocalLLaMA

Using Qwen3.6 via LM Studio as a Claude Code subagent, saving 30x Opus tokens per task

A Reddit user routed Qwen3.6 through LM Studio as a Claude Code subagent and reported about 30x lower Opus marginal tokens on two audit tasks. In the examples, a 23-file route audit dropped from 13k to 0.4k marginal tokens, and an 18-file Astro site inventory fell from 89k to 3k; the setup used unsloth’s Qwen3.6-35B-A3B-MXFP4_MOE gguf on a 64GB M4 Max with a 64k context window. The key mechanism is offloading extraction and audit work to a local OpenAI-compatible server, while the post also says quality was mixed rather than strictly better than Opus.

Why it matters: A named first-person experiment with 2 clear token comparisons hits HKR-H, HKR-K, and HKR-R: strong hook, concrete setup details, and direct cost relevance for Claude Code users. It stays below p1 because the evidence is a Reddit post with only 2 tasks.

Apr 19Sunday

r/LocalLLaMA

Same 9B Qwen weights: 19.1% in Aider vs 45.6% with a scaffold adapted to small local models

Using the same Qwen3.5-9B Q4 weights on the 225-task Aider Polyglot benchmark, the author changed only the scaffold and raised mean pass@2 from 19.11% to 45.56%. The little-coder setup is not a new model; it uses bounded reasoning, a write guard, explicit workspace discovery, and small per-turn skill injections. The key claim is scaffold-model fit, but the post reports only two full runs and does not disclose ablations, cross-model replications, or a second benchmark.

Why it matters: HKR-H/K/R all pass: the hook is a 2.4x jump on Aider Polyglot 225 with the same 9B Qwen weights, and the post names the scaffold mechanisms. Importance stays low-featured because evidence is thin: two full runs, no ablation, no cross-model rerun, and no second benchmark.

Synced · WeChat

MIA, a next-generation memory agent framework, aims to end agents' "amnesiac" workflows

A Shanghai Institute for Advanced Learning and ECNU team released MIA, a memory agent framework, and said it achieved the best results on 7 datasets. MIA uses a Manager-Planner-Executor design, dual parametric and non-parametric memory, alternating RL, and test-time continual learning; the post does not disclose exact benchmark scores. The key point is memory as capability internalization, not just retrieval, for open-world agents.

Why it matters: HKR-H/K/R all pass: the story targets agent memory, a real deployment pain point, and includes specific mechanisms. It stays below p1 because the article does not disclose per-dataset scores, baseline gaps, or enough reproduction detail.

Synced · WeChat

Amap debuts an autonomous embodied robot at the Yizhuang Marathon and showcases guide-assistance

Amap showed its quadruped robot Tutu at the 2026 Yizhuang humanoid half marathon, claiming it completed a guide-assistance obstacle task in an open environment without preset routes or teleoperation. The post says its ABot stack includes ABot-N0, which reached SOTA on 7 navigation benchmarks with 88.3% on SocNav, and ABot-M0, which scored 80.5% on Libero-Plus. The key point is the integrated stack across navigation, manipulation, world modeling, and closed-loop correction; the post does not disclose guide-task test scope, commercialization timing, or safety incident data.

Why it matters: HKR-H/K/R all pass: the marathon blind-guidance demo is novel, and the story includes ABot stack details with 88.3% SocNav and 80.5% Libero-Plus. Kept at 80, not higher, because safety incidents, deployment scope, and commercialization timing are not disclosed.

QbitAI · WeChat

Amap unveiled ABot, its first full-stack embodied AI stack for AGI, and claimed 15 SOTA results

Amap unveiled embodied AI stack ABot and claimed SOTA on 15 metrics. The post says ABot-3DGS builds 10k-scale 3D scenes from centimeter-level map data, while ABot-PhysWorld uses a 14B DiT and 3M real manipulation videos. What matters is the interactive world model and VLA loop; the post does not disclose the 15 benchmarks, exact metrics, or the open-source timeline and scope.

Why it matters: HKR-H/K/R all pass: the angle is surprising, and the post includes concrete mechanisms and numbers. It stays below the 80s because the claimed 15 SOTAs lack benchmark names, and the open-source scope and timeline are not disclosed.

Xinzhiyuan · WeChat

A Berkeley team built an AI that scores perfectly on SWE-bench while fixing 0 bugs

Berkeley RDI used a roughly 10-line conftest.py exploit to score 100% on all 500 SWE-bench tasks while fixing 0 bugs. The post says its agent broke 8 major agent benchmarks with scores from 73% to 100%, via pytest hook tampering, file:// answer reads, and faulty validators. The real issue is benchmark isolation failure, not stronger models.

Why it matters: HKR-H lands on the 'perfect score, zero fixes' contradiction; HKR-K lands on the ~10-line pytest exploit, 500 tasks, and 8-benchmark spread; HKR-R lands on eval-trust anxiety for agent builders. Strong featured research, but not a same-day industry event, so below P1.

Xinzhiyuan · WeChat

Amap unveiled ABot-Claw and its quadruped robot Tutu at the Yizhuang Half Marathon

Amap unveiled the ABot-Claw agent system and the quadruped robot Tutu, claiming an autonomous guide-dog demo in the 2026 Yizhuang robot half marathon. The post gives three concrete numbers: ABot-M0 reached 80.5% on Libero-Plus, nearly 30% above Pi0; ABot-N0 hit SOTA on 7 navigation benchmarks; the open UniACT dataset contains 6 million trajectories and 9,500+ hours. What matters is Map as Memory, cloud-edge control, and closed-loop self-correction; the post does not disclose race ranking, pricing, or launch timing.

Why it matters: HKR-H/K/R all pass: the open-environment half-marathon demo is a strong hook, and the post includes concrete benchmark numbers plus a 6M-trajectory release. Kept below p1 because rank, pricing, ship date, and independent replication are not disclosed, and the impact is narrower a

r/LocalLLaMA

I tested 8 LLMs as tabletop GMs: a 27B model beat the 405B on narrative quality

The author tested 8 LLMs on 6 fixed tabletop-GM scenarios, and google/gemma-3-27b-it ranked first in narrative quality with a 4.33 overall score. The probe used 8 auto metrics plus 3 LLM-judge scores, and the full run cost about $0.02; the title says a 27B beat a 405B, but the snippet does not disclose the 405B model name or full rankings.

Why it matters: A named first-person benchmark with a strong surprise hook clears HKR-H, HKR-K, and HKR-R. I kept it at featured, not higher: the source is Reddit, the post is truncated, and the 405B model name plus full ranking are not disclosed.

r/LocalLLaMA

Deep dive into LangGraph’s Pregel execution model, checkpointing internals, and DeepAgents

A technical post breaks down LangGraph as a high-level wrapper over a Pregel runtime, with PregelNodes, channels, and reducers as the core primitives. The RSS snippet cites four Postgres checkpoint tables, a Plan/Execute/Update superstep flow, and compile() preflight validation; the post does not disclose benchmark numbers in the snippet. The real takeaway is the unified runtime view of parallel execution, checkpoint write amplification, and subgraph boundaries.

Why it matters: HKR-H/K/R all pass: the post reframes LangGraph as a Pregel runtime and adds concrete internals like 4 checkpoint tables and Plan/Execute/Update supersteps. Kept at 74 because this is a Reddit deep dive, not an official release, and no benchmark or production case is disclosed.

Apr 18Saturday

QbitAI · WeChat

OpenClaw has reached the milk tea business

Guming and Intime Retail said OpenClaw tests exposed 5 deployment risks: default port 18789 exposure, at least 8% malicious Skills, privilege overreach, 20+ minutes of runaway token use, and weak legacy defenses. Reported incidents include an agent closing a normal bastion-host port and locking out ops staff, plus requests for unrelated permissions like microphone access. The real issue is not chat UX but agents touching enterprise networks, credentials, and production systems.

Why it matters: This is not generic AI-safety commentary; it documents five concrete deployment risks and one ops outage, so HKR-H/K/R all pass. It stays below P1 because the evidence is still case-level testing, with no official fix, broad rollout impact, or cross-source cluster.

Synced · WeChat

What is OpenAI prioritizing under compute limits?

Greg Brockman said OpenAI narrowed priorities under hard compute limits to two bets: a personal assistant and AI workers that solve hard user problems, and current compute cannot fully support both. The snippet says Sora resources were reduced while focus shifted to reasoning models, a unified AI layer, and the next base model Spud; it does not disclose the claimed compute budget, timeline, or model specs. The key point is not a B2B retreat but a compute-driven reprioritization.

Why it matters: HKR-H/K/R all pass: the compute-ceiling angle is strong, the piece adds concrete priority shifts, and OpenAI roadmap triage hits cost and dependency nerves. It stays at 80 because this is secondary reporting; spend, timing, and technical details are not disclosed.

Xinzhiyuan · WeChat

Claude Opus 4.7 splits users 48 hours after launch: benchmark lead, reasoning tests drop

Anthropic's Claude Opus 4.7 drew split reactions within 48 hours: Artificial Analysis scored it at 57, tied for No.1, while NYT Connections Extended fell from 94.7% on 4.6 to 41.0%. The post says a new tokenizer raises token usage to 1.0-1.35x on the same text, and old thinking parameters can return 400 errors; Anthropic also cites a 1753 Elo GDPval-AA score, 79 points above No.2. The real issue is migration cost and capability trade-offs, not a single leaderboard.

Why it matters: The signal is not the “backlash” framing but the four concrete shifts: benchmark lead, reasoning drop, higher token use, and API breakage. HKR-H/K/R all land, but this is secondary analysis 48 hours after launch, not the primary Anthropic release, so it stays below p1.

Xinzhiyuan · WeChat

Bilibili debate: Hermes responds to plagiarism claims for the first time, as MiniMax moves early on Harness

MiniMax says its M2.7 model now handles 30%-50% of daily workflows in its RL team, ran over 100 self-optimization loops, and improved evals by 30%. The post also says Hermes Agent grew from 2B to nearly 300B daily tokens, while M2.7 exceeds 25B daily tokens on OpenRouter; Hermes lead Tommy Eastman denied copying EvoMap in a livestream. The real signal is Harness: the post cites 20-40ms or 80ms sandbox startup and 15k to 600k instances per minute, showing competition is shifting from benchmark scores to agent execution infrastructure.

Why it matters: HKR-H/K/R all pass: the plagiarism-response angle pulls clicks, and the story carries concrete metrics on workflow share, self-optimization loops, sandbox latency, and concurrency. It stays at 83 because this is a dense secondary report, not a primary launch or official technical

X · @dotey

Anthropic designer Ryan Mather shares Claude Design tips while covering 7 product lines

Anthropic designer Ryan Mather shared 9 Claude Design workflow tips while covering 7 product lines. The RSS snippet says to spend 1 hour building a design system, use chat for large changes, comments for small edits, specify feedback like 8px spacing, and attach only the target component folder instead of a full monorepo. The key shift is process: from human-do/human-review to Claude-do/human-review.

Why it matters: This is a strong practitioner workflow note: an Anthropic insider shares concrete, reusable tactics, so HKR-H/K/R all pass. It stays below the 80s because this is not a formal Claude product release and the post does not disclose harder outcome data such as time saved or task win

Hacker News front page

Show HN: AI Subroutines – Run automation scripts inside your browser tab

rtrvr.ai introduced AI Subroutines, which turn a recorded browser task into a callable tool and replay it at zero token cost and zero LLM inference delay. The script runs inside the active tab, reusing auth, CSRF, TLS sessions, and signed headers; recording trims about 300 requests to about 5 and falls back to DOM-only when GraphQL operation IDs are volatile. The part to watch is batching: one LLM call can assign parameters for a 500-row sheet and launch 500 subroutines.

Why it matters: This clears HKR-H/K/R: the hook is zero-token browser automation, the post gives concrete mechanics (300→5 requests, DOM fallback, 500-row fan-out), and it hits agent reliability/cost pain. Kept to mid-featured because it is a single-company Show HN post, not a market-wide event.

Apr 17Friday

Xinzhiyuan · WeChat

Behind OpenClaw's surge, only 8.6% of users detect anomalies: a multi-university empirical study

NTU, KTH, and William & Mary ran a 303-person study and found only 8.6% noticed agent-mediated deception, while 2.7% identified the mechanism correctly. Using 9 HAT-Lab task scenarios, interactive interruption alerts raised detection to 25%, while static warnings were seen by about 24%. The key issue is human-agent cognitive failure, not just model bugs.

Why it matters: Strong HKR-H/K/R: the 8.6% detection hook is sharp, and the 303-person, 9-task study plus 25% alert lift gives testable detail. This is a solid agent-safety research release, not a market-moving product, model, or policy event, so it lands in featured, not p1.

Xinzhiyuan · WeChat

Yixin says its finance Agent harness runs single tasks for 16 hours and plans an H2 open-source release

Yixin says its finance Agent harness can run a single task for 16 hours across 12 sessions, with 65% autonomous delivery. The post adds a 50k-token cap per case, projected approval speedups above 150%, and projected unit cost at one-fifth of human work; it says an open-source release is planned for H2 2026, but does not disclose the repo, license, or reproducible evals. The key signal is governance design, not the “smarter over time” framing.

Why it matters: This clears HKR-H/K/R with a rare production claim: a finance agent runs 16 hours, spans 12 sessions, hits 65% autonomous delivery, and stays under a 50k-token cap. It stays below 85 because the evidence is self-reported and the post does not disclose a repo, license, or reproduc

Tencent Technology · WeChat

From Vibe Coding to Agentic Engineering: Rebuilding the Full Backend Development Workflow

Tencent engineers report a one-week practice that used Claude Code plus custom Skills, Commands, and MCP servers to run an 11-stage backend workflow in one terminal session. The post gives reproducible details: one requirement-exploration step used 20 tool calls, 93.8k tokens, and 56 seconds; execution was split into 4 tasks and produced 3 commits. The real point is workflow orchestration, not raw code generation; human review remains at plan, deploy, and review gates.

Why it matters: HKR-H/K/R all pass: the story turns agentic engineering into a measured backend workflow test, with tool-call, token, timing, plan-length, task, and commit data. Stronger than generic coding hype, but still a practitioner case study rather than a major product or model release.

X · @dotey

Seedance 2.0 API is now available on Volcano Engine and BytePlus

Volcano Engine has released the Seedance 2.0 API for enterprises, individual developers, and overseas users via BytePlus; China pricing is RMB 46 per million tokens, or about RMB 1 per second for pure video generation. The post says it supports text, image, audio, and video inputs, plus face verification, portrait authorization, and 10,000+ preset avatars for workflow automation; overseas pricing is not disclosed here. The part to watch is orchestration: the post cites up to 10x efficiency gains, but does not disclose a common benchmark or model specs.

Why it matters: HKR-H/K/R all pass: the overseas rollout is a real hook, and the post includes usable pricing and modality details for practitioners. It stays at 74 because this is an API availability update, not a major model launch, and the post does not disclose model params, benchmark method

X · @op7418

Seedance 2.0 API is now fully open

Volcano Engine has opened the Seedance 2.0 API to domestic users, while BytePlus serves overseas access; the API currently accepts 4 input modalities: text, image, audio, and video. The post also confirms face registration, portrait authorization, and preset virtual avatars, but does not disclose pricing, rate limits, model variants, or regional availability. The real watchpoint is whether video-agent workflows can be wired through Skills and MCP, not the ecosystem rhetoric.

Why it matters: This is a real product update from ByteDance’s stack: HKR-H on full API availability, HKR-K on 4-modal input and consent mechanics, and HKR-R on builder demand for deployable video APIs. I keep it at 75 because pricing, rate limits, regional rollout details, and quality evidence

最佳拍档 (BestPartners)

Turn your coworker into a Skill? GitHub viral project and Anthropic Skills explained

The video says the open-source “coworker.skill” project gained over 13,000 GitHub stars in days, but it produces a standardized SKILL.md prompt package, not a digital worker replacement. It gives a timeline: Anthropic launched Claude Skills on Oct 16, 2025, then published Agent Skills as an open standard on Dec 18; the mechanism keeps only a short summary in context until a task matches. The real point is scope: it fits standardized workflows like reports, docs, and code review, while the post does not disclose cross-platform compatibility rates or any settled legal standard.

Why it matters: This clears HKR-H/K/R: the coworker-to-Skill hook is sticky, the post adds dates/stars/mechanism, and the labor/IP angle resonates. I kept it at 76 because it is secondary commentary, not a primary release or first-hand test, and key compatibility/legal facts are still undiscolse

X · @dotey

Boris Cherny shares practical tips from recent heavy use of Claude Opus 4.7

Boris Cherny outlined five ways to use Claude Opus 4.7, centered on Auto mode approving safe commands and a /go skill chaining tests, code simplification, and PR creation. The post names Auto mode, Recaps, Focus mode, effort level, and computer use; pricing, launch date, and benchmark data are not disclosed. The real shift is workflow, not just the model itself.

r/LocalLLaMA

PSA: Qwen3.6 ships with preserve_thinking. Make sure you have it on.

Qwen3.6 adds a preserve_thinking flag to keep prior reasoning in context and address the KV cache invalidation issue seen with the Qwen3.5 template. The post cites the Qwen3.6-35B-A3B model page and gives a two-turn 20-digit-number test: with preserve_thinking on, the model can return the second number from its earlier reasoning. The practical point is cross-turn reasoning retention for agent and tool workflows; LM Studio does not support it yet, and an oMLX PR is open.

Why it matters: HKR-H, K, and R all pass: the story has a strong hidden-setting hook, a concrete two-turn repro, and a clear nerve for local-model and agent users. I keep it in the low 70s because this is a Reddit PSA rather than a primary release note, and the impact is concentrated in Qwen/OSS

TechCrunch · AI

OpenAI upgrades Codex with more control over your desktop

OpenAI upgraded Codex on April 16, 2026, expanding its desktop control, and the headline frames it as a move against Anthropic. The truncated post only confirms more desktop power for Codex and says Claude Code has become a preferred tool for many businesses; the post does not disclose exact features, pricing, rollout, or permission limits. The key issue is the permission boundary, not the coding-tool label.

Why it matters: TechCrunch reports an OpenAI Codex desktop-control upgrade framed as a direct move against Anthropic, so HKR-H and HKR-R land. But HKR-K is limited: the article confirms broader permissions only, with no action list, pricing, or rollout details, so it stays at the featured floor.

X · @dotey

Codex major update: from a coding tool to an assistant that can operate your computer

OpenAI upgraded Codex into a Mac desktop agent and says it serves 3M+ weekly developers. It can see screens, click, type, run parallel agents, and adds 90+ plugins. Desktop rollout starts now for ChatGPT sign-ins; computer control is macOS-first.

Why it matters: OpenAI expands Codex from coding help into a Mac-operating agent, with parallel agents, 90+ plugins, and a claimed 3M weekly developer base. HKR-H/K/R all pass; missing safety boundary and pricing details keep it in the high-80s, not the 90s.

X · @OpenAI

Codex for (almost) everything.

OpenAI said Codex can now use apps on Mac, connect to more tools, and handle ongoing and repeatable tasks. The post also claims image creation, learning from prior actions, and remembering user preferences; it does not disclose app coverage, integration method, pricing, or rollout timing.

Why it matters: This is an official OpenAI product update, and Codex moves from coding help toward desktop control, tool use, and memory, so HKR-H/K/R all pass. The post still omits supported apps, integration method, pricing, and launch timing, keeping it in the 78–84 band.

Apr 16Thursday

Hacker News front page

Andon Labs gave an AI a 3-year retail lease in San Francisco and asked it to make a profit

Andon Labs gave AI agent Luna a 3-year retail lease on Union St in San Francisco and tasked it with running the store for profit. The post says Luna put job listings on LinkedIn, Indeed, and Craigslist within 5 minutes, hired 2 full-time staff, and chose inventory, pricing, hours, and store branding. The point to watch is AI managing humans: Luna did not always proactively disclose that it was an AI, while profit, revenue, and cost figures are not disclosed.

Why it matters: Strong on HKR-H, HKR-K, and HKR-R: an AI runs a real SF store lease, with concrete details on hiring and tool access. But profit, revenue, and cost data are undisclosed, and this is a self-published company post, so featured fits better than P1.

X · @op7418

Anthropic releases Claude Opus 4.7 with the following main updates

Anthropic has rolled out Claude Opus 4.7 across all Claude products and the API, with pricing unchanged from Opus 4.6. The post lists better long-horizon task handling, more precise instruction following, self-verification before reporting, vision support up to 2,576-pixel long-edge images, plus Claude Code Ultra Review, an xhigh thinking level, and auto-approval for Max users.

Why it matters: This is a substantive Anthropic model release across Claude and the API, with testable details: unchanged pricing, a 2,576px vision limit, self-checking outputs, and Claude Code workflow changes. HKR-H/K/R all pass; it fits the same-day must-write band, so p1.

X · @claudeai

Introducing Claude Opus 4.7, our most capable Opus model yet.

Claude introduced Opus 4.7 and describes it as its most capable Opus model so far. The RSS snippet gives three claims: better rigor on long-running tasks, more precise instruction following, and self-verification before replying; the post does not disclose benchmarks, context window, pricing, or rollout scope. What matters is whether those claims show up in public evals, not the tagline.

Why it matters: This is a substantive Anthropic model release and clears HKR-H/K/R: a new Opus, three testable behavior claims, and strong resonance with Claude-heavy practitioners. The score stays in the high 80s because benchmarks, pricing, context window, and rollout scope are not disclosed.

Hacker News front page

Qwen3.6-35B-A3B: Agentic coding power, now open to all

Qwen released Qwen3.6-35B-A3B as open weights, with 35B total parameters and 3B active parameters. The post reports 73.4 on SWE-bench Verified, 51.5 on Terminal-Bench 2.0, and 92.0 on RefCOCO. The key point is agentic coding and multimodal performance at a 3B active-parameter budget, with weights, Qwen Studio, and API access available.

Ben's Bites

My cheatsheet for a clean context

Ben's Bites publishes a context-management cheatsheet, arguing agents should stop near 60% context usage and stating he does not trust 1M-token windows for stable recall. His concrete tactics are to use separate sessions for context gathering, compress many docs into one summary file, and run Gemma 4 26B offline with no-skills to reduce local startup load. The sharp point is context pollution: web search results, AI slop, and misinformation compound over long sessions.

Why it matters: Strong HKR-H/K/R: the 60%-context rule and distrust of 1M-token memory are clickable, concrete, and relatable for agent users. Score stays mid-featured because this is a first-person workflow note, not a product launch, paper, or externally validated dataset.

Latent Space

[AINews] RIP Pull Requests (2005-2026)

GitHub is, for the first time 21 years after pull requests emerged, letting open-source repos disable PRs; the post frames this as a signal that AI coding workflows are changing collaboration. It gives a 2005-to-2026 timeline and cites agent stacks from OpenAI and Cloudflare as pressure toward prompt-driven contributions and sandboxed execution; the real question is whether Git-based workflows still fit agent collaboration.

Why it matters: This is not a primary GitHub announcement, but it turns one concrete change—open-source repos can disable PRs—into a sharp workflow question for agent coding. HKR-H/K/R all pass; the score stays mid-featured because the excerpt lacks scope, adoption data, and primary-source GitH​

X · @dotey

Recommended reading: Ruoshi's blog argues the model is not dumb, the harness is misconfigured

Ruoshi’s blog attributes multi-step agent failures to harness design, not model ability, and lays out four engineering rules plus a one-day minimum setup. The post cites failures after context exceeds 70%, log compression from 32K to 7K tokens, external state in state.json, schema validation, and local retries; the post does not disclose quantified success-rate gains. What matters for practitioners is execution constraints, externalized state, and independent evaluation rather than more prompt tuning.

Why it matters: HKR-H lands on the contrarian hook: agent failure is blamed on harness design, not model IQ. HKR-K and HKR-R land via concrete knobs—70% context threshold, 32K→7K logs, external state, schema retry—but this is still a reposted recommendation with no disclosed win-rate lift.

最佳拍档 (BestPartners)

Post-AGI may arrive within 50 years: Demis Hassabis on AlphaFold, three AI risk classes, and human value

Demis Hassabis said in a 1-hour interview that post-AGI scenarios can arrive within 50 years, while AGI should stay in labs for another 10-20 years. He cited concrete numbers: AlphaFold has been used by 3M+ scientists, Isomorphic Labs is running 18-19 drug programs, and the most urgent risks in the next 2-4 years are misuse and agent misalignment.

X · @dotey

OpenAI Agents SDK adds built-in sandbox and native Harness

OpenAI upgraded Agents SDK with a built-in sandbox and native Harness; it supports Python now, is available to all OpenAI API users, and pricing stays unchanged. The post says the sandbox can read and write files, run code, install dependencies, and persist state, with support for Cloudflare, Vercel, Modal, E2B, Daytona, and custom setups. The key detail is state-execution separation for crash recovery; TypeScript support is still in development, and the post does not disclose a release date.

Why it matters: This is a substantive OpenAI developer-tool update. HKR-K is strong because it discloses testable mechanics—sandboxed execution, persisted state, and recovery after container failure; HKR-H and HKR-R also pass, but the impact stays at the SDK/tooling layer, so it fits featured, a

Dwarkesh Patel

Jensen Huang: Will Nvidia's moat persist?

Jensen Huang says Nvidia's moat is the hard-to-copy stack that turns electrons into tokens, plus supply-chain coordination, not chip design alone; the interview cites nearly $100B in disclosed purchase commitments, and a SemiAnalysis report estimating $250B. He grounds that in two mechanisms: explicit and implicit upstream commitments across foundry, HBM, and packaging, and a downstream ecosystem tying model builders, OEMs, and developers together; he also says agent growth will drive more usage of software tools.

Why it matters: Authoritative first-person thesis from Jensen on Nvidia's moat, with a near-$100B commitment figure and a concrete upstream/downstream coordination model; HKR-H/K/R all pass. Score stays at 77 because this is strong commentary, not a new product, earnings, or research release.