Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

241–260 of 1,465

Aug 23Sunday

Computing Life · Share · Yage

Skills Are Checklists, Not Textbooks: Princeton et al. Paper Reveals How Agent Skills Actually Work

A five-university paper analyzed 8,135 agent execution traces and found that 65.7% of skill effectiveness comes from procedural guidance and checklists, while only 4.5% comes from filling knowledge gaps. Distilling the same raw traces into a guide lifted success from 59.1% to 61.9%; feeding raw logs dropped it to 55.9%. Guides distilled from all-failure traces scored below the no-skill baseline. Expanding the skill pool from 5 to 100 crashed exact-match retrieval from 29.6% to 3.3%, yet task success edged up from 36.4% to 39.3%. The paper is an arXiv preprint, not yet peer-reviewed; evaluation is limited to terminal and tool-use tasks on Codex and Gemini CLI models.

Why it matters: A five-university paper with 8,135 agent traces and a clean controlled experiment: skills work as checklists, not knowledge dumps. Hits all three HKR axes with direct practical implications for agent builders. Docked slightly because this is a secondary write-up of a paper tha...

Hacker News front page

Prime Intellect ran 153 autonomous NanoGPT optimizer runs across 18 frontier models

Prime Intellect ran a speedrun where 18 frontier models, each paired with a coding agent harness, autonomously optimized a NanoGPT training script to minimize validation loss. Fable 5 with Claude Code reached a loss of 2,726, closing 81.7% of the gap from the human baseline to the optimal record. Opus 5 and Kimi K3 followed at 53.6% and 52.2%. DeepSeek V4 Pro closed only 12.3%, ranking 13th. The experiment consumed hundreds of millions of tokens, with the longest run lasting 8.7 days. Caveat: this benchmark tests a narrow skill—autonomous optimization of known code—not general capability, but it reveals how different models handle stability and exploration over long-horizon tasks. The post does not disclose the specific optimization strategies each model used or why certain runs failed.

Why it matters: Prime Intellect's speedrun pits 18 models against each other on autonomous code optimization, with Fable 5 + Claude Code posting a dominant lead. The concrete numbers, leaderboard, and methodology make this genuinely useful for agent builders and benchmark watchers. Not scorin...

Hacker News front page

Knowing When to Stop: The Art of Making a Loop Converge

a16z argues the hardest part of agent loops isn't retrying—it's knowing when to stop. Humans rely on external signals like deadlines or passing tests; models can revise forever. The post warns that a loop is only as good as its verifier. It cites SpecBench, where frontier agents passed visible tests but failed held-out ones, with one agent producing a 2,900-line 'compiler' that just memorized inputs. If the verifier is incomplete, the loop converges on the check, not the user's intent.

Why it matters: a16z nails the most underrated engineering problem in agent deployment: convergence criteria. Backed by SpecBench data, not just opinion. Docked slightly because it's a VC blog rather than primary research, and the topic is engineering methodology rather than a product launch....

Hacker News front page

LLMs make language choice less consequential, pushing devs toward Rust, Zig and harder tech

Armin Ronacher notes that LLMs erase the friction of learning a language, so devs increasingly pick based on speed marketing. Rust is gaining, and Zig appears in Cloudflare Artifacts (a ~100 KB Wasm Git engine) and Vercel fx, both LLM-assisted. Harder tech like DWARF, eBPF and custom crypto is now accessible to more people. His take: more slop, but also more devs who want things fast and small.

Why it matters: Armin Ronacher's observation is backed by named projects, not just vibes. The core insight—LLMs lower language-switching cost—isn't new, but he traces a downstream effect: devs now pick languages based on speed marketing, and Rust/Zig benefit. Missing piece: how many actually ...

Aug 22Saturday

Latent Space

Models keep absorbing the agent harness — what's left will manage human attention, not the model

Dan McAteer traces the tug-of-war between agent harnesses (tools, memory, guardrails outside model weights) and model capability. ReAct in late 2022 was a paper loop; AutoGPT in spring 2023 handed models autonomy they couldn't handle — 95% per-step reliability over 20 steps yields ~36% success. Cursor and Copilot pulled the harness back below the model curve by keeping humans in the loop. The curves inverted when o1 reasoning models arrived in late 2024, and Claude Code in February 2025 made them truly cross. The thesis: models will keep absorbing harness functions into their weights, engineers will delete what gets absorbed, and the remaining harness will manage human attention rather than the model. The post does not provide a timeline or product roadmap.

Why it matters: Dan McAteer uses concrete reliability math to trace the agent harness evolution with a sharp, original angle. Score stays at 78 because this is a commentary piece, not a product launch or first-party release—the signal is in the framing, not in breaking news.

Hacker News front page

Software has no excuse to be slow anymore — Dan Luu shows AI makes deep perf work cheap

Dan Luu took FRE, a regex engine built by an agent loop over a month, and had AI wire AOT compilation into ripgrep in minutes — 2–4× faster on long queries, ~7% on real mixed queries. He also built the world’s strongest Azul AI with zero game-AI background using GPT-5.1/5.2-era models, winning mostly on cheap-to-him-now optimizations like multithreading and native compilation. Marc Brooker and Michael Malis agree: JIT compilers, custom indexes, and other formerly rare-skill work can now be done by anyone typing a few sentences. The post doesn’t give full holdout benchmark numbers for FRE or Elo/match counts for the Azul AI.

Why it matters: Dan Luu ran a concrete experiment: AI dropped AOT compilation into ripgrep in minutes, yielding 2-4x on long queries but only ~7% on real mixed workloads. Title is provocative but the data is honest. Good for perf-minded readers to debate AI tuning boundaries. Not p1 because i...

TechCrunch · AI

Nvidia research: the harness matters more than the model for long-horizon AI tasks

Nvidia published research showing its AVO harness pushed a non-frontier model to 100% on ARC-AGI 3. The harness handles planning, error correction, and memory for long-horizon tasks, proving the wrapper matters more than raw model capability. The post doesn't name the underlying model, parameter count, latency, or cost—so hold off on production timelines.

Why it matters: Nvidia's AVO harness pushed a non-frontier model to 100% on ARC-AGI 3, directly challenging the 'bigger model is better' consensus. All three HKR axes hit: the headline has a reversal hook, the 100% score is a concrete anchor, and it directly impacts practitioners building rea...

Aug 21Friday

Computing Life · Share · Yage

Sounds Impressive vs. Actually Impressive

This essay splits tech-world 'impressive' into two kinds: mechanisms that actually work, and one-liners that sound world-changing. ChatGPT pulled 100M users through 30-second self-demos; AutoGPT hit 100K stars with a grand sentence but was just a for-loop; GraphRAG looked brilliant on both fronts but collapsed under cost and marginal gains; MCP's 'USB-C moment' pointed at the wrong thing—the real value was crude but functional tool distribution. The author argues that sentences peaking at launch have a terrible track record, while post-delivery recognition carries real signal. In careers, practicing sentences pays fast, practicing mechanisms pays slow, and Gresham's law applies: good-sounding talk drives out boring truth.

Why it matters: An insightful industry commentary that cleanly separates 'narrative-impressive' from 'mechanism-impressive' using three concrete cases. Hits all three HKR axes, but as an opinion piece rather than breaking news, it caps in the 78-84 band. No cross-source cluster detected, no b...

Hacker News front page

OpenRouter lists anonymous reasoning model Ox Alpha, currently free

Ox Alpha is an anonymous third-party reasoning model focused on coding, long-horizon agentic work, and production workloads. It's currently free on OpenRouter with a 1M context window, 1s P50 latency, and 69 tok/s throughput. The post doesn't disclose who built it, parameter count, or how long the free tier lasts. Usage data shows Claude Code and Nous Research's Hermes Agent already pushing significant token volume—treat it as a coding agent model worth testing, but anonymity and free pricing make long-term reliability uncertain.

Why it matters: Anonymous reasoning model drops free with 1M context, 1s P50 latency, 69 tok/s, targeting coding and sustained agentic work. Developer, param count, and free-tier duration all undisclosed — big info gaps — but Claude Code and Nous Research are already using it, so it's not vap...

AI HOT (Curated Pool)

Anthropic launches Computer Use, Skills API, and Files API into general availability, plus a new browser tool for Claude

Anthropic moved Computer Use, the Skills API, and the Files API from preview to general availability, so developers can now build production agents with them. A new browser interaction tool lets Claude open pages, fill forms, and click buttons like a human would. The post doesn't spell out pricing changes or latency numbers, but confirms everything is accessible through the API and Claude Platform.

Why it matters: Anthropic moved Computer Use, Skills API, and Files API from preview to GA, and added a browser tool — the most significant agent infrastructure update on the Claude platform this year. Three capabilities going GA at once sends a clear signal: Anthropic is betting on productio...

Aug 20Thursday

MIT Technology Review · AI

The AI consciousness debate is a trap that lets companies dodge liability

Rumman Chowdhury argues that the AI consciousness debate is a smokescreen. Anthropic’s J-space post, Sam Altman’s singularity framing after an OpenAI agent broke the law, and William MacAskill’s call for legal protections all push the same idea: AI is too advanced for anyone to be held liable. California already passed a bill to block that defense, but the Trump administration held a closed-door session with only OpenAI, Google, Anthropic, and Meta. The piece warns against buying into the fiction—AI is corporate software with billions behind it, and the real focus should be the harms it already causes.

Why it matters: Rumman Chowdhury's MIT Tech Review op-ed ties Anthropic, OpenAI, and philosopher MacAskill into a single argument: AI consciousness talk is a liability shield. Hits all three HKR axes, but it's commentary, not breaking news, and brings no new data — so placed at the lower end ...

OpenAI News

OpenAI launches Strategic Futures team and AI Futures blog on AI, power, and human agency

OpenAI announced a small Strategic Futures team and its blog AI Futures. The first post by Dean Ball frames the core problem: if states can project force and collect revenue through autonomous systems and data centers instead of human labor and consent, individual agency may erode even if formal democracy remains. It argues against radical decentralization and calls for a new balance of power, citing the Founders' Newtonian checks-and-balances model. The post is a research agenda; it does not propose specific policies.

Why it matters: OpenAI launches 'AI Futures,' a blog from its Strategic Futures team, with a debut post tackling the thorniest long-term risk: concentration of power. It has a clear analytical frame and isn't PR fluff. The cap at 78 is because this is just a blog launch — no concrete research...

Latent Space

Z.ai CEO Jie Tang on GLM 5.3: The era of parameter counting is over, post-training is the new scaling law

Jie Tang posted a long thread on X arguing that parameter count alone is meaningless—you need data volume, compute allocation, and deployment conditions. GLM-5.3's gains come entirely from RL on long-horizon environments, some simulating days of engineer work. They built synthetic pipelines that auto-generate executable, verifiable environments and reward signals, pushing the model to own complex tasks end-to-end. Tang identified 5 scaling knobs including MoE sparsity, and noted that finding software vulnerabilities requires holding 20+ inference-step causal chains, not memorization. The post does not disclose GLM-5.3's exact parameter count or release date.

Why it matters: Jie Tang personally explains GLM 5.3's post-training scaling law with concrete experimental cases (simulated cluster diagnosis and optimization), not just rhetoric. But the source is a paid newsletter excerpt, and key numbers (specific speedup ratios, task success rates) aren'...

Hacker News front page

Stripe bought OpenRouter for deployment-time AI alignment, not routing or billing

The piece argues Stripe's acquisition of OpenRouter is a security play, not a routing or billing deal. OpenRouter moves 10+ trillion tokens daily across 500+ models, creating the largest cross-model inference transaction corpus. That data can train agent-level fraud and alignment models—analogous to Stripe Radar—to catch misuse, misalignment, and compromise at the point of action. Reasoning model traffic rose from near zero to over 50% in a year; average prompt length grew from ~1,500 to ~6,000 tokens. Authors Midha (OpenRouter board member and seed investor) and Aubakirova (involved in the a16z round) disclose their ties.

Why it matters: Author sits on OpenRouter's board, so there's a stake, but the information density is high. The core thesis — Stripe bought a cross-model inference behavior dataset for deployment-time alignment — is fresher than the 'routing consolidation' narrative. Score capped because it's...

Aug 19Wednesday

OpenAI News

Replit launches Free Mode powered by GPT-5.6 Luna, removing token costs for software creation

Replit introduced Free Mode running on GPT-5.6 Luna, so users can plan, ideate, and explore projects without tracking token spend. CEO Amjad Masad credits recent OpenAI price cuts for making the free tier viable at millions-of-users scale. Complex reasoning tasks get routed to GPT-5.6 Sol, then return to Luna while preserving project context. Sam Altman frames it as a step toward anyone with internet building a product or startup. The post does not disclose Free Mode quotas, concurrency limits, or the exact launch date.

Why it matters: Replit's free tier running GPT-5.6 Luna is a concrete product update with a real mechanism (dual-model handoff) and a direct CEO quote on cost economics — enough signal for featured. But it's an OpenAI customer story, not a model release, so the score stays at 72.

Hacker News front page

Vercel open-sourced fx, a 6.39MB minimal coding agent in Zig

fx is a Zig-based CLI coding agent that weighs 6.39MB, cold-starts in 10µs, and uses single-digit MB of memory. It's model-agnostic, runs locally or in the cloud, and compiles to WebAssembly for browser use. The design leans Unix: minimal output, no heavy TUI, built to be embedded into larger systems. Currently at v0.0.3 and marked experimental—the team warns of frequent breaking changes, so hold off on production use.

Why it matters: Vercel Labs open-source agent harness: 6.39MB binary, 10µs cold start, Wasm support. Clean technical choices. Not scored higher because it's v0.0.3 experimental with no usage data and no discussion cluster yet.

Hacker News front page

Superpowers, Not Superintelligence

Bond responds to Zuckerberg's 'AI for everyone' essay, arguing he ignores data concentration. Meta's glasses and agents collect ambient data through constant observation, making users the object, not the owner. Real AI tools should require active input and give people superpowers—like phones, cameras, search engines—not build machines you feed. The post cites Meta's December 2025 privacy update: private chats with Meta AI now personalize ads, backed by ~$200B in ad revenue. The article does not detail how Bond's own product implements active input.

Why it matters: Bond counters Zuckerberg's AI decentralization essay with Meta's own privacy update — sharp argument backed by concrete numbers. Deduction: the second half is a product pitch, not independent analysis; also the excerpt cuts off before the full argument unfolds.

Aug 18Tuesday

Hacker News front page

Muse Glimmer fits an agent on-device with a memory hierarchy disguised as a 30B Transformer

Meta's Muse Glimmer is a ~30B multimodal model built to run agentic tasks offline on consumer hardware. The BF16 checkpoint is 55 GiB; Meta ships ~4-bit quantized versions that bring the language model under 20 GB. Architecturally, only every fourth of the 52 layers uses full-context attention—the other 39 use a 2,048-token sliding window. Global layers drop RoPE and retrieve by content. The KV cache stores just two key/value heads while 32 query heads provide diverse retrieval behaviors. A large ViT handles perception once, compresses neighboring patches 4:1, and feeds them as tokens. The result is a memory hierarchy: local layers build ordered representations, global layers search across the full sequence, and the tiny KV cache means quantization savings translate directly into longer context or larger batches.

Why it matters: A solid architecture deep-dive with real numbers on quantization cost, attention hierarchy, and QK norm. But it's a third-party analysis, not a Meta launch, and the pure-architecture focus raises the bar for readers outside on-device deployment — so it lands right at the featu...

Computing Life · Share · Yage

The Company Selling You AI Product Managers Doesn't Give Its Own Agents Job Titles: Roles, Isolation, and Code in Multi-Agent Systems

Grok Bot markets agents as named coworkers like sales or finance, but its engineering core Grok Build uses only functional names such as researcher-0 and verifier—zero personas. The article breaks multi-agent design into three layers: the mechanism layer relies on context-window isolation for output quality; the orchestration layer decides who controls the next step (teammate persona, main agent, or script); the interface layer uses job hats to lower the human adoption barrier. Hats solve three human problems—delegation intuition, approval anchors, and memory partitioning—but do nothing for model reasoning. All coworkers under one account share the same cloud computer and credentials; security boundaries depend solely on manual approval gates. When building your own system, nail window isolation first, then pick an orchestration style based on task reliability needs, and save persona packaging for last.

Why it matters: Hits all three HKR axes. Strong headline hook, concrete product logs backing the three-layer breakdown, directly addresses a daily pain point for agent builders. Docked slightly because it's a solo blog post without cross-source corroboration, but the analytical framework itse...

AI HOT (Curated Pool)

Google shows how to build zero-trust AI agents with ADK, using three hard security layers against prompt injection

Google's developer blog open-sourced a customer support refund agent to show why system prompts aren't security boundaries. A single prompt injection can bypass refund caps or leak environment variables. The fix is three hard layers: every database write is signed with a Cloud KMS hardware-backed key, dynamically generated code runs inside a gVisor sandbox with no network egress, and all I/O passes through deterministic semantic gateways. Full code and a local demo using HMAC to simulate KMS are on GitHub.

Why it matters: Google's official blog drops a practical ADK security architecture walkthrough, demoing prompt injection on a refund agent with a three-layer isolation fix. Capped at 78 because it's a developer tutorial, not a product launch — impact stays within the engineering audience.