Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1121–1140 of 1,465

May 5Tuesday

r/LocalLLaMA

Benching Local Qwen as a Codex Validator, Co-agent, and Challenger

robert896r1 tested Qwen3.6 27B GGUF beside Codex as a coding validator and released a reproducible eval suite. The runs covered Bartowski, Unsloth, 65k/128k context, and q8/f16 KV cache; three 128k profiles tied for best, with no measured q8 KV accuracy loss in this suite. The useful signal is the sidecar eval: missed directives, overbuilding, UI judgment, and long-context misses, not a universal leaderboard.

Why it matters: HKR-H/K/R all pass: a reproducible sidecar eval with concrete Qwen/Codex conditions beats a normal Reddit tip. Source authority and event scale keep it in the 72–77 band, not a same-day must-write.

May 4Monday

Hacker News front page

Sierra Raises $950M at $15B Valuation

Sierra raised $950M at a $15B valuation. The RSS snippet does not disclose investors, round type, use of funds, or product metrics. The signal is customer-agent valuation, not a model update.

Why it matters: HKR-H/K/R all pass: the $950M and $15B figures make this a strong agent-market story. Limited sourcing on investors, round, product metrics, and use of funds keeps it in the 78–84 band.

Financial Times · Technology

Blackstone and Goldman among backers for $1.5bn JV with Anthropic

Blackstone and Goldman are among backers of a $1.5bn joint venture with Anthropic. The consulting firm will advise Wall Street firms on AI deployment across portfolios; the post does not disclose ownership, products, or timeline.

Why it matters: HKR-H/K/R all pass: a $1.5bn Anthropic-linked JV backed by Blackstone and Goldman is a strong commercialization signal. Missing equity structure, product details, and timeline keep it below 85.

Import AI (Jack Clark)

Import AI 455: Automating AI Research

Jack Clark argues that no-human-involved AI R&D has a 60%+ chance of arriving by the end of 2028, citing SWE-Bench gains from Claude 2 at about 2% to Claude Mythos Preview at 93.9%, plus METR task horizons rising from 30 seconds in 2022 to 12 hours in 2026.

Why it matters: HKR-H/K/R all pass: Jack Clark anchors a >60% end-2028 automated-AI-R&D claim in SWE-Bench and METR numbers. This fits the 85–94 band for a notable figure’s AI-timeline essay, below model-release magnitude.

r/LocalLLaMA

Deep research report with Hermes Agent and qwen3.6-35b-a3b Q6_K

A Reddit user used Hermes Agent and qwen3.6-35b-a3b Q6_K to produce a 21-page research report. The run took 6 loops and over 5 hours on an RTX 4060, at about 28 tokens/s. The repo includes prompts, scripts, intermediate artifacts, and the final report.

Why it matters: HKR-H/K/R all pass: this is a local-agent experiment with hardware, runtime, speed, and artifacts. Reddit source limits reach, so it stays in the 72–77 featured-threshold band.

QbitAI · WeChat

DeepSeek-TUI, a “DeepSeek Claude Code,” reaches 2.3k GitHub stars

DeepSeek-TUI reached 2.3k GitHub stars; the Rust project is MIT-licensed. It targets DeepSeek V4 with a 1M-token context, RLM up to 16 V4 Flash subtasks, MCP, Shell, Git, and three control modes. Watch cache misses: uncached tokens cost 10x cached tokens.

Why it matters: HKR-H/K/R all pass: the hook is a DeepSeek-flavored Claude Code, with 2.3k stars, 1M tokens, 16 subtasks, and a 10x cache-miss cost gap. Impact is developer-specific, so it sits in the 72–77 band.

Xinzhiyuan · WeChat

Claude token rankings: Disney employee hits 460,000 calls in 9 days; Meta burns 60T monthly

Xinzhiyuan says Disney tracks Claude use via an AI Adoption Dashboard, with one employee making about 460,000 calls in 9 workdays. It also says Meta used 60 trillion tokens in 30 days, worth about $9B by public API pricing; the post does not show raw tables. The key issue is that input rankings are not outcomes.

Why it matters: HKR-H/K/R all pass: the hook is concrete usage shock, the post gives dashboard mechanics and token figures, and the nerve is enterprise Claude cost control. Kept at 74 because the data is secondhand and no raw table is disclosed.

r/LocalLLaMA

Pushing a 5-Year-Old 6GB VRAM Laptop to Its Limits: Qwen3.6-35B-A3B

Reddit user abhinand05 ran Qwen3.6-35B-A3B on a 5-year-old Asus ROG Zephyrus G14, reaching about 23 t/s plugged in and 10+ t/s unplugged. The setup uses RTX 2060 Max-Q 6GB, 24GB DDR4, Ryzen 7, plus llama-server configs for 64k and 128k context. The key detail is the mix of CPU MoE, KV-cache quantization, and ngram speculative decoding.

Why it matters: HKR-H/K/R all pass: the old-laptop angle is clicky, the post gives speeds and configs, and local-LLM cost resonates. It remains a single Reddit run, not a broader release.

May 3Sunday

r/LocalLLaMA

Local LLM Benchmark for Backend Generation via Function Calling: GLM vs Qwen vs DeepSeek

AutoBe posted a controlled backend-generation benchmark and says qwen3.5-35b-a3b matches gpt-5.4 on DB/API design. One shopping-mall run uses 200–300M tokens, costing $1,000–$1,500 per model at GPT 5.5 pricing. The key caveat is n=4 projects and self-scoring harness bias.

Why it matters: HKR-H/K/R all pass, but Reddit sourcing, n=4 projects, and self-eval harness bias keep it at the low featured band. Concrete cost and test constraints carry the score.

r/LocalLLaMA

LLM proxy that lets Claude Code talk to any model

DataNebula released open-source rosetta-llm, letting Claude Code call multiple providers through one gateway. It translates Anthropic Messages, OpenAI Chat, and OpenAI Responses, and round-trips encrypted reasoning via the signature field. The key detail is thinking-block fidelity for multi-turn agent prompt-cache hits.

Why it matters: HKR-H/K/R all pass, but this is a Reddit open-source tool post with no adoption, stars, or benchmark data disclosed. Score stays in the mid-weight tooling band, not 78+.

r/LocalLLaMA

Upskill: skill registry your agent consults before it starts, with 10k+ indexed skills

Autoloops released Upskill, an open-source skill registry with 10k+ indexed skills for agents. Search combines Postgres full-text search, 1024-dim embeddings, and reranking by stars, installs, and feedback. LLM adversarial review blocked hundreds of skills at index time.

Why it matters: HKR-H/K/R pass: a useful open-source agent registry with concrete retrieval and safety mechanics. Source authority is low and adoption is unproven, so it stays in the 72–77 featured band.

Synced · WeChat

Why CTOs at Billion-Dollar Companies Are Joining Anthropic as Engineers

Jiqizhixin lists at least six CTOs who joined Anthropic as individual contributors. Cases include Workday, You.com, Box, Super.com, and Adept AI from Jan 2025 to Apr 2026. The key issue is career leverage, not just AGI mission talk.

Why it matters: HKR-H/K/R all pass: the career-status reversal is clickable, the post gives 6 cases, and it touches AI talent competition. No hard exclusion, but it is commentary, not a model or product release.

Xinzhiyuan · WeChat

Claude Code helps Anthropic double revenue pace in two months

Semi Analysis says Anthropic’s ARR reached $44B, adding $35B over 12 months. Claude Code hit $2.5B annualized revenue by Feb 2026, while inference gross margin rose from 38% to over 70%. The key test is keeping enterprise usage, coding-agent revenue, and inference margin together.

Why it matters: HKR-H/K/R all pass: SemiAnalysis gives hard ARR, Claude Code revenue, and inference-margin numbers. Not a model launch, but it materially shifts the view of Claude Code monetization.

Xinzhiyuan · WeChat

Google Vantage uses AI role-play to assess collaboration under pressure

Google Research and NYU tested Vantage with 188 US participants aged 18-25 on conflict resolution and project management. Its four-layer agent pipeline generates scenarios, applies pressure, extracts behavior, and scores against rubrics; AI-human agreement matched expert-expert Kappa of 0.45-0.64. The key gap is transfer beyond lab settings; the post says Vantage remains a Google Labs research experiment.

Why it matters: HKR-H/K/R all pass: the Vantage study has a strong hook, concrete sample size, and evaluator-risk resonance. It stays in the low featured band because it is still a Google Labs experiment with a narrow cohort.

QbitAI · WeChat

DeepSeek V4’s biggest omission

DeepSeek V4’s technical report omits Engram while listing mHC, CSA, HCA, Muon, and FP4. Engram was open-sourced by DeepSeek and Peking University in January, inserting lookup modules between Transformer layers 2 and 15; its 27B test raised MMLU by 3.4 and Multi-Query NIAH to 97.0%. The engineering signal is CXL pooling: 8 servers shared a 4TB memory pool with under 5% throughput loss.

Why it matters: HKR-H/K/R all pass: the omitted-Engram angle is clickable, with layer ranges, benchmark deltas, and CXL memory-pool numbers. It is analysis, not the V4 launch itself, so 78–84 fits.

May 2Saturday

r/LocalLLaMA

I built Semvec: A constant-cost semantic memory for LLMs, looking for testers

A developer released Semvec, replacing unbounded chat history with fixed-size semantic state. Its 48-turn benchmark claims about 76% token reduction, with identical input footprint at turn 10 and 10,000. It supports OpenAI-compatible LLMs, MCP, Claude Code, Cursor, and multi-agent shared state.

Why it matters: HKR-H/K/R all pass, but this is a Reddit self-release with author benchmarks only. Treat it as an interesting indie memory tool, not a same-day industry story.

Hacker News front page

Show HN: Filling PDF Forms with AI Using Client-Side Tool Calling

SimplePDF released a Copilot demo that fills PDF forms via client-side tool calling; SimplePDF has 200k+ monthly users. PDFs stay in the browser, with parsing, rendering, and field detection local. The demo uses a DeepSeek V4 Flash proxy by default, with BYOK, cloud, or LM Studio options.

Why it matters: HKR-H/K/R pass: the client-side PDF-agent angle is specific, with a clear privacy mechanism and builder relevance. It sits in the 72–77 band as a useful product demo, not a major platform release.

QbitAI · WeChat

Apple Support App Accidentally Shipped Claude.md, Revealing Internal Claude Code Use

Apple Support v5.13 shipped a Claude.md file on May 1 and was pulled within 24 hours. The file describes Juno AI and Live Agents switching through a Protocol layer, with client, agent, and assistant messages handled in one flow. The key issue is release review; the post does not disclose how the file entered production.

Why it matters: HKR-H/K/R all pass, but this is still an app-packaging incident, not a model or platform release. Apple scale and Claude.md details clear the featured bar; the review-chain failure is not disclosed.

TechCrunch · AI

Replit's Amjad Masad on the Cursor deal, fighting Apple, and why he'd rather not sell

Replit grew from $2.8M in 2024 revenue to a billion-dollar annualized target. The excerpt says Cursor is reportedly discussing a $60B SpaceX acquisition; the post does not disclose Masad's full Apple or sale comments.

Why it matters: HKR-H/K/R pass: TechCrunch has the Cursor $60B hook, Replit revenue target, and coding-tool exit tension. The excerpt lacks Masad’s full Apple and sale comments, so this stays in the 72–77 band.

May 1Friday

r/LocalLLaMA

MiMo-V2.5-Pro: the actual best open-weights model

Reddit user cjami benchmarked Xiaomi MiMo-V2.5-Pro in autonomous Blood on the Clocktower games. It scored 88% as Good and 48% as Evil, with 183,639 output tokens per game, $0.99 cost, and a 0.4% tool-call error rate. The key comparison is Kimi K2.6: 580,000 tokens, $2.65, and 10–15 hours per game.

Why it matters: Single Reddit benchmark limits authority, so this is not a model-release story. HKR-H/K/R all pass via a named test with win rates, token counts, cost, and tool-error data, placing it in the 78–84 featured band.