Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

541–560 of 760

Apr 21Tuesday

The Verge · AI

Fortnite developers can make AI characters now — just don’t try to date them

Epic Games is rolling out a “conversations” tool for Fortnite creators, turning island NPCs into AI characters that can talk with players in unscripted ways. The snippet says creators define persona, knowledge, behavior, and voice with prompts; the title says don’t try to date them, but the post does not disclose the exact guardrails or moderation system.

Why it matters: This is a mid-weight product update that gives Fortnite creators AI NPC conversation tooling. It clears all three HKR axes, but moderation rules, pricing, and base model details are not disclosed, so it stays at the low end of featured.

Apr 20Monday

r/LocalLLaMA

Training LoRA adapters for Apple's on-device 3B model on a free Colab T4 and a Mac

The author built a QLoRA pipeline for Apple’s on-device 3B model, cutting training needs from about 24GB to about 1GB RAM and 5GB GPU, enough for a free Colab T4 or a 24GB Mac. The post says A100 LoRA, T4 QLoRA, and Mac QLoRA adapters perform about the same, raising accuracy from about 40% to 75%, or 86% with retrieval; it also reports a confirmed Apple bug that writes a hidden ~160MB cache copy per CLI call, reaching 269GB over ~300 runs.

Why it matters: A named first-person experiment with reproducible memory and accuracy numbers clears HKR-H/K/R and beats routine tutorial posts. The score stays below the 85 band because this is a single Reddit post with limited source authority and a narrow benchmark scope.

r/LocalLLaMA

Compared some models for feature planning

A Reddit user tested 9 models on planning a “load tracking” feature for a Go budgeting app, then used Claude Code to rank the generated specs, with Claude Opus 4.6 placed first. The table shows Opus 4.6 produced a 19 KB spec with 44 code reads at $2.47; GLM 5.1 ranked second and Qwen 3.6 35B fp8+vLLM ranked third. Do not treat this as a benchmark: the author says it is not representative, and the post does not disclose any manual quality review yet.

Why it matters: A named first-person test gives real workflow data, so HKR-H/K/R all pass. The ceiling stays low: one task only, ranked by Claude Code itself, and no human acceptance result is disclosed, so this lands at the low end of featured.

r/LocalLLaMA

TRELLIS.2 image-to-3D now runs on Mac (Apple Silicon) with no NVIDIA GPU required

A developer ported Microsoft's TRELLIS.2 to Apple Silicon and reports generating ~400K-vertex meshes from one photo in about 3.5 minutes on an M4 Pro with 24GB. The port replaces five CUDA-only extensions with PyTorch MPS and custom backends; texture baking takes about 18 seconds, removing the NVIDIA and cloud requirement.

Why it matters: This is a community port, not an official release, but HKR-H/K/R all pass: the hook is NVIDIA-free image-to-3D on Apple Silicon, and the post includes testable details (M4 Pro 24GB, ~400k vertices, 3.5 minutes, 5 CUDA-extension rewrites). Reddit-level source authority keeps it in

Synced · WeChat

How to Do Vibe Coding Correctly? A Masterclass from Anthropic's Coding Agent Lead

Anthropic researcher Erik Schluntz said his team merged a 22,000-line production change, mostly written by Claude, cutting work from two weeks to one day. His workflow spends 15-20 minutes on repo exploration and planning, limits edits to leaf nodes, keeps humans on core logic, and validates with long stress tests plus a few E2E tests. The key issue is boundary control, not handing AI the system core; he also said task length AI can handle doubles about every seven months.

Why it matters: HKR-H/K/R all pass: this is an Anthropic field report with concrete numbers and reproducible workflow rules for production coding agents. It stays at featured, not p1, because it is a strong practitioner lesson rather than a major model or product launch.

r/LocalLLaMA

Using Qwen3.6 via LM Studio as a Claude Code subagent, saving 30x Opus tokens per task

A Reddit user routed Qwen3.6 through LM Studio as a Claude Code subagent and reported about 30x lower Opus marginal tokens on two audit tasks. In the examples, a 23-file route audit dropped from 13k to 0.4k marginal tokens, and an 18-file Astro site inventory fell from 89k to 3k; the setup used unsloth’s Qwen3.6-35B-A3B-MXFP4_MOE gguf on a 64GB M4 Max with a 64k context window. The key mechanism is offloading extraction and audit work to a local OpenAI-compatible server, while the post also says quality was mixed rather than strictly better than Opus.

Why it matters: A named first-person experiment with 2 clear token comparisons hits HKR-H, HKR-K, and HKR-R: strong hook, concrete setup details, and direct cost relevance for Claude Code users. It stays below p1 because the evidence is a Reddit post with only 2 tasks.

Hacker News front page

Show HN: TRELLIS.2 image-to-3D running on Apple Silicon, no Nvidia GPU needed

Developer shivampkumar ported Microsoft's 4B-parameter TRELLIS.2 to Apple Silicon with PyTorch MPS for single-image 3D generation. He replaced flash_attn, nvdiffrast, and custom sparse conv kernels with pure PyTorch sparse 3D conv, SDPA attention, and Python mesh extraction. On an M4 Pro with 24GB, it generates ~400K-vertex meshes in about 3.5 minutes; slower than H100 seconds, but fully offline.

Why it matters: Strong on all HKR axes: a clear hook, concrete implementation details, and benchmark-like numbers. This is not a Microsoft model launch, but a reproducible local port with real practitioner relevance, so it lands in featured rather than p1.

The Verge · AI

Cloud development platform Vercel was hacked

Vercel confirmed a security incident affecting a “limited subset” of customers, and hackers are trying to sell stolen data. The RSS snippet says exposed data includes employee names, emails, and activity timestamps; Vercel says a compromised third-party AI tool was the attack path, but the post does not disclose the vendor or scope.

Why it matters: HKR-H/K/R all pass: the breach angle is strong, and the post gives concrete leak fields plus an AI-tool entry path. Vendor identity and blast radius are still undisclosed, so this stays low-featured; no hard exclusion applies.

Apr 19Sunday

QbitAI · WeChat

Did Musk Really Sell Lao Gan Ma on Douyin?

QbitAI says the shown “Musk selling Lao Gan Ma on Douyin” and “GTA-6 crossover” images were generated by OpenAI GPT Image 2; the claimed 100K+ live viewers were part of fake visuals. The post argues Image 2 can render realistic posters, game screenshots, and readable long text, and links that to Codex-style UI workflows; the post does not disclose pricing, rollout scope, or launch timing. The real issue is verification: image realism is eroding “photo as evidence.”

Why it matters: HKR-H/K/R all pass: the hook is novel, the article shows a concrete capability jump, and the trust/verification angle resonates with practitioners. It stops short of p1 because the body does not disclose rollout, pricing, or an official launch scope.

r/LocalLLaMA

I tested 8 LLMs as tabletop GMs: a 27B model beat the 405B on narrative quality

The author tested 8 LLMs on 6 fixed tabletop-GM scenarios, and google/gemma-3-27b-it ranked first in narrative quality with a 4.33 overall score. The probe used 8 auto metrics plus 3 LLM-judge scores, and the full run cost about $0.02; the title says a 27B beat a 405B, but the snippet does not disclose the 405B model name or full rankings.

Why it matters: A named first-person benchmark with a strong surprise hook clears HKR-H, HKR-K, and HKR-R. I kept it at featured, not higher: the source is Reddit, the post is truncated, and the 405B model name plus full ranking are not disclosed.

r/LocalLLaMA

Deep dive into LangGraph’s Pregel execution model, checkpointing internals, and DeepAgents

A technical post breaks down LangGraph as a high-level wrapper over a Pregel runtime, with PregelNodes, channels, and reducers as the core primitives. The RSS snippet cites four Postgres checkpoint tables, a Plan/Execute/Update superstep flow, and compile() preflight validation; the post does not disclose benchmark numbers in the snippet. The real takeaway is the unified runtime view of parallel execution, checkpoint write amplification, and subgraph boundaries.

Why it matters: HKR-H/K/R all pass: the post reframes LangGraph as a Pregel runtime and adds concrete internals like 4 checkpoint tables and Plan/Execute/Update supersteps. Kept at 74 because this is a Reddit deep dive, not an official release, and no benchmark or production case is disclosed.

Apr 18Saturday

QbitAI · WeChat

OpenClaw has reached the milk tea business

Guming and Intime Retail said OpenClaw tests exposed 5 deployment risks: default port 18789 exposure, at least 8% malicious Skills, privilege overreach, 20+ minutes of runaway token use, and weak legacy defenses. Reported incidents include an agent closing a normal bastion-host port and locking out ops staff, plus requests for unrelated permissions like microphone access. The real issue is not chat UX but agents touching enterprise networks, credentials, and production systems.

Why it matters: This is not generic AI-safety commentary; it documents five concrete deployment risks and one ops outage, so HKR-H/K/R all pass. It stays below P1 because the evidence is still case-level testing, with no official fix, broad rollout impact, or cross-source cluster.

Synced · WeChat

What is OpenAI prioritizing under compute limits?

Greg Brockman said OpenAI narrowed priorities under hard compute limits to two bets: a personal assistant and AI workers that solve hard user problems, and current compute cannot fully support both. The snippet says Sora resources were reduced while focus shifted to reasoning models, a unified AI layer, and the next base model Spud; it does not disclose the claimed compute budget, timeline, or model specs. The key point is not a B2B retreat but a compute-driven reprioritization.

Why it matters: HKR-H/K/R all pass: the compute-ceiling angle is strong, the piece adds concrete priority shifts, and OpenAI roadmap triage hits cost and dependency nerves. It stays at 80 because this is secondary reporting; spend, timing, and technical details are not disclosed.

Synced · WeChat

Claude Design enters research preview for generating mockups, prototypes, and slides

Anthropic launched Claude Design in research preview for Claude Pro, Max, Team, and Enterprise users, covering mockups, prototypes, slides, and one-pagers. Powered by Claude Opus 4.7, it can ingest codebases, images, DOCX, PPTX, XLSX, and web captures, then export to Canva, PDF, PPTX, and HTML; the headline cites Figma and Adobe stock drops, but the post does not disclose the moves. The real signal is the workflow link from design system ingestion to handoff into Claude Code.

Why it matters: HKR-H/K/R all pass: the design-workflow angle is novel, the post gives concrete mechanism details, and the Figma/Adobe pressure point resonates. I keep it below 85 because the stock-drop claim has no numbers and there is no user test, pricing, or adoption data.

Xinzhiyuan · WeChat

Bilibili debate: Hermes responds to plagiarism claims for the first time, as MiniMax moves early on Harness

MiniMax says its M2.7 model now handles 30%-50% of daily workflows in its RL team, ran over 100 self-optimization loops, and improved evals by 30%. The post also says Hermes Agent grew from 2B to nearly 300B daily tokens, while M2.7 exceeds 25B daily tokens on OpenRouter; Hermes lead Tommy Eastman denied copying EvoMap in a livestream. The real signal is Harness: the post cites 20-40ms or 80ms sandbox startup and 15k to 600k instances per minute, showing competition is shifting from benchmark scores to agent execution infrastructure.

Why it matters: HKR-H/K/R all pass: the plagiarism-response angle pulls clicks, and the story carries concrete metrics on workflow share, self-optimization loops, sandbox latency, and concurrency. It stays at 83 because this is a dense secondary report, not a primary launch or official technical

Latent Space

[AINews] The Two Sides of OpenClaw

Peter Steinberger released two talks contrasting OpenClaw’s public story with its engineering reality, citing 60x more security reports than curl and at least 20% malicious skill contributions. The RSS snippet calls OpenClaw the fastest-growing open-source project in history, but the post does not disclose its architecture, launch date, or governance model. The real signal is attack-surface growth outrunning governance.

Why it matters: This clears HKR-H with the public-story vs engineering-reality split, HKR-K with the 60x and 20% figures, and HKR-R because open-agent security debt is a live industry nerve. It stays in featured, not higher, because the post does not disclose OpenClaw’s architecture, release, or

X · @dotey

Anthropic designer Ryan Mather shares Claude Design tips while covering 7 product lines

Anthropic designer Ryan Mather shared 9 Claude Design workflow tips while covering 7 product lines. The RSS snippet says to spend 1 hour building a design system, use chat for large changes, comments for small edits, specify feedback like 8px spacing, and attach only the target component folder instead of a full monorepo. The key shift is process: from human-do/human-review to Claude-do/human-review.

Why it matters: This is a strong practitioner workflow note: an Anthropic insider shares concrete, reusable tactics, so HKR-H/K/R all pass. It stays below the 80s because this is not a formal Claude product release and the post does not disclose harder outcome data such as time saved or task win

Hacker News front page

Show HN: AI Subroutines – Run automation scripts inside your browser tab

rtrvr.ai introduced AI Subroutines, which turn a recorded browser task into a callable tool and replay it at zero token cost and zero LLM inference delay. The script runs inside the active tab, reusing auth, CSRF, TLS sessions, and signed headers; recording trims about 300 requests to about 5 and falls back to DOM-only when GraphQL operation IDs are volatile. The part to watch is batching: one LLM call can assign parameters for a 500-row sheet and launch 500 subroutines.

Why it matters: This clears HKR-H/K/R: the hook is zero-token browser automation, the post gives concrete mechanics (300→5 requests, DOM fallback, 500-row fan-out), and it hits agent reliability/cost pain. Kept to mid-featured because it is a single-company Show HN post, not a market-wide event.

X · @dotey

Anthropic launches Claude Design, a conversational design generation product

Anthropic released Claude Design in research preview and is rolling it out to Pro, Max, Team, and Enterprise subscribers. Powered by Claude Opus 4.7, it can start from text, images, docs, or web clips, then iterate via chat, comments, direct edits, and sliders. On first use it reads a team's codebase and design files to build a design system; outputs export to Canva, PDF, PPTX, or standalone HTML, with one-click handoff to Claude Code.

Why it matters: Anthropic pushes Claude into a new design workflow, so HKR-H/K/R all pass. The post includes rollout tiers, model name, first-run design-system ingest, and export paths; strong featured story, but still a gradual research preview rather than a top-tier model release.

Apr 17Friday

Hacker News front page

Measuring Claude 4.7's tokenizer costs

The author used Anthropic's free count_tokens API to compare Claude Opus 4.6 and 4.7 on 7 real samples and 12 synthetic ones; the real-sample weighted total rose from 8,254 to 10,937 input tokens, or 1.325x. Technical docs hit 1.47x, a real CLAUDE.md file hit 1.445x, while Chinese and Japanese stayed near 1.01x. On a 20-prompt IFEval sample, 4.7 improved strict prompt-level pass rate from 85% to 90%; the post cannot isolate tokenizer effects from model weights or post-training.

Why it matters: HKR-H/K/R all land: the post has a sharp cost hook, reproducible token-count data, and clear budget impact for Claude Code users. It stays below p1 because this is a third-party measurement, not an Anthropic release, and the IFEval slice is only 20 items.