Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

121–140 of 760

May 29Friday

AI HOT (Curated Pool)

Xiaomi Open-Sources Controllable Video Foley Model ControlFoley

Xiaomi’s large model application team open-sourced ControlFoley, a controllable video Foley model supporting three tasks: text-guided video dubbing, text-controlled video dubbing, and reference-audio-controlled video dubbing, with code, model weights, and an online demo released.

Why it matters: ControlFoley clears HKR-H/K/R with controllable video Foley plus code, weights, and demo. It is a useful multimodal-audio release from Xiaomi, but not a flagship foundation-model launch, so it sits near the featured threshold.

QbitAI · WeChat

Tencent unveils Code Craft, an AI game creation platform for beginners and developers

Tencent Games unveiled Code Craft, an AI game creation platform that turns natural-language prompts into runnable 2D or 3D games, with a planning knowledge base, Skill system, visual tuning panels, and more than 20,000 free cloud assets; the post does not disclose release timing, pricing, model details, or supported engines.

Why it matters: HKR-H/K/R pass on the Tencent game-creation hook, runnable 2D/3D output, and 20,000+ assets. Pricing, access scope, and model limits are not disclosed, so it stays in the lower featured band.

r/LocalLLaMA

Liquid AI releases LFM2.5-8B-A1B

Liquid AI released LFM2.5-8B-A1B with a 128K context window, 38T pre-training tokens, large-scale reinforcement learning, doubled vocabulary for non-Latin tokenization, and availability on Hugging Face.

Why it matters: HKR-H/K/R pass: 8B/A1B, 128K context, and 38T tokens are concrete hooks for local inference. No benchmarks, license, or deployment limits are disclosed, so it stays in the mid featured band.

Latent Space

Anthropic raises $65B Series H, releases Opus 4.8 and Dynamic Workflows

Anthropic announced a $65B Series H at a $965B post-money valuation, disclosed a $47B revenue run rate, and released Claude Opus 4.8 plus Claude Code Dynamic Workflows as a research preview for parallel subagent orchestration.

Why it matters: HKR-H/K/R all pass: this combines a frontier-lab financing event with an Anthropic model and Claude Code workflow release. I score using the summary’s $65B raise and $965B post-money valuation because the title’s dollar figure conflicts with it.

AI HOT (Curated Pool)

Cursor team releases Developer Habits Report

Cursor’s report says developers’ weekly code output rose from about 3.6K to 8.6K lines, while AI agents increased tool calls per session by roughly 30%.

Why it matters: HKR-H/K/R all pass: Cursor’s own report gives concrete 3.6K→8.6K and +30% figures for AI coding work. It is not a product launch or cross-source event, so 78–84 fits better than the must-write band.

Ruan YiFeng's Weblog

Technology Enthusiasts Weekly Issue 398: Token Costs Are Hard to Afford

Peter Steinberger posted one month of usage showing 7.6 million requests and 603 billion tokens, with CodexBar estimating a $1.3 million value under preset rates rather than his actual spend as an OpenAI employee.

Why it matters: HKR-H/K/R all pass: the CodexBar case turns token economics into concrete usage and cost. This is strong practitioner commentary, not a model or platform release, so it fits the 72–77 featured band.

AI HOT (Curated Pool)

StepFun Releases Step 3.7 Flash, Focused on Agent Efficiency

StepFun released the open-source Step 3.7 Flash model with a 198B-parameter MoE architecture, about 11B active parameters, a 256K context window, and a 67.1 score on ClawEval-1.1.

Why it matters: HKR-H/K/R all pass: the release has a clear sparse-model hook, concrete context and benchmark numbers, and practitioner resonance around open agent efficiency. Official-post sourcing and no independent eval keep it in the 78–84 band.

AI HOT (Curated Pool)

Skill distillation

Skill distillation has Opus 4.7, GPT-5.1, and Gemini 3 Pro write standardized SKILL.md procedure files, while local Qwen 35B and Gemma 26B models execute those files step by step.

Why it matters: HKR-H/K/R pass: the agent-skill distillation pattern is concrete and practitioner-relevant. The summary lacks success rates, cost data, or task outcomes, so it sits at the featured threshold, not must-write.

Latent Space

The Age of Async Agents — Cognition's Walden Yan and OpenInspect's Cole Murray

Latent Space discusses async coding agents with Cognition’s Walden Yan and OpenInspect’s Cole Murray, citing Devin’s 7x merged PR growth and an increase from 16% to 80% of commits across Cognition repos.

Why it matters: HKR-H/K/R all pass: the Cognition repo numbers make this more than agent rhetoric. It stays in the 78 band because it is an interview/trend piece, not a major model or product release.

TechCrunch · AI

Anthropic releases Opus 4.8 with new Dynamic Workflows tool

Anthropic released Opus 4.8 with a Dynamic Workflows tool for coordinating swarms of subagents. The RSS snippet does not disclose pricing, context window size, benchmarks, or a rollout schedule.

Why it matters: HKR-H/K/R all pass: an Anthropic model release plus an agent orchestration tool fits the 85–94 same-day band. Missing price, context window, and rollout detail keep it below the top of the band.

May 28Thursday

AI HOT (Curated Pool)

Perplexity Computer Now Integrates with Microsoft Office

Perplexity Computer is now available in Microsoft Excel, Word, PowerPoint, and Outlook, letting users access Computer from the app sidebar to coordinate work, draft documents, model data, create presentations, and handle email.

Why it matters: HKR-H/K/R pass: Perplexity brings Computer into four Office apps with sidebar workflows, a useful product fact and competitive hook. Price, permission model, enterprise rollout, and measured results are not disclosed, so it stays at the featured threshold.

The Verge · AI

CNN sues Perplexity over “verbatim” copycat articles

CNN sued Perplexity in a New York court on Thursday, alleging its AI answer tools generate “verbatim” copies of CNN work and provide users with information locked behind CNN’s subscription wall.

Why it matters: HKR-H/K/R all pass: CNN vs Perplexity has a clear conflict, concrete allegations, and licensing-risk resonance. It is a notable copyright front, but only a lawsuit filing, not a ruling or product change.

QbitAI · WeChat

A New Paradigm for GUI Agent Trajectories: FSMs Generate Trajectories at $0.04 Each

AutoWebWorld synthesized 29 web environments, 875 pages, and 11,663 verified trajectories at about $0.04 per trajectory, using FSM-defined states, preconditions, and transitions to verify GUI agent tasks instead of human labeling or an LLM judge.

Why it matters: HKR-H/K/R all pass: $0.04 per trace, 11,663 verified traces, and FSM state checks give concrete hooks for GUI-agent data and eval cost. The source is not a top lab release, so it stays in the 78–84 research-tool band.

AI HOT (Curated Pool)

Mistral AI Releases Search Toolkit

Mistral AI released the public preview of Search Toolkit, an open source framework that combines data ingestion, retrieval, and evaluation behind shared interfaces for cloud, on-premises, or edge deployment.

Why it matters: HKR-K and HKR-R pass: Mistral combines RAG ingestion, retrieval, and evaluation in an open-source Search Toolkit with cloud, local, and edge deployment. HKR-H is weak, so this sits at the featured threshold for a mid-weight product update.

Mistral AI

Mistral upgrades Le Chat into unified agent Vibe, covering office work and coding

Mistral upgraded Le Chat into a unified AI agent called Vibe, with one license covering both office work and coding. Existing chats, settings and plans all carry over. Work Mode supports enterprise knowledge search, structured data analysis, document and report generation, scheduled multi-step tasks and reusable skills, and connects to Google Workspace, Outlook, SharePoint, Slack, GitHub and more.

Why it matters: It discloses Vibe's Work Mode, coding mode and CLI updates in full, so readers can judge how it plugs into existing workflows.

Alibaba Technology · WeChat

AI-Native Project Management: Two Git Repos Replace Weekly Updates, Insights, and Metrics Reports

Zhou Zhiwei describes a project-management setup that uses two Git repositories, an AI coding assistant, Shell, and Python to replace at least 80% of manual weekly-update chasing, data moving, chart generation, and engineering-metrics reporting.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives an 80% replacement claim and a two-repo mechanism, and it hits engineering-management toil. This is a strong practical workflow piece, not a model or platform launch.

AI HOT (Curated Pool)

Mistral AI launches physics AI model for industrial engineering

Mistral AI integrated the Emmi AI team and launched a physics AI foundation model for industrial engineering, with the post saying it can learn from geometry, boundary conditions, or measurement data and predict full physical fields on a single GPU in seconds.

Why it matters: HKR-H/K/R pass: a major model lab entering physics simulation with a concrete single-GPU seconds claim. The score stays in the lower featured band because model name, benchmarks, pricing, and access are not disclosed.

The Verge · AI

YouTube Will Let You Ask AI to Make a Custom Video Feed

YouTube is rolling out an AI custom video feed that uses a user-entered prompt or suggested options to build a personalized homepage feed around interests, moods, or topics; the feature supports English first and is available to signed-in users in the US on YouTube’s mobile app and desktop site.

Why it matters: HKR-H/K/R pass: promptable YouTube feeds are a concrete consumer-AI UX shift with rollout conditions. No model, ranking mechanism, or performance data is disclosed, so this stays in the mid-weight product-update band.

Financial Times · Technology

Kirkland & Ellis to Spend $500mn Building Its Own AI Technology

Kirkland & Ellis plans to spend $500mn building its own AI technology platform for the “collective intelligence” of its lawyers; the post does not disclose model architecture, vendors, or launch timing.

Why it matters: FT source authority and a $500mn in-house AI budget give strong HKR-H/K/R for legal AI adoption. Missing model architecture, vendors, and launch timing keep it in the 72–77 enterprise-adoption band.

AI HOT (Curated Pool)

Grok Build 0.1 on API

xAI released Grok Build 0.1 in public beta through the xAI API for agentic coding tasks, with throughput above 100 tokens per second and pricing at $1 per million input tokens and $2 per million output tokens.

Why it matters: HKR-H/K/R all pass, but this is a 0.1 public-beta API and pricing launch; benchmarks, context window, and task success rates are not disclosed. It fits a solid mid-weight product update at 78, featured not p1.