Claude Code’s Next Era — Thariq Shihipar, Anthropic
Anthropic 的 Thariq Shihipar 在 Latent Space 播客中谈 Claude Code 的下一阶段,包括 Ask User Question、artifacts、Claude Tag、Projects 和可自定义 harness 的 Claude Mods。
Anthropic 的 Thariq Shihipar 在 Latent Space 播客中谈 Claude Code 的下一阶段,包括 Ask User Question、artifacts、Claude Tag、Projects 和可自定义 harness 的 Claude Mods。
Simon Willison 反驳「MCP 一直是坏主意」的观点,认为该文忽略了 MCP 当下的价值:若运行 Claude Code、Codex、Meta Muse、OpenClaw 等拥有无限制互联网访问的完整终端智能体,确实几乎无需 MCP,直接调用 API 即可。
GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,但审查力度应按风险分级;Skills 与 MCP 解决不同问题,前者是打包的团队经验,后者是连接工具与数据的标准,可组合使用;RAG 并未死亡,它为模型提供训练数据之外的相关信息,减少 token 浪费并让回答更有依据。
Thariq shared 10 Claude Code tips that shift review from checking outputs to steering the right task, with concrete practices including full upfront context, /goal, Workflows for parallel tasks, self-checking, and comparison reports.
Why it matters: This is a strong Claude Code workflow tutorial, with concrete tactics around task calibration, /goal, and Workflows self-checks. It lands in the 72–77 tutorial band; the insider source and all three HKR hits justify featured.
The article reverse-engineers Claude Design from Anthropic’s open-source Design plugin and describes a six-layer structure; the snippet only discloses mechanisms such as workflow decomposition, aesthetic injection, evaluation transfer, and connector abstraction.
Why it matters: HKR-H/K/R all pass, but this is third-party reverse engineering rather than an Anthropic launch. It fits the high-quality Claude/agent mechanism analysis band just above featured threshold.
The author retired an OpenClaw AI assistant after more than one month of use; the post says it required self-hosting, API setup, and keeping one home computer running 24 hours a day.
Why it matters: HKR-H/K/R all pass, but this is a personal weekly write-up, not a model or platform release. The month-long OpenClaw use and 24/7 PC requirement make it just clear the featured threshold.
Xu Xiaobin argues that agent-based development compresses the intent-to-code loop from weeks or months to minutes, using a weekly-report system, a multi-role agent development setup, and image-repository provisioning as examples; the article identifies mismatches in Git, CI, code review, release flows, permissions, harness setup, and dry-run validation.
Why it matters: HKR-H/K/R all pass, but this is infrastructure commentary rather than a model or product launch. The named cases and week/month-to-minutes claim put it in the 72–77 featured band.
@mvanhorn shared an agent engineering workflow centered on a Research→Plan→Work loop, plan.md constraints, and 22 practical tips; the snippet says it covers planning, parallel execution, input methods, and remote control, but the post does not disclose the full tool stack list.
Why it matters: HKR-H/K/R all pass, but this is a practitioner methods post, not a model or product release. The full tool stack is not disclosed, so it sits at the featured threshold.
The Claude Code engineering team described process changes after making agentic coding the default at Code w/ Claude SF 2026: JIT planning, asking Claude first for context collection, Claude handling style and tests in code review, and humans focusing on legal and safety judgments.
Why it matters: First-party Claude Code workflow post with concrete engineering mechanisms and strong HKR-H/K/R fit. It is not a model or major product release, so it stays in the 78–84 band.
An Anthropic developer shared a Claude Code understanding-verification workflow with 8 steps, using incremental teaching, user restatement, checklists, and quizzes to confirm the human can defend the problem, solution, and impact before moving to the next stage.
Why it matters: HKR-H/K/R all pass: a concrete Claude Code workflow with an 8-step verification loop and a strong oversight hook. It is a practical tutorial, not a product release, so it sits at the lower featured band.
ReinforcedKnowledge analyzes ByteDance’s verl RLHF loop, covering DataProto plus rollout, reward, advantage, and update paths. The author stopped a private fork because near-daily upstream changes made sync cost exceed refactoring work, and describes an NCCL hang fixed on one node by setting NCCL_SOCKET_IFNAME=lo.
Why it matters: Niche but useful RL post-training field report, not an industry release. HKR-H comes from the fork-then-quit twist; HKR-K has verl’s five paths and NCCL_SOCKET_IFNAME=lo; HKR-R hits the cost of maintaining open-source training forks.
Deezer reports that more than 50,000 AI-generated songs are uploaded each day, while Recording Academy CEO Harvey Mason Jr. says AI is now present in every recent music session he has attended and Grammy rules still bar AI music from the industry’s highest honors.
Why it matters: HKR-H/K/R all pass, but this is a podcast-style policy discussion rather than a model, product, or binding regulation story. The concrete signal is the 50,000/day Deezer figure plus the Grammy eligibility conflict.
The author used Claude Opus 4.8 to turn Nonviolent Communication into an AI Skill through a six-step workflow, taking about 45 minutes, using roughly 300,000 tokens, and costing under RMB 20.
Why it matters: HKR-H/K/R all pass: this is a numbered first-person Claude workflow with concrete cost and token details. It stays in the lower featured band because it is a personal tutorial, not an Anthropic release or model update.
Peter Steinberger posted one month of usage showing 7.6 million requests and 603 billion tokens, with CodexBar estimating a $1.3 million value under preset rates rather than his actual spend as an OpenAI employee.
Why it matters: HKR-H/K/R all pass: the CodexBar case turns token economics into concrete usage and cost. This is strong practitioner commentary, not a model or platform release, so it fits the 72–77 featured band.
Skill distillation has Opus 4.7, GPT-5.1, and Gemini 3 Pro write standardized SKILL.md procedure files, while local Qwen 35B and Gemma 26B models execute those files step by step.
Why it matters: HKR-H/K/R pass: the agent-skill distillation pattern is concrete and practitioner-relevant. The summary lacks success rates, cost data, or task outcomes, so it sits at the featured threshold, not must-write.
Latent Space discusses async coding agents with Cognition’s Walden Yan and OpenInspect’s Cole Murray, citing Devin’s 7x merged PR growth and an increase from 16% to 80% of commits across Cognition repos.
Why it matters: HKR-H/K/R all pass: the Cognition repo numbers make this more than agent rhetoric. It stays in the 78 band because it is an interview/trend piece, not a major model or product release.
Zhou Zhiwei describes a project-management setup that uses two Git repositories, an AI coding assistant, Shell, and Python to replace at least 80% of manual weekly-update chasing, data moving, chart generation, and engineering-metrics reporting.
Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives an 80% replacement claim and a two-repo mechanism, and it hits engineering-management toil. This is a strong practical workflow piece, not a model or platform launch.
Lv Ruofan proposes the Agent Room model: multiple agents share context, a task ledger, Memory, Runtime, and Artifacts, and two software-engineering cases show the system moving from workflow automation toward collaborative judgment rather than predefined task routing.
Why it matters: HKR-H/K/R all pass, but this is a methodology piece rather than a model launch or open-source framework. Concrete Agent Room mechanisms and 2 R&D sites put it in the 72–77 featured band.
The author proposes writing a Skill before asking AI to execute a task; each Skill should include three elements—success criteria, observed pitfalls, and deterministic tools—and can be organized through index.md plus AGENTS.md or CLAUDE.md for reuse.
Why it matters: HKR-H/K/R pass via a concrete Skill-first workflow and reusable agent practice. No model release, product capability, or experiment numbers, so it sits at the featured threshold.
Yage argues that users should externalize work before execution by writing reusable Skills for Claude Code, Codex, and Cursor. The post gives an Outlook email example: spend about 30 minutes documenting username, phone approval, and client choice, then have AI read that file on later runs.
Why it matters: HKR-H/K/R all pass, but this is a workflow tutorial rather than a product or model release. The concrete Skill mechanism and Outlook example clear the featured floor; weak source authority keeps it at 72.
Hugging Face’s post frames an agent as three layers: Model, Scaffolding, and Harness; Scaffolding defines behavior through prompts and tool descriptions, while Harness runs model calls, tool calls, and control loops.
Why it matters: HKR-H/K/R pass: the Hugging Face post gives a concrete agent-stack taxonomy. It clears featured on practitioner relevance, but lacks a release, benchmark, or deployment case, so it stays at the threshold.
HomoAgens1 runs a local ReAct orchestration loop on Qwen3.6-35B-A3B, with about 3B active parameters, a 12GB GPU, 30 expert offload, and 40 tokens/s prompt generation; smaller dense models fail first on tool-call discipline, inventing arguments or repeating bad calls, while reasoning is not identified as the first break point.
Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment rather than a formal release. The VRAM, speed, and failure-mode details put it at the 72 featured threshold.
Salesforce has adopted a headless architecture that lets salespeople update data through AI; the post says MCPs, HTML, audio, and web interfaces can be generated dynamically by context, but it does not disclose implementation metrics or adoption numbers.
Why it matters: HKR-H/K/R all pass, but this is a software-form thesis without user metrics, launch timing, or a reproducible test. It fits the insightful-commentary band, not a must-write release.
Cursor summarizes lessons from building cloud agents: after migrating to Temporal, reliability rose above 99.9%, and the platform processes more than 50 million operations per day.
Why it matters: HKR-H/K/R all pass: Cursor is central to coding agents, and the post gives Temporal, 99.9%+ reliability, and 50M daily operations. Not a launch, so it stays at low-end featured.
Zhan Xupeng published a roughly 50-minute article on Agent theory and a personal assistant implementation, covering memory, ReAct planning, progressive skill loading, subagents, and harness-level fault recovery.
Why it matters: HKR-K/R pass via concrete agent mechanisms and practitioner reliability pain; HKR-H is weak because the headline is a standard tutorial frame. This fits the quality-tutorial threshold, not the 78+ news band.
A Reddit user tested Claude structured outputs and reported 90–95%+ first-pass parse rates with tool_use, typed schemas, enum constraints, and stepwise validation, versus 65–70% for prompt instructions followed by regex or JSON parsing and retries.
Why it matters: HKR-H/K/R all pass: the post has a clear reliability contrast and concrete 90–95%+ vs 65–70% numbers. Source authority is limited to one Reddit experiment, with sample and task details not disclosed, so it stays at low featured.
A Reddit user says an API-log audit showed Cursor and Claude Code recursively grep about 40 files in 10k-plus-line repositories, sometimes load 2k-line files for 5-line edits, and spend roughly 30k tokens on tool definitions and logs before generating code.
Why it matters: HKR-H/K/R all pass: the hook is contrarian, the API-log numbers are concrete, and coding-agent context waste is a live practitioner pain. Reddit single-post sourcing and no shared logs keep it at the featured threshold.
Former Microsoft executive Matt Veloso said Microsoft generated about $30 billion from its AI partnership between 2023 and 2025, while related costs reached $100 billion; he also said actual usage among paid Copilot users is below 3%.
Why it matters: HKR-H/K/R all pass: a former executive gives concrete Microsoft AI cost, revenue, and Copilot usage numbers. Kept at 80 because this is a single former-exec claim, not an official Microsoft disclosure.
The author used Claude Code to generate a single-file HTML project plan page in 2 minutes, with a dark theme, timeline, and collapsible tables; the comparable Notion template previously took 30-40 minutes.
Why it matters: HKR-H/K/R all pass: the post has a concrete Claude Code workflow hook, a 2-minute vs 30-40-minute comparison, and clear practitioner resonance. Scope is small, so it sits at the featured threshold.
A Reddit user compared 11 Hermes Agent alternatives across open-source and managed options; OpenClaw is listed with 347k GitHub stars, 24+ integrations, and 9 CVEs in four days, while TrustClaw uses OAuth-only sandboxed execution and Perplexity Computer requires a $200/month Max tier.
Why it matters: HKR-H/K/R all pass: this is a practical agent-tool comparison with 11 items and concrete integration/security figures. Reddit single-post sourcing limits confidence, so it stays near the featured threshold.
Peter Steinberger used 603 billion tokens across 7.6 million requests in 30 days, with the bill exceeding $1.3 million; he said disabling fast mode cut the price by 70%, and OpenAI does not charge him for the tokens.
Why it matters: HKR-H/K/R all pass: the story has a sharp cost hook, concrete usage numbers, and strong practitioner resonance. It is a first-person bill disclosure, not an OpenAI pricing or product launch, so it sits just above the featured threshold.
Anthropic published Founder’s Playbook, arguing that AI tools such as Claude Code reduce prototyping cost but increase startup failure risk across the Idea, MVP, Launch, and Scale stages through false validation, confirmation bias, agentic technical debt, and founder decision bottlenecks.
Why it matters: HKR-H/K/R pass: the Anthropic founder playbook has a sharp counterintuitive angle, a four-stage mechanism, and clear founder resonance. It stays near the featured floor because no dataset or reproducible test is disclosed.
Claude API prewarms prompt cache with the system prompt, skips output, then hits cache on the real request.
Why it matters: HKR-H/K/R all pass: this is a concrete Claude API latency mechanism, not a vague product tease. It clears featured, but it is a mid-weight inference update rather than a major model or capability release.
A Reddit user ran Qwen 3.6 27B on a dual RTX 3090 Ubuntu setup, reporting 48GB VRAM, a 262k context window, no NVLink, about 4000 pp/s prompt processing, and 113 tk/s generation.
Why it matters: All HKR axes pass, and this is a first-person local-inference run with concrete numbers. Source is a single Reddit post with limited reproducibility detail, so it sits at the low featured threshold.
Andrej Karpathy says 90% of AI coding bills is wasted on unnecessary context, including repeated full-repository sends, expensive models for simple tasks, and missing prompt caching.
Why it matters: HKR-H/K/R all pass via the 90% claim, named waste mechanisms, and practitioner cost pain. It reaches featured, but stays at 72 because the post gives no billing sample or reproducible test.
Anthropic engineer Thariq argued for using HTML instead of Markdown and gave 5 reasons; the post says HTML generation takes about 2 to 4 times longer than Markdown.
Why it matters: HKR-H/K/R all pass, but this is a developer format debate rather than a model or product launch. Named Anthropic/Karpathy context and the 2-4x time figure clear the featured threshold at the low end.
The author used AI to build a HealthKit export app in about 5 minutes and run multivariate regression, finding that the last post-dinner AI usage time correlated negatively with sleep duration; after avoiding AI at night, average sleep increased by 1 hour and 40 minutes.
Why it matters: HKR-H/K/R all pass: a first-person quantified experiment links post-dinner AI use to shorter sleep, then reports +1h40m after stopping. Personal-blog scope keeps it below major industry-update territory.
The author used AI to build a HealthKit export tool and run multivariate regression, found that post-dinner AI use correlated negatively with sleep duration, and added 1 hour 40 minutes of average nightly sleep after avoiding AI for several weeks.
Why it matters: HKR-H/K/R all pass: the personal reversal is clickable, the HealthKit/regression setup adds testable detail, and sleep loss hits AI practitioners directly. Scope is anecdotal, so it stays at the featured floor.
Simon Willison demonstrates using an LLM command in a script shebang line, with fragments generating SVG, the -T option calling llm_time, and a YAML template defining Python tools to compute 2344×5252+134 and return 12,310,822.
Why it matters: HKR-H/K/R all pass: Simon Willison shows a reproducible LLM-in-shebang workflow with concrete flags. Impact stays within CLI/script automation, not a model or platform release, so it sits in the low featured band.
Karpathy argues that LLM output is moving from Markdown toward richer HTML, while interactive neural video still has an open problem: how to combine neural generation with precise traditional software.
Why it matters: HKR-H/K/R pass: Karpathy gives a fresh UI frame, a concrete Markdown→HTML→neural-video path, and a builder-facing product question. Single X post with no data keeps it at the featured floor.