Skip to content

Anthropic / Claude

Everything Anthropic: the Claude models, Claude Code, its safety research agenda and company news.

Latest picks

941–960 of 1,304

May 26Tuesday

Xinzhiyuan · WeChat

Chinese agent SkyClaw targets Opus 4.6-level performance with free trial

Kunlun Tech released SkyClaw-v1.0 and SkyClaw-v1.0-lite with a 2-4 week free trial, claiming SkyClaw-v1.0 input costs are 1/24 of DeepSeek V4 Pro and about 1/43 of Sonnet 4.6.

Why it matters: HKR-H/K/R all pass: SkyClaw-v1.0 has a sharp cost hook, concrete trial and pricing ratios, and budget resonance. Source facts remain vendor claims, so it stays at the low featured band.

New York Times Chinese

Pope Leo XIV Challenges Silicon Valley and Warns of AI Risks

Pope Leo XIV issued the 42,300-word encyclical Magnifica Humanitas, warning that AI amplifies the power of people with economic resources, expertise, and data access, and calling for regulation and transparency.

Why it matters: HKR-H/K/R all pass: the hook is unusual, the article gives a 42,300-word encyclical and a concrete power-concentration claim, and the topic hits regulation and safety accountability. Not a model, product, or company-moving event, so 78 featured.

AI HOT (Curated Pool)

OpenAI GPT-5.6 Reportedly Set for Next Month With 1.5M-Token Context

Developers found an unannounced OpenAI GPT-5.6 entry in Codex backend logs under the codename iris-alpha, with a 1.5 million-token context window, about 43% higher than GPT-5.5’s 1.05 million-token limit.

Why it matters: HKR-H/K/R all pass: the Codex-log leak, 1.5M-token window, and 43% increase are concrete and practitioner-relevant. It stays below 85 because this is not an official GPT-5.6 launch.

r/LocalLLaMA

Update on a 12×32GB SXM V100 Cluster for Local Legal Drafting

A lawyer runs a local legal-drafting pipeline across 16 GPUs, with Qwen3.5-122B-A10B reaching about 50 tok/s on four V100s, while a verifier blocks ungrounded citations, dates, and Bates numbers before any final document is used.

Why it matters: HKR-H/K/R all pass: this is a first-person local-LLM experiment with concrete numbers, not a vendor post. Reddit source limits authority, so it stays at the low featured band rather than p1.

AI HOT (Curated Pool)

Anthropic co-founder Chris Olah speaks at Pope Leo's encyclical launch

Chris Olah raised three AI governance questions at the Vatican, saying frontier labs face commercial, research, and geopolitical pressures that can conflict with doing the right thing, and that external oversight is essential.

Why it matters: HKR-H/K/R all pass, driven by the Vatican-Olah hook and the concrete claim about commercial, research, and geopolitical pressure. It lacks the full question list or a new policy mechanism, so it stays below the 78–84 band.

May 25Monday

Financial Times · Technology

Tech Giants Need Oversight to Protect National Security

The FT headline says tech giants need oversight for national security. The snippet names Anthropic and SpaceX and proposes one presidentially nominated, Senate-confirmed director on their boards, but the post does not disclose an implementation mechanism.

Why it matters: HKR-H/K/R pass: the FT piece names Anthropic and SpaceX and gives a concrete board-seat proposal. It is commentary, not enacted policy, and implementation details are not disclosed, so it sits at the featured threshold.

AI HOT (Curated Pool)

Harness, Scaffold, and AI Agent Terminology Explained

Hugging Face’s post frames an agent as three layers: Model, Scaffolding, and Harness; Scaffolding defines behavior through prompts and tool descriptions, while Harness runs model calls, tool calls, and control loops.

Why it matters: HKR-H/K/R pass: the Hugging Face post gives a concrete agent-stack taxonomy. It clears featured on practitioner relevance, but lacks a release, benchmark, or deployment case, so it stays at the threshold.

May 24Sunday

Xinzhiyuan · WeChat

Anthropic’s Three Cards Surface: Mythos 1 Appears, Opus 4.8 Spotted

Xinzhiyuan says Anthropic’s claude-opus-4.8 appeared in Google Vertex AI, while a 59.8MB Claude Code source-map leak with 512,000 TypeScript lines exposed Sonnet 4.8 references and Mythos 1 clues tied to Claude Code and Claude Security.

Why it matters: HKR-H/K/R all pass, but this is a leak plus Vertex listing, not an Anthropic launch. No capability numbers, pricing, context window, or reproducible evals, so it stays in the 78–84 band.

r/LocalLLaMA

Vision-capable LLMs vs. OCR for long-document QA with charts, images, and tables

The author tested Claude Sonnet 4.5 on 171 questions from 30 image-heavy MMLongBench-Doc PDFs, comparing native PDF vision use with OCR pipelines. Native PDF ranked fifth of six at 52.0% accuracy and cost $0.2552 per query, while LlamaCloud premium with full context reached 59.6% at $0.1885 per query.

Why it matters: HKR-H/K/R pass: the post gives 30 PDFs, 171 questions, accuracy, and per-question cost for long-document QA. Limited sample and Reddit sourcing keep it in the featured-threshold band.

May 23Saturday

AI HOT (Curated Pool)

Anthropic reportedly nears over $30B funding round, with valuation set to overtake OpenAI

Bloomberg reports that Anthropic is nearing a funding round of over $30 billion, expected to close as soon as next week, pushing its valuation above $900 billion, while the company projects second-quarter revenue of $10.9 billion and its first profitable quarter.

Why it matters: HKR-H/K/R all pass: Bloomberg’s reported $30B+ round, $900B+ valuation and $10.9B Q2 revenue make this a same-day Anthropic capital-race story. It is still not officially closed, so it stays in the lower 85-94 band.

AI HOT (Curated Pool)

v2.1.149 release summary

Claude Code v2.1.149 adds categorized /usage reporting, an enterprise allowAllClaudeAiMcps setting for cloud MCP connectors, and fixes three security issues involving PowerShell permission bypass, Git worktree sandbox allowlist overflow, and otelHeadersHelper failures when script paths contain spaces.

Why it matters: Official Claude Code point release with concrete changes but limited blast radius: /usage categories, an enterprise MCP allow switch, and PowerShell bypass fixes hit developer security and governance needs.

AI HOT (Curated Pool)

Claude Auto Mode Adds Pro Plan and Model Support

Claude Auto Mode is now available on the Pro plan and supports Sonnet 4.6 and Opus 4.7; users can start it with Shift+Tab, while the post does not disclose pricing changes or rollout scope.

Why it matters: HKR-H/K/R all pass: official Claude dev channel gives Pro access, two supported models, and a shortcut. This is a mid-weight Claude product update, not a major model or capability release.

AI HOT (Curated Pool)

Project Glasswing: Initial Update

Anthropic says Project Glasswing used Claude Mythos Preview with about 50 partners to find more than 10,000 high or critical vulnerabilities in global critical systems, with independently verified accuracy of 90.6%.

Why it matters: HKR-H/K/R all pass: Anthropic gives concrete numbers—~50 partners, 10,000+ high/critical bugs, 90.6% validation—and the story hits AI-agent security automation and critical-system risk.

Bloomberg Technology

Anthropic to Close Over $30 Billion Round as Soon as Next Week

Anthropic plans to close a funding round of over $30 billion as soon as next week at a valuation above $900 billion, Bloomberg reported, citing people familiar with the matter, which would put it ahead of OpenAI as the world’s most valuable AI startup.

Why it matters: Bloomberg reports Anthropic may close a $30B-plus round next week at a $900B-plus post-money valuation, a frontier-lab capital-structure story. HKR-H/K/R all pass; the deal is not closed, so it stays below the 95 band.

AI HOT (Curated Pool)

Project Glasswing Collaborative AI Cybersecurity Project Reports Results

Anthropic says Project Glasswing and its partners found more than 10,000 high or critical vulnerabilities in key software since the initiative launched last month; the post does not disclose the vulnerability list, reproduction conditions, or remediation status.

Why it matters: HKR-H/K/R all pass: Anthropic ties AI security work to 10,000+ severe flaws. Missing vulnerability lists, reproduction details, and fix status keep it in the featured-threshold band, not p1.

May 22Friday

AI HOT (Curated Pool)

Karpathy’s CLAUDE.md Four Rules Raise AI Coding Accuracy to 94%

Karpathy published a 65-line CLAUDE.md with four rules that raised AI coding accuracy from 65% to 94%, and the file received over 220,000 GitHub stars.

Why it matters: HKR-H/K/R all pass: a notable name, a claimed accuracy jump, and a rules-based Claude Code workflow. It stays below 85 because the body only gives summary-level numbers; task set, evaluation method, and the four rules are not disclosed.

Xinzhiyuan · WeChat

Microsoft, after investing $13B in OpenAI, saw its engineers run up Claude Code costs

Microsoft plans to end Claude Code subscriptions by the end of June for its Experiences and Devices teams and move nearly 100,000 engineers to GitHub Copilot CLI, with the article attributing the change to external token-based billing costs.

Why it matters: HKR-H/K/R all pass: the OpenAI-Claude contrast hooks, the story gives end-June migration, nearly 100k engineers and token-billing, and it hits enterprise coding-agent cost control. Not a model release or official major launch, so 78–84 fits.

Xinzhiyuan · WeChat

Enterprise Agent Operations Begin? Anthropic Updates Architecture, Chinese Tech Firms Have It Running

Alibaba Cloud JVS Crew splits Agent, Environment, and Session into three layers, with sandboxes, snapshot recovery, RBAC, and usage-based billing. Anthropic added self-hosted sandboxes to Claude Managed Agents on May 19, while the article cites 2-week deployments and 5x or 10x efficiency gains in several Chinese customer cases.

Why it matters: HKR-H/K/R all pass, but the facts are an enterprise agent-infra comparison: Anthropic self-hosted sandboxes and Alibaba Cloud JVS Crew architecture. This is featured-level, not a must-write model release.

Hacker News front page

Show HN: Spec-Driven Development Workflow for Claude Code

The sddw author released a Claude Code plugin that splits work into requirements, code analysis, and design specs, then clears context after each step to keep cost and context focused.

Why it matters: HKR-H/K/R all pass for a Claude Code workflow with a concrete spec-and-context mechanism. It stays in the 72–77 featured band because the post lacks benchmarks, adoption data, or an official Anthropic release.

AI HOT (Curated Pool)

v2.1.147 Release Update

Claude Code v2.1.147 adds a Workflow tool, disabled by default, for deterministic multi-agent orchestration, and renames /simplify to /code-review with code-correctness reporting and GitHub PR inline-comment generation.

Why it matters: HKR-H/K/R all pass: the official Claude Code release adds a default-off Workflow tool for deterministic multi-agent orchestration. No performance data, pricing, or scope limits are disclosed, so this stays in the mid product-update band.