Skip to content

Anthropic / Claude

Everything Anthropic: the Claude models, Claude Code, its safety research agenda and company news.

Latest picks

261–280 of 1,304

Sep 3Thursday

Computing Life · Share · Yage

OpenAI Codex's self-wake mechanism: it sets its own alarm to watch CI after fixing code

A system prompt template merged into OpenAI's open-source codex repo in late August reveals how Codex Persistent mode actually works: it's not a 24/7 always-on process, but a wake-check-sleep loop every 1–3 minutes. The template requires the agent to record its goal, latest status, completion condition, and next check time before sleeping, then decide what to do upon waking. One hard rule: persistence does not broaden authorization scope—anything beyond scope requires explicit permission. WIRED reported on this mode earlier, but media headlines saying 'always-on' clash with the code's 'sampled again' language. OpenAI hasn't launched it yet; the backend request still shows 'disabled.' ProAgentBench shows models achieve only 64.4% accuracy in judging when to proactively help, and Anthropic's engineering blog reports a 17% miss rate on real overreach during automated review—two numbers that explain the hold. Tasks suited for it are delivery-type jobs like CI, deployment, and builds that execute for one minute and wait for ten. Open-ended tasks like writing proposals or designs are a bad fit. Three discipline rules from the template can be adopted today: write four-element checkpoints, stay silent when nothing has changed, and prefer deterministic mechanisms.

Why it matters: High information density with concrete sourcing from the open-source repo — reveals the real wake-check-sleep loop and the authorization scope rule. Deduction because this is interpretation of a template, not an official launch; actual product experience is unknown.

AI HOT (Curated Pool)

Claude can now use your computer in the background while you do other things

Claude Cowork and Claude Code can now take over your computer in the background—clicking, typing, opening apps—while you switch to other tasks. The post doesn't disclose latency, permission boundaries, or supported operating systems.

Why it matters: Anthropic shipped background computer use to both Cowork and Claude Code — a substantive product update for the Claude ecosystem. All three HKR axes hit: the UX shift is novel, the dual-product rollout signals productization, and it directly lands with heavy Claude users. Held...

AI HOT (Curated Pool)

Anthropic publishes a guide to effective commerce agent architecture and open-sources a reference implementation

Anthropic's post explains how to turn models like Claude into commerce agents that actually work in production, focusing on architecture, latency, and cost. They also open-sourced a reference implementation called commerce-agents. The full article body isn't available yet—only the title and lede are shown—so specific architecture details, latency figures, and cost breakdowns are still missing.

Why it matters: Official Anthropic guide plus open-source repo hits H and K, but the body is title-only right now — no architecture details, latency numbers, or cost breakdowns are public. Policy says default to the lower band when key facts are missing, so 72 at the featured threshold. If th...

Sep 2Wednesday

Latent Space

Anthropic drops Claude Fable/Mythos 5.1: new SOTA for coding, but 70% more output tokens

Anthropic launched Claude Fable 5.1 and Mythos 5.1 on Sep 1, claiming SOTA on coding and knowledge work. Fable 5.1 hits 55.8% on Terminal-Bench 4.0 and is pitched for autonomous multi-step tasks. Cache read price dropped 75% to $0.25/MTok, but Artificial Analysis found output tokens rose 1.7x, netting a ~20% per-task cost increase. Community speculation suggests Fable and Mythos may share weights with different safety routing—the post doesn't confirm this. Early praise for coding ability is offset by complaints about rate limits, false safeguard triggers, and subscription UX.

Why it matters: Anthropic dropped Claude Fable/Mythos 5.1 with a 55.8% Terminal-Bench 4.0 score, a 75% cache read price cut to $0.25/M tokens, and a 70% increase in output tokens. A capability upgrade plus major pricing shift makes this a same-day must-write. Not a 95 because we only have Lat...

AI Chat-Group Daily (群聊日报)

DeepSeek V4 Flash beats Sol in real-world use; Anthropic drops Fable 5.1

Community members ran two-month SBS comparisons and a week-long 5.1B-token workload on DSH + DeepSeek V4 Flash, concluding it feels better than GPT-5.6 Sol in real tasks. Sol overthinks and produces bloated output; V4 Flash is fast (2.3s first token) and cost ¥362.84 total. A 'subscription gym paradox' theory argues subscription-based harnesses quietly throttle usage while pay-per-token models don't. Anthropic launched Fable 5.1 with 75% cheaper cache reads, but Fable 5 scored below Opus 5. Also: Astra hits Critical cybersecurity tier, Anthropic's $35B compute deal, Qwen 3.8-Max-0902 benchmark run, Microsoft AI secretary setup, and Grok Bot hands-on.

Why it matters: The side-by-side data is solid — 5.1B tokens, ¥362.84 total spend, 2.3s first-token latency — but the source is an anonymized chat log, not an official release or reproducible benchmark. That caps the authority. HKR all hit, so featured is the right tier.

Hacker News front page

Simon Willison tests Claude Fable 5.1's pelican benchmark across five reasoning levels

Simon Willison ran his classic 'SVG of a pelican riding a bicycle' prompt against Claude Fable 5.1 at five reasoning levels. Low and medium produced near-identical outputs with no visible reasoning, taking ~23 seconds and ~10 cents. At xhigh the model spent 7m51s and $1.83, adding real detail. Max ran for 13m54s and $3.30, delivering his best Anthropic pelican yet—blue hat, basket with a fish, feet on pedals—though he still says it lacks the flair of Gemini 3.7 Flash. Separately, Fable 5.1 hit 52.6% on the new Terminal-Bench-Science 0.1 benchmark, up from 24.7% for Fable 5.

Why it matters: Simon Willison ran a controlled five-tier reasoning comparison on Claude Fable 5.1 with concrete latency and cost numbers, making it more useful than the official announcement. Score stays below 85 because this is a personal evaluation rather than a major capability breakthrou...

Hacker News front page

Anthropic banned a paying user for "suspicious signals" — no warning, no human appeal

A long-time Claude Max subscriber at $200/month got banned overnight with a template email citing "suspicious signals" — no clause, no example, no human appeal path. A colleague in the Philippines was banned right after paying $100 for Max. GitHub issues show similar cases in Brazil, Singapore, and Malaysia, some hitting multiple linked accounts within hours of upgrading. The author argues Anthropic's enforcement feels like 2010s Google account bans: automated, opaque, and disproportionately painful for individuals who depend on the product. Enterprise customers get account managers; Max users get a no-reply address and a reference ID. The author filed an appeal but no longer trusts a single frontier lab with their entire workflow, and plans to diversify across other providers and open-weight models.

Why it matters: Multiple paid users across countries report sudden bans after payment — not an isolated glitch. HKR all hit, but this is a user complaint, not an official statement, so capped at 78 due to information asymmetry.

The Verge · AI

Anthropic launches Claude Fable 5.1, up to 45% cheaper for agentic work

Anthropic released Fable 5.1 and Mythos 5.1, directly addressing customer complaints about cost, data retention, and overzealous safeguards. Fable 5.1 outperforms Fable 5 while costing ~25% less typically and up to 45% less for complex agentic tasks, driven by lower pricing on cached data. Every CEO Dan Shipper called it the strongest coding model they've used, now fast, token-efficient, and speaking like a normal person. The post doesn't spell out Mythos 5.1 specs or detailed pricing.

Why it matters: Anthropic drops Fable 5.1 and Mythos 5.1 with a clear cost-reduction story for agent workloads — up to 45% cheaper via cached call pricing. Concrete performance and pricing details make this a strong signal. Held at 85 rather than higher because we only have the headline and s...

TechCrunch · AI

OpenAI's Astra model is on the way — and very good at breaking into computer systems

OpenAI shared safety details on Astra, its first LLM to hit a 'critical cybersecurity threshold.' Astra can find and exploit unknown security flaws without human guidance. OpenAI plans to release it soon but will limit access to its most advanced cyber capabilities. This mirrors concerns Anthropic raised about its Mythos model earlier this year.

Why it matters: OpenAI's first public safety assessment of Astra confirms the model has crossed the autonomous vulnerability exploitation threshold, with a gated release planned. This directly parallels Anthropic's handling of Mythos earlier this year — the second case in 2026 of a top lab re...

AI HOT (Curated Pool)

Claude Fable 5.1 tops Artificial Analysis Intelligence Index, but per-task cost is 20% higher than Fable 5

Artificial Analysis tested Claude Fable 5.1 at max effort and it scored 66, hitting #1 on their Intelligence Index. The trade-off: per-task cost is 20% higher than Fable 5. The post doesn't break down task types or latency—just the headline and a one-line result.

Why it matters: A new Anthropic model tops a third-party benchmark with concrete score and cost data — enough substance to feature. But without task breakdowns or latency numbers, it's a solid news bite, not an 85+ story.

TechCrunch · AI

Anthropic's Fable 5.1 is cheaper and less restrictive

Anthropic bumped Fable and Mythos to 5.1. Fable 5.1 is now cheaper and triggers fewer false-positive safety refusals; it's live today on cloud platforms and the API. Mythos 5.1 remains restricted to registered cybersecurity and life sciences partners. A key change is zero data retention—clients can run the model on their own infra with no data outflows, rolling out this fall. The post doesn't disclose specific price cuts or benchmark comparisons.

Why it matters: Anthropic updates both Fable and Mythos lines simultaneously — Fable gets cheaper with fewer false refusals, Mythos stays gated. Zero data retention is the hardest new fact here, but the post doesn't disclose specific price cuts or refusal-rate numbers, so the score stays at 78.

Hacker News front page

Rewriting 65k lines of Go to Rust with Fable cost $400

The author rewrote a 65k-line terminal editor from Go to Rust using Fable 5 for $400. The method has three steps: extract code into a data representation (state machines, graphs, formulas), operate on that representation, then regenerate code in the target language. Fable's precise data-flow tracing is the key enabler. The post doesn't report compilation pass rate or test coverage, so I'd discount the 'fully autonomous' claim until those numbers surface.

Why it matters: 65k lines Go-to-Rust for $400 with a clever intermediate-representation approach hits H and K. But the post doesn't disclose compile pass rate or test coverage, so 'fully automated rewrite' needs a discount — lands at 72, right at the featured threshold.

AI HOT (Curated Pool)

Claude Fable 5.1 is live on OpenRouter, targeting agentic coding and long-running workflows

Anthropic released Claude Fable 5.1 on OpenRouter as a direct upgrade to Fable 5. The focus areas are agentic coding, long-running workflows, visual code generation, finance, and analytics. The post doesn't disclose benchmark numbers or pricing changes, so I'd wait for third-party evals.

Why it matters: Anthropic model update with four clear focus areas, directly relevant to Claude developers. But no benchmarks, no pricing, no third-party evals in the post — stays at 78, the featured threshold, pending real-world testing.

AI HOT (Curated Pool)

Claude Fable 5.1 lands on Claude Code and Platform, cache reads 75% cheaper

Anthropic shipped Claude Fable 5.1 and Mythos 5.1 together. Pricing matches Fable 5, but API cache reads are 75% cheaper. The model stays autonomous longer on long tasks, flags when it's stuck more proactively, and writes more naturally. The post doesn't disclose latency, context window, or benchmark scores—I'd discount the 'most advanced' claim until numbers land.

Why it matters: Anthropic shipped Fable 5.1 and Mythos 5.1 together with a 75% cache-read price cut — a real cost improvement that heavy Claude Code users will care about. Missing latency, context window, and benchmark numbers keeps it from scoring higher, but the price drop and tooling updat...

Hacker News front page

Anthropic launches Claude Fable 5.1 and Mythos 5.1, cutting price by 25% and targeting coding and scientific research

Anthropic released two models, Fable 5.1 and Mythos 5.1—same underlying model, different safeguards. Fable 5.1 is generally available; Mythos 5.1 is gated behind trusted access programs for cybersecurity and life sciences. Fable 5.1 beats Fable 5 across coding, knowledge work, and long-horizon tasks, while costing ~25% less on typical workloads and up to ~45% less on highly agentic work. Enterprise Frontier Safeguards (EFS) will let customers keep data in their own cloud infra, rolling out in phases from fall 2026; until then, eligible customers get zero data retention. Cybersecurity false positives dropped 60%, and the model can discover vulnerabilities but not build exploits. The post does not disclose parameter count, context window, or training details.

Why it matters: Anthropic's flagship model refresh with a dual-track release (Fable 5.1 for everyone, Mythos 5.1 gated behind trusted projects) is an industry first. Coding and long-horizon tasks beat the previous gen across the board, and bio capabilities are strong enough to require governm...

Hacker News front page

Claude Fable 5.1: same price, stronger at long-running coding and multistep research

Anthropic updated its platform docs for Claude Fable 5.1. Pricing matches Fable 5, with cache reads at a quarter of the cost. The focus is stronger long-running agentic coding, multistep research, and document, spreadsheet, and slide work. Three breaking changes: forced tool use now errors, earlier models can't read its thinking blocks, and editing earlier turns invalidates thinking blocks. Five additive features include mid-conversation effort changes, turn-scoped system messages, and readable progress between tool calls—some marked beta. The post doesn't include benchmark scores or latency figures.

Why it matters: Anthropic ships Claude Fable 5.1 with a 4x cache cost reduction and three breaking changes developers need to watch. Solid product update with direct cost and workflow impact for Claude-heavy users. Not scoring higher because it's a docs-only release so far — no independent be...

Sep 1Tuesday

AI Chat-Group Daily (群聊日报)

Claude Code's journey from 2 likes to global phenomenon, ChatGPT Ads hits $1B run rate

Boris from Anthropic walked through Claude Code's full origin story on Lenny's podcast—the internal launch post got just 2 likes. The team used an 'underfund' principle: deliberately starve projects of headcount but give them unlimited tokens, forcing everything to be 'Claudified.' Boris hasn't manually written a line of code since last November. Separately, ChatGPT Ads hit a $1B annualized run rate in under 200 days, but the analysis argues agents and ads are fundamentally at odds—agents compress decision steps that ads depend on. The group also debated whether solo builders beat teams, using Overcooked as the litmus test.

Why it matters: Claude Code lead's first full retrospective on going from zero to global adoption, with concrete numbers backing the underfund principle and Boris's zero-manual-coding practice. All three HKR axes hit, but the source is a chat-group digest's secondhand summary rather than the ...

Anthropic News

Anthropic launches Enterprise Frontier Safeguards with customer-held data and keys

Anthropic released Enterprise Frontier Safeguards (EFS), which pairs zero data retention (ZDR) privacy with safety monitoring for abuse detection. Data sits in the customer's own cloud infrastructure rather than at Anthropic.

Why it matters: The piece details EFS's data retention and monitoring architecture, so readers can weigh privacy against safety when deploying frontier models.

Computing Life · Share · Yage

Salesforce shipped the same APIs twice—only the productized version got traction

In April 2026, Salesforce opened its entire platform to AI agents via Headless 360 and hosted MCP servers. Over the next four months, almost no enterprise adopted it—developers hit config errors, per-user auth bottlenecks, and missing docs. In August, the same underlying capabilities relaunched as Claudeforce, a Claude plugin with 37 sales-specific skills, one-click admin setup, inherited permissions, and native embedding in the Claude UI. The stock jumped 23% on announcement day, though earnings and a short squeeze drove most of it. The product enters public beta in September; pricing and named external customers are still missing. The article frames the gap as a four-layer stack—interface, semantics, governance, distribution—and argues April only delivered the first layer.

Why it matters: A sharp case study using Salesforce's own A/B test: same pipes shipped twice, only the packaged product got traction. Concrete timeline and failure analysis, not fluff. Slight discount because it's a single-source analysis rather than breaking news, and the enterprise-software...

AI HOT (Curated Pool)

Anthropic launches Claude Fable 5.1 and Mythos 5.1, cache reads drop to $0.25/MTok

Anthropic released Claude Fable 5.1 (successor to Fable 5 for long-running coding and research) and Mythos 5.1 (Project Glasswing only). Both default to 1M context, 128k max output, and the same $10/$50 per MTok pricing as Fable 5. The standout change: cache reads are now $0.25/MTok—2.5% of base input price, far cheaper than other models. tool_choice drops 'any' and 'tool' modes; thinking blocks only work with same-gen or newer models; text watermarking and C2PA credentials are on by default. The post doesn't include benchmark scores or performance comparisons.

Why it matters: Anthropic dropped two new models on the same day — a substantive product update. Fable 5.1 targets long coding and research sessions; Mythos 5.1's gated release adds intrigue. 1M context, 128K output, same price — high info density. The post doesn't detail Mythos 5.1's capabil...