Skip to content

Anthropic / Claude

Everything Anthropic: the Claude models, Claude Code, its safety research agenda and company news.

Latest picks

621–640 of 1,304

Jul 2Thursday

Latent Space

Paul Bakaus on skill engineering and why one-shot AI design is a dead end

Paul Bakaus presented Impeccable at the AI Engineer World’s Fair, an open-source design skill system for coding agents. Instead of one-shot full-site redesigns, users steer output with terms like 'bolder' or 'quieter' that the skill translates into precise design actions. Bakaus calls this 'skill engineering'—compressing expert vocabulary so agents don't converge on generic results. He noted designers now make up at least half of Impeccable's audience, using it as a bridge into code. He rejects full auto mode, arguing the goal is to insert human judgment at the exact point it matters most.

Why it matters: Paul Bakaus introduces 'skill engineering'—packaging designer feedback vocabulary into an open-source instruction set (Impeccable) to steer AI design iteratively rather than one-shot. The concept is novel, backed by a concrete artifact and user data. Score sits at the featured...

Ben's Bites

Fable 5 is back, and there's a new Claude Sonnet 5

Anthropic re-released Fable 5 for paid users with stronger guardrails, available in subscriptions only through July 7 and capped at 50% of usage limits. Scale's benchmark shows it completes 16% of remote work tasks, double Opus 4.8. Claude Sonnet 5 also launched—benchmarked close to Opus 4.8 on agent tasks, cheaper per token but roughly the same cost per task in practice; the author finds it expensive and slow. Google dropped two new models: Nano Banana 2 Lite for fast, cheap images and Omni Flash for video generation and editing. Bridgewater and Thinking Machines trained a financial triage model hitting 84.7% accuracy at 13.8x lower cost than the best frontier model tested.

Why it matters: Anthropic dropped Fable 5's limited return and Sonnet 5 simultaneously — two signals stacked. Scale benchmark provides hard comparable numbers, not pure marketing. Fable 5's 16% task completion rate doubled but absolute number is still low, so not pushing past 90.

AI HOT (Curated Pool)

Fable 5 hits 16.1% automation on freelance jobs in the Remote Labor Index, up from 2.5% eight months ago

The Remote Labor Index tests AI agents on 240 real freelance projects worth $144,000. Fable 5 reached a 16.1% automation rate, nearly double Opus 4.8's 8.3% and well ahead of GPT-5.5's 6.3%. Eight months ago the top score was 2.5%. 22 of Fable 5's projects couldn't be evaluated due to US government access restrictions; even in the worst case its rate would be 14.6%. The study also found AI judges overrate performance badly—GPT-5.5's score was inflated nearly 3x because the AI judge couldn't open professional software to inspect actual deliverables. No model's output passed as finished professional work, but the automation rate has more than quadrupled in under a year.

Why it matters: RLI is one of the few benchmarks using real paid freelance projects; Fable 5 hitting 16.1% — nearly double the runner-up — with a 6x improvement in 8 months is solid. Held below the top band because Fable isn't a tier-1 lab and the post doesn't disclose model size or cost, so ...

Latent Space

AIEWF Day 3: Autoresearch takes the stage, but speakers push back on full autonomy

Day 3 of AIEWF focused on autoresearch. Introspection's Roland Gavrilescu described it as an outer loop where agents maintain the system itself. Anthropic's Thariq Shihipar echoed continuous discovery in his Claude Code keynote, saying models are 'grown, not developed.' Former Google engineering lead Addy Osmani pushed back hard: the outer loop must stay human—inner loop is capability, outer loop is agency. Notion's Geoffrey Litt and Impeccable's Paul Bakaus both argued humans need to understand the code and steer the final 20%. Bakaus stated flatly there will 'never be auto.' Google's Nicole Brichtova added that cultivated expertise sees what average preference misses.

Why it matters: On-the-ground AIEWF report with first-hand quotes from Introspection and Anthropic — not a press release. But it's a conference roundup, not a product launch, so it lands at the featured threshold.

Computing Life · Share · Yage

Claude Science skips the eureka moment and starts as a compute dispatcher

Anthropic's Claude Science desktop app, announced June 30, skips the AI-scientist fantasy and targets the grunt work that eats 80% of researchers' time. It auto-pulls and cleans data from UniProt, PDB, Ensembl, and ChEMBL, then writes SLURM scripts, sets up conda environments, and submits jobs to HPC clusters—retrying on failure. By constraining the model to verifiable execution tasks, it sidesteps hallucination risks. The post does not disclose pricing or a GA date.

Why it matters: Claude Science is a substantive Anthropic product release, and the article nails the positioning — not an 'AI scientist' but a compute orchestrator and data wrangler for research workflows. HKR all hit, but this is a third-party analysis, not a first-party launch post, and the...

Computing Life · Share · Yage

Fable 5's 18-day ban: Anthropic's share went to GLM

Anthropic's Fable 5 was taken offline by US export controls three days after launch, for 18 days. OpenRouter daily token data shows total volume grew from 24T to 32T, but Anthropic's share dropped from 20.7% to 17.6% and its absolute volume shrank. GLM was the biggest winner, jumping from 1.8% to 7.4% share. GLM-5.1 saw a spike on the day GLM-5.2 launched, then collapsed 48 hours later as GLM-5.2 took over. Community sentiment shifted from sympathy for Anthropic to mocking its business strategy. The post notes this only captures API-layer data, not first-party subscriptions, and token volume comparisons overstate GLM's share due to a 10x price gap.

Why it matters: A data-driven attribution of the 18-day Fable 5 ban's competitive impact using OpenRouter daily token data: Anthropic lost 3.1pp share, GLM gained 6.7pp, plus a weird anomaly where GLM-5.1 traffic spiked 4-5x after GLM-5.2 launched. Counterintuitive findings that directly touc...

Computing Life · Share · Yage

Token pricing squeeze: buyers flee, sellers double down, builders fill the gap

Anthropic launched Claude Sonnet 5 at a promo price of $2/$10 per million tokens, rising to $3/$15 in September, while a new tokenizer adds ~30% more tokens per task—nearly doubling real cost over Sonnet 4.6. The next day Palantir's Karp called token billing “completely wrong,” arguing real value should mean outcome-based pricing. On the buy side, Uber burned its full-year AI coding budget in four months after deploying Claude Code and Cursor to 5,000 engineers, then capped per-engineer spend at $1,500/month. Microsoft cut Claude Code licenses for thousands of engineers on the last day of its fiscal year, routing them back to Copilot. GitHub Copilot's switch from flat-rate to token billing triggered a wave of cancellations as users exhausted quotas in a day or two. Sellers can't stop: OpenAI projects a $14B loss in 2026, Anthropic's monthly compute cost runs $1.25B, and Amazon shifted Anthropic payments from per-hour to per-token, pushing its own teams to distill smaller models. Builders in the middle route only 26% of requests to expensive closed-source models, add prompt caching and semantic caching to cut costs 40%–80%, and lean on open-source models whose programming traffic share has pushed Claude Opus down to 4.7%. Tokens won't disappear—they'll recede to a backend meter while front-office billing shifts to per-task, per-seat, or annual budget models.

Why it matters: Anthropic's new pricing collides with buyer flight, backed by Karp's public criticism and Uber's budget blowout — strong signal. HKR all hit, but this is industry analysis rather than a hard news event with cross-source cluster, so 82 featured.

Hacker News front page

Meta Caps Internal AI Token Spending as Costs Near Billions

Meta warned ~6,000 employees that internal AI token costs are on track to hit billions this year. Employees burned through 73.7 trillion tokens in ~30 days, tracked on an internal leaderboard called 'Claudeonomics.' CTO Andrew Bosworth said token volume isn't a measure of impact. Meta is scrapping the leaderboard, rolling out an 'AI Gateway' monitoring dashboard, and will enforce formal token budgets starting in 2027. The company is also steering staff from Anthropic Claude toward its own MetaCode assistant. Uber faced a similar problem—it blew through its 2026 AI coding budget in four months and now caps spending at $1,500 per person per month.

Why it matters: A rare inside look at internal AI spend spiraling at Meta, with concrete numbers and mechanism details. Docked slightly because it's a secondary rewrite of a single-source report (The Information) with no independent verification added by MLQ.

Jul 1Wednesday

AI HOT (Curated Pool)

Meta plans to sell excess AI compute, following SpaceX's playbook

Meta is building a cloud infrastructure business to sell AI compute and model access, putting it in direct competition with AWS, Google Cloud, and Microsoft Azure. The move comes weeks after SpaceX leased its Colossus 1 data center capacity to Anthropic. The pattern suggests data center owners, not model builders, may end up winning the AI race.

Why it matters: Meta selling compute is a notable signal with a fresh angle and a concrete thesis. But the body only has a headline and summary — no pricing, scale, or timeline — so the info density can't support a higher score.

MIT Technology Review · AI

Anthropic launches Claude Science; California's manure carbon math doesn't add up

Anthropic announced Claude Science at an event for pharma execs and biotech founders. It works like Claude Code but for research, autonomously handling computational biology and drug development tasks from short instructions. Anthropic will also use it in-house for rare disease drug research. Separately, the US lifted restrictions on Anthropic's Mythos and Fable models, restoring access today. Another piece digs into California's subsidies for turning cattle manure methane into natural gas—research suggests the carbon offset math is flawed and could lock in more warming.

Why it matters: Anthropic launched Claude Science, extending its Agent model from coding to scientific research, pitched directly to pharma. A significant Claude product line expansion with concrete use cases and internal adoption. Not 90+ because we only have the announcement — no performanc...

AI Chat-Group Daily (群聊日报)

Claude Code found to embed China-user detection; Fable 5 export controls lifted same day

A Reddit reverse-engineering post reveals Claude Code since v2.1.91 silently classifies China-based users via timezone checks and encodes the result into Unicode apostrophe variants in the system prompt. Multiple group members were banned the same day; a reseller said Anthropic targeted Alibaba-related accounts. Meanwhile, the US Commerce Department fully lifted export controls on Fable 5 and Mythos 5. Ford became the top US recall leader after replacing engineers with AI. Sonnet 5 launched at $2/$10 per million tokens but uses a new tokenizer that inflates token counts. WeChat's built-in AI assistant 'XiaoWei' began grayscale rollout, raising privacy concerns as others can invoke it in private chats without consent.

Why it matters: Reddit reverse-engineering post confirms Claude Code uses Unicode steganography to flag Chinese users, with multiple ban reports the same day — high signal density and timeliness. Score capped below 85 because the source is a chat-group digest, not primary reporting, and the p...

Hacker News front page

Anthropic restores global access to Claude Fable 5 after export controls lifted

US export controls on Claude Fable 5 and Mythos 5, imposed June 12, were lifted June 30. Fable 5 returns globally July 1. The controls followed an Amazon researcher's report showing a jailbreak that let Fable 5 identify and demonstrate a software vulnerability exploit. Anthropic confirmed older models including GPT-5.5 and Kimi K2.7 could do the same—no unique Mythos-level capability was exposed. A new safety classifier now blocks the reported technique in over 99% of cases, though false positives on routine coding requests will increase. Anthropic is also working with Amazon, Microsoft, and Google on a shared industry framework for assessing jailbreak severity, and deepening pre-release testing collaboration with the US government.

Why it matters: Anthropic officially announces the lifting of export controls on Fable 5 and its global redeployment — a major status change for a flagship model with strong cross-source signal. All three HKR axes hit: the reversal creates suspense, the timeline and usage numbers are concrete...

Latent Space

Anthropic launches Claude Sonnet 5, but the real story is Fable 5's absence

Anthropic released Claude Sonnet 5 today, calling it the most agentic Sonnet yet with planning, browser/terminal tool use, and a 1M-token context window. Pricing stays at $3/$15 per million tokens, with a promo rate of $2/$10 through late August. The community reaction was muted: benchmarks show it consumed more tokens than Fable, and one test found it cost more than Opus 4.8. The post confirms Fable/Mythos 5 were approved for re-release after government work, but gives no timeline.

Why it matters: Anthropic dropped Sonnet 5 with 1M token context and unchanged list pricing, but tokenizer changes drove 3-6x real consumption—community benchmarks show it may cost more than Fable to run. HKR all hit: the launch is news, the efficiency twist has substance, and Claude users wi...

New York Times Chinese

‘AI Marxism’: How China Is Handling the AI Revolution

The NYT argues China may have an edge in managing AI’s social fallout. After Wuhan taxi drivers protested driverless cabs, Beijing quickly suppressed the outcry but also accelerated policy—its five-year plan now pledges to cushion AI’s job impact. Scholars are developing ‘AI Marxism’ to debate who creates value when machines do the work. The most concrete signal: a Hangzhou court ruled in April that firing an employee after replacing them with AI software is illegal, stating technology should ‘liberate labor.’ The piece contrasts China’s state-driven, job-preserving approach with a US model that lets companies pursue superintelligence largely unchecked. The post does not disclose specific unemployment figures or a timeline for the proposed ‘AI unemployment insurance.’

TechCrunch · AI

Trump drops export restrictions on Anthropic's Mythos and Fable models

The US lifted a license requirement for exporting Anthropic's Mythos and Fable models, widely seen as the most advanced AI released so far. Anthropic said public access resumes July 1. The models were added to the export control list on June 12, making compliance impractical at scale and forcing Anthropic to cut off all public access. Commerce Secretary Lutnick said Anthropic agreed to proactively detect security risks, work with the US government on release standards, and report malicious activity. Anthropic had already publicly pledged to do most of this months earlier, which is partly why security experts were skeptical of the restrictions.

Why it matters: Anthropic's flagship models had export controls imposed and lifted within 18 days, with concrete safety cooperation terms attached — a same-day must-write. Score stays below 90 because the post doesn't disclose specific capability changes for Mythos/Fable, and policy details r...

AI HOT (Curated Pool)

Anthropic embedded steganography in Claude Code to identify users in China

Community reverse engineering found Claude Code reads the local timezone and ANTHROPIC_BASE_URL env var, then checks against an encrypted list of 147 Chinese company domains—Meituan, ByteDance, Moonshot AI included. On match, it tweaks date formatting in the system prompt by 2–3 bits of steganography and sends the marker back to Anthropic's servers. The post doesn't say whether Anthropic has responded, but this directly undermines developer trust.

Why it matters: Anthropic product exposed using steganography to flag Chinese users, involving a list of 147 company domains — method is specific, evidence chain is solid. HKR all hit, industry-shaking level. Not 95+ because only community reverse engineering so far, no official Anthropic res...

Computing Life · Share · Yage

Claude Code embeds steganographic marks for China endpoints, raising enterprise control-plane concerns

A security researcher reverse-engineered Claude Code 2.1.196 and found it embeds environment classification into the system prompt date string using visually similar Unicode characters. The trigger is a custom ANTHROPIC_BASE_URL combined with China timezone or matching 147 known domains and 11 lab keywords including deepseek, moonshot, and zhipu. Separately, GitHub issue #62061 revealed a mechanism for the client to pull extra system prompts from Anthropic's server. Together these show a privileged agent's instruction layer can change without enterprise visibility. The article argues against banning Claude Code and instead recommends treating it as a privileged dev runtime: disable bypass, sandbox execution, build audit surfaces, separate control planes, and maintain model and client fallbacks.

Why it matters: The reverse-engineering work surfaces a concrete mechanism with a character mapping table — not speculation. The story hits security, trust, and geopolitics simultaneously, clearing all three HKR axes. Not scoring higher because it's a single-source reverse-engineering report ...

Hacker News front page

Anthropic restoring access to Claude Fable 5 and Mythos 5 from tomorrow

The US Department of Commerce lifted export controls on Claude Fable 5 and Mythos 5. Anthropic will begin restoring access tomorrow. The post doesn't spell out why the controls were imposed, the exact restoration timeline, or model capability details.

Why it matters: Anthropic announces Commerce lifted export controls on Claude Fable 5 and Mythos 5, access resumes tomorrow. Strong on suspense and user resonance, but the post lacks details on why controls were imposed, scope of restoration, or model specifics—knowledge signal is thin, keepi...

Hacker News front page

US Commerce Department lifts export controls on Anthropic's Claude Fable 5 and Mythos 5

Anthropic tweeted that the US Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5. The post body is just a link to the tweet—no details on which specific controls were removed, when this takes effect, or why these two models were restricted in the first place. All we can confirm right now is what the title says; the rest needs an official follow-up.

Why it matters: Anthropic's official account announcing a policy shift on two named models — direct signal for compliance and cross-border teams. Score held back because the body is just a tweet link with no details on effective date, scope, or original rationale; the headline carries all the...

MIT Technology Review · AI

Anthropic launches Claude Science, a flagship product for AI-driven research

Anthropic launched Claude Science, positioning it alongside Claude Code as a flagship product. It writes code, runs experiments on compute clusters, and prioritizes reproducibility—aimed at computational biology and drug discovery. A live demo showed it identifying drug candidates for phenylketonuria. Anthropic will also use it for in-house rare-disease research. Harvard physicist Matthew Schwartz previously rated Opus 4.5's research ability at the level of a second-year grad student; Claude Science productizes that capability.

Why it matters: Anthropic flagship product launch with a clear positioning and live demo — a same-day must-write. Not above 90 because only a single MIT Tech Review report so far; pricing, availability, and multi-source confirmation are still missing.