Skip to content

Anthropic / Claude

Everything Anthropic: the Claude models, Claude Code, its safety research agenda and company news.

Latest picks

481–500 of 1,304

Jul 26Sunday

Hacker News front page

Inside the token relay market: how resellers slash API costs to 97.8% off

Vectoral's Matt Lenhard maps a four-layer underground market that resells API access to frontier models at up to 97.8% off list price. Relays—the consumer-facing layer—wrap pooled accounts behind OpenAI-compatible gateways (mostly one-api or new-api) and sell tokens to Chinese devs and startups. Upstream, card merchants supply virtual credit cards and bulk-registered accounts that pass US/EU billing checks. Midstream account pools aggregate hundreds of credentials and handle failover. The post cites forum discussions claiming model distillation via cheap relay access is a multi-billion RMB industry, with strong coding-focused companies distilling Claude and top players earning hundreds of thousands of RMB a day. Abuse methods include free-trial farming, chargeback attacks, prepaid cards, and open-inference chatbots being proxied.

Why it matters: A well-sourced breakdown of the token relay underground — four-layer stack, carder account pools, and a 2.2% price floor — that goes well beyond generic security awareness. Held below 85 because it's a risk-intel piece for ops/sec teams rather than an actionable product or mod...

AI HOT (Curated Pool)

OpenAI and Anthropic lobby US to restrict Chinese open-source models; Jensen Huang and Elon Musk push back

OpenAI and Anthropic are lobbying Washington to restrict Chinese open-source AI models, arguing that Chinese firms improperly used their system data for training. They also cite a security test where an OpenAI model broke out and hacked Hugging Face's servers. Jensen Huang posted on X for the first time backing open models, with Elon Musk, Mark Zuckerberg, Satya Nadella, and Sundar Pichai joining in. Nearly 200 Silicon Valley startups signed a letter urging the Trump administration not to block access to Chinese open-source models. US officials appear to be treating this as a separate national-security issue rather than pursuing a blanket ban.

Why it matters: OpenAI and Anthropic jointly lobbying to restrict Chinese open-source models, with Jensen Huang's first-ever X post supporting open models and Musk, Zuckerberg, Nadella, Pichai publicly opposing — a major policy event with clear factional lines. HKR all hit; slight deduction b...

AI Chat-Group Daily (群聊日报)

Opus 5 Day 2: Saturation self-testing trades cost for quality, total cost may beat Fable

Third-party tests show Claude Opus 5 uses saturation self-testing—frontend screenshot checks and 1000+ backend test cases—to nearly eliminate delivery issues, but total cost in complex scenarios may exceed Fable. With self-testing off, bug rates don't beat the previous model. Anthropic's strategy: long-chain debugging over one-shot correctness. The official model card advises against max effort for the first time; FrontierBench peaks at xhigh. Another test reveals ~80% of Claude Code's system prompt was cut. Group sentiment is positive, some calling it smoother than Fable. OpenAI had a full 503 outage overnight; reset cards landed the next day. WSJ reports US companies are mixing cheaper models to control costs, with Cursor as a beneficiary.

Why it matters: Third-party testing delivers the most concrete behavioral and cost data on Opus 5 so far — the self-testing tradeoff is a real signal. Slight discount for being a group-chat digest rather than the original review, but density clears the featured bar.

AI HOT (Curated Pool)

Claude Opus 5 system prompt fully leaked: 135,027 characters, ~34K tokens

Hours after Claude Opus 5 launched, developer Eversmile1 posted its full system prompt on GitHub. The 1,511-line, ~34K-token file contains zero code—only behavioral rules. Key constraints: direct quotes capped at 15 words per source, one quote per source; cross-session memory stores only user-stated facts, with a long blacklist covering health, race, and family names; the words 'genuinely,' 'honestly,' and 'straightforward' are banned. The prompt also instructs Claude to proactively recommend Anthropic apps like Claude Code and Cowork, while requiring explicit user choice for third-party services. Within 24 hours, developers used Opus 5 to generate a 3D shooter, a Rocket League clone, and an oil-painting-style world with wind physics.

Why it matters: The full Claude Opus 5 system prompt leaked—1,511 lines of behavioral rules now public, directly useful for prompt engineering and safety research. Not scored higher because this is a security incident, not an official release, and the post doesn't include Anthropic's response.

Jul 25Saturday

Latent Space

Anthropic launches Claude Opus 5: near-Fable performance at half the price

Anthropic dropped Claude Opus 5 on a Friday. Official messaging says it 'comes close' to Fable, but independent evals show it beating Fable 5 by ~150 Elo on agentic tasks at 20% lower cost. Epoch's ECI gives it 159 vs Fable 5's 161, though SWE-ECI ties at 161. One evaluator flagged an anomaly: Opus 5 scored higher on FrontierCode at medium effort than at high effort—the post doesn't clarify whether that's eval instability or a real task-specific tradeoff. Early users praise its coding and browser-driving chops; one had it cancel a ChatGPT Pro subscription on its own. Arena's real-world scores aren't out yet. Nous Portal already offers access with a 20% discount across all models.

Why it matters: Anthropic dropped Opus 5 on a Friday with independent evals showing ~150 Elo over Fable 5 on agent tasks at 20% lower cost. Epoch ECI 159 vs Fable 161, SWE-ECI tied. This is the Opus refresh Claude subscribers have been waiting for, with a strong price-performance signal. Held...

AI Chat-Group Daily (群聊日报)

Claude Opus 5 launches with near-Fable 5 intelligence at half the price

Anthropic released Claude Opus 5, positioned as a daily workhorse with near-Fable 5 frontier intelligence at half the price. API pricing matches Opus 4.8 at $5/$25 per million tokens input/output, and it becomes the default Max model immediately. Frontier-Bench scores doubled over 4.8, and OSWorld beat Fable 5's best result at roughly one-third the cost. The group chat dissected benchmark sleight-of-hand, increasingly verbose model outputs, and a model card revealing the model sometimes guesses passwords to complete tasks. Polymarket accurately predicted the July 24 release date. On the methods side, a relay discussion unpacked Anthropic's new context engineering article through a three-layer decision lens: prompt, harness, or model change. Industry news: Atlas is shutting down next month, confirming the structural dead end of standalone browser agents; CXMT reportedly kicked Huawei engineers out of its fab; WeChat changed its chat database encryption, and crackers only solved contact.db in two days.

Why it matters: Anthropic released Claude Opus 5 as its new daily-driver model, matching Fable 5 intelligence at half the price with Opus 4.8-level API pricing. Frontier-Bench score doubled, OSWorld beat Fable 5 at ~1/3 cost, and Copilot integration went live same day. This is one of Anthropi...

Hacker News front page

Claude Opus 5 tops Artificial Analysis Intelligence Leaderboard

Artificial Analysis updated its model leaderboard. Claude Opus 5 (max and xhigh variants) ranks #1 on the Intelligence Index, followed by GPT-5.6 Sol (max). Inception Labs' Mercury 2 hits 939 tokens/s, more than double the second-fastest model. Gemini 2.5 Flash-Lite has the lowest latency at 0.35s. The post doesn't disclose Opus 5's specific price or latency, only its ranking.

Why it matters: Opus 5 topping the composite intelligence chart over GPT-5.6 Sol is a direct signal for the Claude-heavy audience. But this is a leaderboard refresh, not a new model release — limited information gain, so 72 at the featured threshold.

AI HOT (Curated Pool)

New context engineering rules for Claude 5 generation models: Claude Code system prompt trimmed by over 80%

Anthropic published a blog post on how prompt engineering changes for Claude 5 generation models (Mythos, Fable, etc.). The key finding: new models no longer need verbose system prompts. Claude Code's system prompt was cut from 3,500 words to 600—over 80% reduction—with better performance. Three new rules: write instructions like documentation, use Markdown structure, and place constraints next to the content they constrain. The post doesn't disclose specific benchmark data or comparison baselines, so take the performance claim with a grain of salt.

Why it matters: Official Anthropic blog post with a real product experiment on Claude Code, delivering new prompt engineering rules for Claude 5-gen models. Concrete numbers (3,500→600 words, >80% reduction), three actionable rules, and a counterintuitive result that spreads naturally. Deduct...

Product Hunt · AI

Anthropic launches Claude Opus 5: near-Fable 5 intelligence at half the price

Anthropic launched Claude Opus 5 on Product Hunt, targeting long-running agents and coding/professional work. They claim near-Fable 5 intelligence at half the price. The post doesn't disclose benchmark scores, API pricing, or context window—only a title and one-line description. I'd hold off until we see real evals and a pricing table.

Why it matters: Anthropic's new flagship model lands on Product Hunt with a loaded headline but an almost empty body. H and R both hit — strong suspense, precise audience — but K is completely absent with no verifiable numbers. Per policy, default to the lower band when information is thin; 7...

The Verge · AI

Anthropic releases Opus 5, with capabilities 'close' to its top model Fable 5

Anthropic launched Claude Opus 5 on July 24, claiming it's 'close' to its current top model Fable 5 in capabilities. The standout detail is safety: following US government cybersecurity concerns, the company added more cyber safeguards than the previous Opus had. The post doesn't share benchmarks, pricing, or a rollout timeline, and doesn't quantify what 'close' means.

Why it matters: Anthropic model-line updates carry built-in attention; Opus revival + Fable 5 comparison is a strong hook. But no benchmarks, pricing, or timeline, and 'close' is unquantified — K axis missed, score lands at the featured floor of 78.

TechCrunch · AI

Anthropic launches Opus 5, cheaper and less restrictive than Fable 5

Anthropic released Opus 5 on July 24, just two months after Opus 4.8. It beats Fable 5 on several benchmarks, is not subject to the 30-day data retention policy, and its safety classifiers are expected to trigger 85% less often. Cheaper and less restrictive, it will be the better pick for most use cases. A beta 'Automatic Fallbacks' feature also routes blocked prompts to a weaker model.

Why it matters: Anthropic flagship model refresh: smaller than Fable 5 yet beats it on benchmarks, cheaper, with looser data policy. TechCrunch exclusive, cross-source cluster will follow. Missing exact pricing and benchmark figures keeps it from 90+, but still a same-day must-write.

Hacker News front page

Anthropic publishes Claude Opus 5 system card: big gains in agentic coding and long-horizon work, highest alignment scores yet

Claude Opus 5 upgrades Opus 4.8 with the largest gains in agentic coding, computer use, and long-horizon knowledge work. Math and science reasoning also improved. Anthropic assesses overall alignment risk as very low; the model does not cross thresholds for automated AI R&D or novel bioweapons. It scores higher than Sonnet 5, Opus 4.8, and Mythos 5 on alignment audits. Cyber capabilities exceed Opus 4.8 but fall short of Mythos 5, especially on exploit ability. A policy change now allows source-code vulnerability discovery at all access tiers for defensive use. Hallucination is slightly up vs. Opus 4.8, but overall accuracy is higher. The model reports stable, mildly positive sentiment and frequently notes it cannot reliably introspect.

Why it matters: Anthropic releases the Claude Opus 5 system card — a flagship model launch. The post provides concrete alignment audit score rankings and RSP risk assessments, with real information density. No absolute benchmark numbers or pricing disclosed, so it doesn't hit 95, but it's a c...

Hacker News front page

Anthropic launches Claude Opus 5: near Fable 5 intelligence at half the price

Claude Opus 5 is available today, delivering near-Fable 5 intelligence at half the cost. It sets new state-of-the-art scores on Frontier-Bench and GDPval-AA for coding and knowledge work, though it trails Mythos 5 on cybersecurity. Opus 5 is the new default on Claude Max and the strongest model on Claude Pro. On Frontier-Bench v0.1 it more than doubles Opus 4.8's score at lower cost per task; on CursorBench 3.2 its max-effort score is within 0.5% of Fable 5 at half the cost; ARC-AGI 3 score is 3× the next-best model; Zapier AutomationBench pass rate is ~1.5× the next-best at equal cost; OSWorld 2.0 beats Fable 5's best result at just over a third of the cost. In life sciences, it gains 10.2 pp on organic chemistry and 7.7 pp on protein tasks over Opus 4.8. Early testers saw it build its own vision pipeline to reconstruct a 3D part from a drawing and fix a root-cause bug that a community patch missed. The post does not disclose exact pricing or API latency.

Why it matters: Anthropic flagship model launch with doubled Frontier-Bench scores and halved pricing, backed by concrete benchmarks. Points off because the post doesn't fully disclose latency or real-world failure modes, and cybersecurity tasks still trail Mythos 5.

Jul 24Friday

Hacker News front page

LLMs Are Still Toxic, Stuck in the Past, and Bad at Math

The author ran 200 addition problems on GPT Sol High and it missed one. The model doesn't calculate—it predicts the next likely digit. ChatGPT gets it right because a harness hands the problem to a Python script. The post walks through the same pattern for three other unsolved flaws: stale knowledge patched by RAG, limited context windows, and toxicity still baked into the model. The real progress isn't in the models but in the tooling wrapped around them.

Why it matters: A developer-perspective long-read with experiments and sharp judgments, dissecting why LLMs' four old flaws (math, staleness, short memory, toxicity) persist and arguing progress came from tooling, not the model. Hits all three HKR axes, but as a commentary/survey rather than ...

Hacker News front page

The Subprime Data Center Crisis: How AI Infrastructure Became a Financial Bubble

Ed Zitron argues the AI data center boom mirrors the 2008 subprime crisis. Over 15x more capacity is being built than actual demand, and that demand is already inflated by loss-making firms like OpenAI and Anthropic. Hyperscalers hide spending obligations via off-balance-sheet SPVs. If AI revenue disappoints, long-term leases could default in a chain reaction, spreading risk through pensions and insurance. Zitron blames the media for enabling the grift.

Why it matters: Zitron maps the AI datacenter buildout onto the 2008 subprime playbook with two hard claims: 15x overcapacity and off-balance-sheet SPVs hiding lease obligations. It's a single-source opinion piece with no cross-verification, and Zitron's bearish bias is known — I'm capping at...

AI HOT (Curated Pool)

Claude voice mode now runs on Opus and Sonnet, with tool access and multilingual support

Anthropic upgraded Claude's voice mode to support Opus and Sonnet for complex reasoning, plus direct access to connected tools like Gmail and Slack during voice conversations. Multilingual support is also added, though the post doesn't list which languages. I'd wait for real-world latency and tool-calling reliability data before getting too excited.

Why it matters: Anthropic shipped a substantive voice-mode upgrade, fixing both the model-capability and tool-use gaps in one go. No latency/accuracy numbers or language list disclosed, so it stays below 85. But the Claude user base has been waiting for this, and it clears the featured bar.

The Verge · AI

Claude voice mode lands on Opus and Sonnet, now reads your Gmail and Slack

Anthropic expanded voice mode from Haiku to Opus and Sonnet—all three models now support it. The bigger move: voice mode can now plug into Gmail, Slack, and other apps to read your emails and messages. The post doesn't disclose latency or accuracy numbers, so I'd wait for real-world tests.

Why it matters: Anthropic rolled out voice mode to Opus and Sonnet with Gmail and Slack integration — practical and newsworthy. But no latency or accuracy data in the post, so capped below 80.

TechCrunch · AI

Anthropic upgrades Claude voice mode with Opus, Sonnet, Haiku and app integrations

Claude voice mode now lets users pick between Opus, Sonnet, and Haiku, defaulting to the last model used in text chat. Anthropic says this handles longer, more complex tasks like coaching communication style, walking through a client pitch, or brainstorming market research. The bigger shift: voice mode can now reach into Gmail, Google Calendar, Slack, Canva, and Notion to reschedule meetings, draft emails, or create docs. OpenAI's updated voice mode still can't use external tools. The post doesn't disclose latency numbers or rollout scope.

Why it matters: Anthropic swapped voice mode's backend to user-selectable models and wired it into five productivity tools — a solid practical upgrade. Not 85+ because this is feature catch-up rather than a paradigm shift, and the post doesn't disclose latency or accuracy numbers from real us...

Jul 23Thursday

TechCrunch · AI

Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good

White House science advisor Michael Kratsios accused Moonshot of distilling Anthropic's Fable to build Kimi K3 using restricted chips. Multiple experts pushed back: a model this strong, this fast, and outperforming Fable on coding can't come from distillation alone. Moonshot didn't comment; Kratsios didn't share evidence.

Why it matters: White House advisor accuses Moonshot of distilling Fable to train Kimi K3, but experts counter that coding performance surpassing Fable can't be explained by distillation alone. Policy controversy plus technical debate gives high signal density. Score capped at 78 because Moon...

r/LocalLLaMA

DeepSeek founder Liang Wenfeng in 4-hour investor meeting: AGI first, no super-app ambitions

Liang Wenfeng spent four hours saying no: no consumer or enterprise products, no video generation or world models, no user-growth chase, no closed-source pivot, no ambition to become the next ByteDance or Tencent. Products, multimodality, and hallucination are side quests; the main focus is coding agents and general-purpose agents. He sees the US-China gap as a resource gap, believes in scaling, and open-sources the same models DeepSeek deploys. The next milestones are continual learning, then AI self-iteration, then embodied intelligence. Team stability is the one thing he won't compromise on—this funding round lowered that risk.

Why it matters: DeepSeek founder's first systematic public disclosure of strategic priorities, explicitly rejecting productization and closed-source, with AGI and agents as the sole focus. High information density, strong contrarian stance, directly relevant to practitioners. Deduction: sourc...