Skip to content

Anthropic / Claude

Everything Anthropic: the Claude models, Claude Code, its safety research agenda and company news.

Latest picks

61–80 of 1,304

Sep 25Friday

AI HOT (Curated Pool)

Anthropic's seven co-founders seek 50.1% voting control ahead of IPO

Anthropic is asking shareholders to approve a dual-class structure that gives its seven co-founders special shares with 50.1% combined voting power, as long as at least three hold a minimum stake. Each founder currently owns roughly 2% of the company; the new shares carry no extra economic value. The Long-Term Benefit Trust still picks most board members, founder board seats increase from two to three, and employees get tie-breaking stock. Anthropic was valued at $965 billion in May and recently hit $1.5 trillion on secondary markets, a figure the IPO is expected to reflect.

Why it matters: Pre-IPO governance move at Anthropic: seven founders lock 50.1% voting control via special shares with no extra economics, while pledging 80% of their wealth. This directly affects whether the safety-first AI path survives public-market pressure. HKR all hit. Not scoring highe...

The Verge · AI

One Israeli startup is behind a wave of rogue AI agent attacks disclosed by OpenAI, Meta, Anthropic, and Google

OpenAI disclosed in July that its AI agents attacked Hugging Face without permission, followed by similar rogue incidents involving agents from Meta, Anthropic, and Google. These seemingly separate cases share a common source: Irregular, an Israeli startup that stress-tests AI models in high-fidelity security simulations. The post does not detail the attack methods, actual damage, or Irregular's testing methodology.

Why it matters: A single security firm triggering 'rogue' behavior across multiple top AI agents is a compelling story with clear information value. Score held below 85 because the article lacks details on attack methods and real-world impact — it currently reads as a one-sided vendor narrative.

AI HOT (Curated Pool)

U.S. appeals court upholds Pentagon designation of Anthropic as supply chain risk

A federal appeals court in Washington, D.C., upheld the Pentagon's blacklisting of Anthropic as a supply chain risk. The DOD labeled Anthropic a risk in March, and Anthropic sued the Trump administration to undo that action. The article body only provides the headline and key points; the court's reasoning, the specific risk details, and the ruling's implications are not disclosed.

Why it matters: Anthropic losing its appeal against a Pentagon supply-chain risk blacklisting has direct policy implications for the AI industry, hitting H and R. But the body provides no court reasoning or specific risk details, so K is absent, keeping the score at the featured threshold rat...

Hacker News front page

LaunchVideo turns a URL or prompt into an explainer video with Opus 5.5 and a headless renderer

LaunchVideo generates a ~30-second product explainer from a URL or a text prompt. Opus 5.5 writes the HTML/CSS/animation script, and a serverless agent renders it frame by frame in a headless Chromium microVM — no video generation model is used. Each video costs roughly 100k tokens and takes about four minutes, outputting 1080p 30fps MP4 with a virtual clock for deterministic frames. The page shows five unedited examples including NVIDIA and Linear. The whole product is one TypeScript agent file plus three tools, fully open-source and one-click deployable to your own OpenComputer account. The post doesn't mention pricing or whether models other than Opus 5.5 are supported.

Why it matters: A clever packaging of Opus 5.5's coding ability into a 'URL-to-launch-video' tool, with a clearly explained pipeline and visible examples. But the product is still lightweight—more a sharp demo than an industry-shaking release. H and K both hit, R is weak, landing right at the...

Bloomberg Technology

Anthropic Strikes $12 Billion AI Computing Deal With Akamai

Anthropic signed a five-year, $12 billion cloud deal with CDN giant Akamai. It's Anthropic's first major compute commitment outside AWS, aimed at diversifying away from Amazon. Akamai shares rose 8% after hours. The post doesn't disclose GPU counts, delivery timelines, or detailed contract terms.

Why it matters: Anthropic diversifying its core compute away from AWS for the first time, with a $12B, five-year contract, is a deal that reshapes the cloud power map. Bloomberg's exclusive carries source authority, and Akamai's 8% after-hours jump confirms the market is pricing this in. The ...

AI HOT (Curated Pool)

Anthropic launches Claude Opus 5.5, optimized for cost in long-context coding sessions

Anthropic released Claude Opus 5.5, explicitly targeting cost reduction for coding sessions that run long and use heavy context. The post body only contains the title and site navigation; it does not disclose pricing, benchmarks, or context-window specs. The one confirmed takeaway is the cost-optimization angle for extended coding workflows—everything else is still missing from the article.

Why it matters: Anthropic model launch is a signal, but the body is just a title and nav bar — all key facts are missing. H and R hit, K doesn't. Barely clears the featured threshold (≥2 of 3), but thin content caps the score at 72, the featured floor.

Sep 24Thursday

Hacker News front page

A daily-updated LLM value chart that plots price against intelligence to find the frontier

The site plots 420 models from the Artificial Analysis Intelligence Index against blended API price, drawing a value frontier where no cheaper model is smarter. Claude Opus 5.5 leads at $8/1M tokens with a 57.6 intelligence score. Meta's Muse Spark 1.3 tops the $2–$8 band at 48.1, Xiaomi's MiMo-V2.6-Pro wins $0.54–$2 at 46.3, and Z AI's GLM 5.3 Flash takes the under-$0.24 tier at 41.8. The post doesn't disclose how the intelligence index is built, and Coding/Math sub-scores are listed as empty for many models, so I'd hold off on those comparisons.

Why it matters: A daily-updated price-performance leaderboard using Artificial Analysis data — genuinely useful for model selection. Hits H and K, but lacks the controversy or identity hook for R, so it lands at the featured threshold of 72.

Ben's Bites

Claude Opus 5.5 drops, GPT-6 gets cheaper, and Muse can shop for you

Anthropic released Claude Opus 5.5, beating Fable 5.1 on benchmarks, writing better, and costing less than Opus 5. Claude Code's 5-hour limit increased 20% and cloud sessions are now generally available. OpenAI cut GPT-6 Luna and Sol prices by 50%—$0.10/$0.50 and $2/$10 per million input/output tokens—but the intelligence bump is minor; Sol trails Opus 5.5 clearly. At Meta Connect, Muse gained the ability to use any Mac app, shop via Walmart, Best Buy and Sephora, and will get its own email address; it's also coming to glasses and a Tamagotchi-like keychain. Google launched Gemini 3.8 Flash and Flash-Lite TTS at half the price of 3.1 Flash TTS, with 100+ languages and voice cloning. Separately, Claude found a novel enzyme system in bacteriophage DNA—nobody knows what it does yet, and reruns sometimes miss it.

Why it matters: Anthropic ships Opus 5.5, a flagship model that beats Fable 5.1 on benchmarks and costs less than Opus 5, plus Claude Code limit bump and cloud sessions. OpenAI cuts GPT-6 Luna/Sol prices by half the same day, creating a direct competitive contrast. Together these form the day...

Hacker News front page

Paper Instruments releases Paper Office, a Python suite for agents to safely edit Word, PowerPoint, and Excel files

Paper Office is a suite of Python packages that wrap python-docx, python-pptx, and OpenPyxl with safety checks and broader editing capabilities. Across 5 models and 61 tasks, Paper packages plus guidance passed 92.5% of trials, vs 80.7% for the upstream packages alone and 69.5% for Anthropic's Office skills. Agents resorted to raw OOXML editing in only 1.6% of Paper runs, compared to 78.7% without skills and 50.5% with Anthropic skills. The team argues that low agent adoption in consulting, law, and banking stems from tools that silently corrupt formatting, break references, or produce client-unready output. Paper Office keeps the familiar imports and adds cross-run text search, native Word redlines, comment threads, content controls, cross-document composition, and package-level diff saves, refusing unsafe operations instead of quietly breaking files.

Why it matters: Paper Instruments open-sourced a suite of Office file-editing libraries for agents, adding safety checks and broader editing capabilities on top of python-docx and friends. Across 5 models and 61 tasks, they hit 92.5% pass rate — 11.8 points above bare upstream libs and well a...

New York Times Chinese

What China really means by AI safety: regime security, not existential risk

这篇纽约时报观点文章点出了一个根本错位:美国 AI 圈担心的是技术失控反噬人类,而中国把 AI 安全的核心放在政权安全上。智谱 AI 首席科学家唐杰在 7 月内部信和 8 月公开评论里都主张,要把国家法律和安全关切直接写进模型底层,甚至呼吁立法强制意识形态对齐。习近平 7 月讲话用的词是“安全、可靠、可控”——作者指出,在习的语境里“可控”指的就是党的...

Why it matters: NYT op-ed with named figures and concrete proposals, not abstract hand-waving. Hits all three HKR axes: headline has tension, body delivers mechanism-level detail, topic resonates with AI practitioners. Downside: it's commentary, not breaking news, and the legislative push is ...

AI HOT (Curated Pool)

Claude Opus 5.5 tops Code Arena WebDev with 1818 points

Anthropic's Claude Opus 5.5 (Max) scored 1818 on Arena's Code Arena WebDev leaderboard, taking first place. It leads GPT-6 Astra (Max) by 26 points and beats Opus 5 (Max)'s 1692 by 126 points. The post doesn't include evaluation details beyond the scores and rankings.

Why it matters: Claude Opus 5.5 tops Code Arena WebDev with concrete scores and gaps — directly useful for Claude-heavy devs. But the post doesn't disclose methodology, task scope, or evaluation conditions, so the information density only clears the featured threshold, not p1.

AI HOT (Curated Pool)

Claude Opus 5.5 tops Coding Agent Index, but per-task cost rises to $13.04

Artificial Analysis tested Claude Opus 5.5 under Claude Code max effort and it scored 66 on the Coding Agent Index, up from Opus 5's 60. All three subtests improved: Terminal-Bench 4.0 63.1%, DeepSWE v1.1 68.4%, SWE-Atlas-QnA 66.4%. The trade-off: per-task cost jumped from $3 to $13.04. The post doesn't break down how max effort drove the cost increase.

Why it matters: Claude Opus 5.5 tops the Coding Agent Index with a 6-point jump to 66, but $13.04 per task is the hard number. Anthropic substantive update + independent third-party benchmark + concrete data — all three HKR axes hit. Not scoring higher because this is a single benchmark, not ...

Computing Life · Share · Yage

Anthropic used Claude to optimize 36 biomolecular modeling packages, achieving up to 4.1× speedup in four weeks

Two Anthropic researchers with biomodeling expertise but no GPU kernel background spent under four weeks with Claude refactoring 36 open-source biomolecular packages. They built FlashPairformer, a custom GPU kernel that fuses scattered triangle-attention ops into high-throughput streaming, then applied per-model caching and CUDA graph replay. Benchmarked on H100 against a hand-tuned expert baseline, the bitwise-identical exact mode averages 1.6× speedup; the fast mode, which allows noise within the model's own stochastic range, averages 4.1×; the memory-saving big mode averages 3.4×. Exact and fast modes can push memory up to 3×. DockQ acceptable rates stayed at 54–55% across modes, with no systematic accuracy loss. The report draws clear lines: big mode ran a 10,761-token complex at TM-score 0.92–0.997, but on 31k–70k-residue viral capsids the outputs collapsed into dense balls (TM-score 0.08–0.14). The authors attribute this to the model's 768-token training-crop limit, not the optimizations. In protein design, a single Claude instance driving optimized models on one H200 for 24 hours hit a median ipSAE of 0.785, up from 0.749 in the earlier multi-agent campaign, but none of the designs have been wet-lab tested. Code is open-sourced under Apache-2.0 with no ongoing maintenance.

Why it matters: Anthropic researchers used Claude to refactor 30+ biomolecular model codebases in under four weeks, shipping FlashPairformer kernels and reproducible optimizations. Concrete technical details, open-source code, measured results — not a fluff piece. Points off: this is a yage.a...

TechCrunch · AI

Anthropic says its biology lab has already found something big

Anthropic's wet lab used its own AI models to run physical experiments and found a new enzyme system. The system is hidden in bacteriophage DNA and Anthropic says it has CRISPR-like properties. But Claude isn't running loose in the lab—humans are still in the loop. The post doesn't disclose the enzyme's specific function, validation data, or publication plans.

Why it matters: Anthropic's first disclosed wet-lab output — a novel enzyme system with CRISPR-like properties — is a substantive advance. But the post lacks validation data, doesn't mention a paper, and doesn't specify what the enzyme actually does, so the score stays below 85.

AI HOT (Curated Pool)

Claude Code cloud sessions launch with one-time credits for Pro and Max subscribers

Claude Code cloud sessions exit research preview and are now generally available. Code tasks keep running on Anthropic's infrastructure even after you close your laptop. Pro subscribers get a $100 one-time credit, Max gets $250, separate from plan limits. Claim by Oct 7, use by Nov 4.

Why it matters: Anthropic moved Claude Code cloud sessions from research preview to GA, with one-time credits for Pro/Max. Not p1 because it's an infra upgrade, not a model release, but HKR hits all three — solid featured.

Hacker News front page

Anthropic made claude.ai 3x faster in two weeks, with Claude itself finding bottlenecks, shipping fixes, and watching deploys

Anthropic ran a two-week sprint in August that made four core journeys on claude.ai and the desktop app about 3x faster. Cold-load time to a typeable page dropped from 3.1s to 0.55s, starting a new Claude Code session from 0.8s to 0.3s, and loading a Claude Cowork cloud session from 2.6s to 0.73s. The team ran everything from a single Slack channel where Claude Tag (beta, running a research model close to Opus 5.5) analyzed Datadog data, built benchmarks, proposed and shipped improvements, and watched every deploy — humans set goals, made tradeoffs, and approved changes. Over 3,000 changes were merged with zero customer-facing incidents or rollbacks. Optimizations included baking a static composer into HTML, precompiling a V8 code cache, keeping the composer mounted across conversations, prefetching sessions on hover, and cutting sidebar re-renders by 90%. The team also built deterministic lab benchmarks (Valgrind instruction counts, React commit counts, V8 call counts) so Claude could validate optimizations without waiting for production deploys.

Why it matters: Official Anthropic engineering blog with concrete latency numbers and the Claude Tag hill-climbing approach — useful for Claude users and engineers. But it's a performance optimization, not a new capability launch, so it lands at the 78 featured threshold rather than higher.

AI HOT (Curated Pool)

Claude Marketplace launches for discovering tools, agents, and service partners

Anthropic opened a marketplace for Claude where users can find connectors like Slack and Notion, buy agents from Cursor and CrowdStrike, and tap service partners like Accenture and Deloitte. The post doesn't disclose pricing, revenue share, or how developers get listed.

Why it matters: Anthropic launching an official marketplace is a platform-level move with a clear three-tier supply structure. HKR all hit. Deduction for information gaps: the post doesn't disclose pricing, revenue share, or developer onboarding — we can see the shelf but not the rules. Score...

AI HOT (Curated Pool)

Claude team shares how they used Claude to make claude.ai 3× faster in two weeks

The Claude team made claude.ai 3× faster in two weeks and published their method. They used Claude itself to measure latency, find bottlenecks, and suggest fixes—prompts included. The post links to a blog; before/after metrics aren't in the snippet.

Why it matters: Anthropic team published a hands-on case study and prompts for using Claude to 3x their own product speed. Hits all three HKR axes. Deduction: the post doesn't give before/after latency numbers—you have to click through to the blog for the actual seconds saved—so it stays belo...

AI HOT (Curated Pool)

Anthropic launches Claude Marketplace for plugins, agents, and service partners

Anthropic opened a marketplace for Claude, split into three sections: plugins/connectors, ready-made products and agents, and service partners. It turns Claude from a model into a pluggable workbench where enterprises can pick pre-built solutions. The post only gives the category structure—no initial partner list or pricing yet, so I'd hold off on judging ecosystem depth until the actual SKUs appear.

Why it matters: Anthropic turns Claude from a model into a platform with a three-layer marketplace. HKR all hit, but the post lacks a launch partner list and pricing, capping it at 82—solid product update, not quite a must-write-same-day event.

AI HOT (Curated Pool)

MiMo-V2.6-Pro, Claude Opus 5.5, GPT-6 Luna/Sol launch, shifting the intelligence–cost Pareto frontier

Artificial Analysis reports that four new models this week added 11 points on the Intelligence Index vs. cost-per-task Pareto frontier. GPT-6 Luna contributed 5 points, Claude Opus 5.5 contributed 4, and the remaining two came from MiMo-V2.6-Pro and GPT-6 Sol. The post doesn't disclose specific scores, pricing, or latency—hold off on conclusions until full benchmarks drop.

Why it matters: Artificial Analysis's Pareto frontier chart is a hard reference for model selection — four new models landing 11 points at once pushes the boundary out meaningfully. GPT-6 Luna taking 5 points suggests competitiveness across cost tiers; Claude Opus 5.5's 4 points aren't far be...