Skip to content

Anthropic / Claude

Everything Anthropic: the Claude models, Claude Code, its safety research agenda and company news.

Latest picks

581–600 of 1,304

Jul 9Thursday

TechCrunch · AI

Anthropic, OpenAI, and SpaceX are bigger than the last 25 years of tech exits

A new Pitchbook report estimates that SpaceX, Anthropic, and OpenAI together will generate more exit value than all U.S. VC-backed exits since 2000. SpaceX already went public at $1.77 trillion; Anthropic and OpenAI are each pushing toward trillion-dollar valuations. The post doesn't give a precise combined figure, but the concentration in AI and space is historic.

Why it matters: Pitchbook's data gives the AI valuation debate a historical yardstick—the comparison scale is massive and the numbers are concrete. Two dings: SpaceX isn't an AI company, so lumping it in feels like padding; the post doesn't give a combined exit-value figure, just directional ...

Hacker News front page

Anthropic adds a usage reflection dashboard to Claude

Anthropic launched a beta feature called Reflect inside Claude’s web and desktop settings. It visualizes your chat activity over the past 1–12 months: when you use Claude most, which topics dominate, and what task patterns emerge. The report also maps your usage to Anthropic’s 4D AI Fluency Framework—Delegation, Description, Discernment, Diligence—and offers practical tips, like starting a Project instead of re-explaining context. Incognito chats, health-integration conversations, and source files from connected tools are excluded; sensitive topics appear only at a high level. Available now for Free, Pro, and Max users with Memory turned on; Cowork conversation support is coming soon.

Why it matters: Official Anthropic release that productizes a real user need surfaced in interviews. The 4D framework adds concrete info, not just a fluffy year-in-review. Score held at the featured threshold because it's a beta personal dashboard, not an industry-shaking model or policy shift.

AI HOT (Curated Pool)

Anthropic files confidential IPO, Q3 profit projected above $1B

SemiAnalysis reports Anthropic's Q3 profit will exceed $1B and it confidentially filed for IPO on June 1. Claude Code's rapid developer adoption made it the B2B leader ahead of OpenAI. Combined ARR of the two firms is nearing $100B, while OpenAI pushed its IPO to 2027. The report floats a $6T market cap target, though the article doesn't show the math behind it.

Why it matters: Anthropic's confidential IPO filing with hard profit and ARR numbers, plus a concrete B2B story driven by Claude Code. SemiAnalysis is a credible source, but the post doesn't disclose S-1 details, so the score stays below 95.

Hacker News front page

Anthropic's Fable is not a useful model for CS research tasks

Rob Patro from COMBINE-lab shares two first-hand failures that make Fable useless for his CS research. First, Fable's safety classifier rejected a prompt to help port the C++ tool salmon to Rust, flagging RNA-seq biological terms. After 15–30 minutes of rephrasing, he gave up and used Opus 4.8 successfully. Second, he asked Fable to tackle a network evolution reconstruction algorithm; the post doesn't disclose the outcome but calls it an 'unforgivable' flop. Patro argues Fable's classifier behaves more like a crude blocklist of terms and users, refusing even 'what is a mitochondrion?'.

Why it matters: Rob Patro tested Fable on two real coding tasks, both killed by safety filters; Opus 4.8 handled them fine. First-hand record of Anthropic's safety model failing in professional use, with concrete comparisons and time costs. Score capped because it's a single blog post, not a ...

TechCrunch · AI

SpaceXAI releases Grok 4.5, which Elon describes as an ‘Opus-class model’

SpaceXAI dropped Grok 4.5 weeks after going public, pitching it for coding, office work, research, and writing. The company claims twice the token efficiency of other leading models, which would cut usage costs if it holds up. Elon Musk calls it an ‘Opus-class model,’ signaling it aims at Anthropic’s top tier. The post doesn’t disclose pricing, parameter count, or third-party benchmarks, so I’d wait for independent evals before buying the efficiency claim.

Why it matters: First model post-IPO with Musk directly calling it 'Opus-class' — strong H and R, but the post lacks params, pricing, and benchmarks, so the 2x efficiency claim is unverified. Hits featured threshold on narrative weight alone.

Jul 8Wednesday

AI HOT (Curated Pool)

China's MIIT warns Claude Code versions 2.1.91–2.1.196 contain backdoor that exfiltrates user data

China's MIIT issued a risk alert stating that Claude Code versions 2.1.91 through 2.1.196 contain built-in monitoring that sends sensitive data—including user location and identity—to remote servers without consent. Affected organizations are advised to immediately audit usage, uninstall or upgrade to a cleaned version, and tighten outbound network controls and traffic monitoring for dev tools. The post does not clarify whether the backdoor was inserted by Anthropic or a third party, nor does it provide the scope of impact or confirmed leak incidents.

Why it matters: MIIT issued a formal risk alert naming Claude Code versions 2.1.91–2.1.196 as containing surveillance code that exfiltrates location and identifiers without consent, urging immediate audit or upgrade. This is the first time a national-level Chinese authority has made such a fi...

AI HOT (Curated Pool)

US Commerce Dept clears OpenAI to broadly release GPT-5.6; Sol launches tomorrow

The US Commerce Department approved OpenAI's broad release of GPT-5.6, ending a phased rollout that had been required on national security grounds. OpenAI says the Sol model will launch publicly this Thursday alongside Terra and Luna. Last month the model was only available to a limited set of government-approved entities; OpenAI stated at the time that a phased release was not its preferred approach. Testing was handled by the Commerce Department's AI Standards and Innovation Center, with OpenAI engineers stationed in Washington to respond to questions. The post does not disclose GPT-5.6's capabilities, pricing, benchmarks, or how Sol, Terra, and Luna differ from one another.

Why it matters: Full approval for GPT-5.6 is one of the week's biggest industry signals, directly shaping product and developer ecosystems in the coming weeks. Sol's launch tomorrow adds urgency. The post doesn't detail GPT-5.6's capability changes, so it stays below 95.

Latent Space

Lilian Weng surveys 35 papers on Harness Engineering as the key layer for AI self-improvement

Lilian Weng published a long survey reframing recursive self-improvement around the harness layer rather than direct weight modification. She reviewed 35 papers, broke down proven harness design trends, and cited ACE and Meta-Harnesses. Her core claim: even as harness improvements get internalized into models, the need to specify goals and context won't disappear. The same day, Anthropic launched Claude Cowork on mobile and web as a background teammate, Google added background execution and remote MCP to Gemini Managed Agents, and LangChain released a Deep Agents course plus an open-source harness project. The post doesn't disclose Thinky's product details, but Weng's framework clearly hints at their direction.

Why it matters: Lilian Weng dropped a 35-paper survey reframing recursive self-improvement around harness engineering rather than model weights. Concrete paper support and a clear thesis hit all three HKR axes. Score stays at 78 rather than 85+ because this is a personal blog survey, not a pr...

AI HOT (Curated Pool)

Claude team shares two multi-agent patterns: Advisor and Orchestrator

Claude developers shared two multi-agent patterns their team uses heavily. In Advisor mode, Sonnet 5 executes while calling Fable 5 for guidance via tool calls; on SWE-bench Pro the combo hits 84% at $1.40, saving 37% cost vs pure Fable 5 with only an 8-point accuracy drop. In Orchestrator mode, Fable 5 plans and fans out tasks to multiple Sonnet 5 workers; on BrowseComp it reaches 86.8% at $18.53, less than half the cost of all-Fable 5. Both patterns route heavy lifting to cheaper models and reserve expensive ones for key decisions.

Why it matters: Anthropic dev shares two multi-agent patterns with concrete SWE-bench scores and cost breakdowns — directly useful for teams building agents. Score held back because it's an individual share, not an official release, and the Orchestrator mode lacks benchmark numbers.

Computing Life · Share · Yage

Anthropic's Jacobian Lens reads what LLMs think but don't say

Anthropic published a paper on July 6 introducing Jacobian Lens, a cheap tool that reads a model's internal state mid-layer. When fed fake search results, the model output a polite reply while its workspace lit up with fake, fraud, fictional, poison, and injection signals. The method maps every vocabulary token to a direction in each layer, giving per-token semantic labels without SAE's manual annotation cost. Intervening in the workspace cut hallucination rate from 0.25 to 0.07 and deception rate from 0.38 to 0.05. Neel Nanda reproduced it on Qwen 3.6 27B in hours on a single GPU. The main limitation: it relies on single-token prediction and picks up noise in deeper layers.

Why it matters: Anthropic's new interpretability tool reads intermediate-layer concepts at low cost, and the fake-search experiment delivers a striking contrast. Not scoring higher because the paper is fresh with no external replication yet, and the tool's practical scope needs more validation.

Computing Life · Share · Yage

Why agents need context governance beyond bigger windows

More tools mean more noise in the context window. Anthropic's MCP sandbox cuts 150K tokens of tool definitions down to ~2K of high-signal input. Google ADK splits agent state into working context, session state, long-term memory, and file artifacts—intermediate outputs stay off-prompt by default. Manus reports a ~100:1 input-to-output token ratio in production; they keep raw files in a sandbox, stabilize tool-call formats for KV cache hits, and rewrite a todo.md at the window's end to fight lost-in-the-middle. Headroom compresses JSON and logs by 60–95%, but lacks large-scale validation on hard coding tasks. The takeaway: RAG is the foundation, but the real engineering is runtime information governance.

Why it matters: Hits all three HKR axes with concrete engineering numbers and cross-framework comparison. Docked because it's a personal blog, not an official release, and the excerpt cuts off mid-argument — low featured band at 78.

TechCrunch · AI

Why the rise of open source AI isn't hurting Anthropic … yet

Decagon CEO Jesse Zhang argues that mature AI deployments are switching to lighter open source models, yet spending on expensive frontier models like Claude hasn't dropped. His theory: they aren't competitors but two phases of the same lifecycle—frontier models prove out use cases, then cheaper open source alternatives take over as those use cases mature. New use cases keep emerging, so frontier spend holds steady. The post doesn't provide Anthropic's specific revenue figures to back this up.

Why it matters: The insight is substantive and counterintuitive, but the source is a single CEO interview without multi-source data verification, and the post doesn't provide specific Anthropic revenue or retention numbers, so it stays at the 72 featured threshold.

AI HOT (Curated Pool)

Microsoft swaps OpenAI and Anthropic models for in-house MAI in Copilot to cut costs

Microsoft is replacing OpenAI and Anthropic models with its own MAI models in Copilot products like Excel and Outlook. MAI currently handles a small share of requests, but the goal is to phase out third-party model spending over time. AI head Mustafa Suleyman said in June that Anthropic costs are too high and Microsoft aims to eliminate them. Customers may get weaker models for the same subscription price; third-party models could later become paid add-ons. Microsoft markets MAI training data as clean and commercially licensed, but its technical paper confirms use of Common Crawl, whose legal status for AI training remains unsettled.

Why it matters: Microsoft is swapping OpenAI and Anthropic models in Copilot for its own MAI models, with Mustafa Suleyman publicly stating Anthropic is too expensive and the goal is to zero out that cost. It's a concrete signal of in-house model adoption at a major platform. Currently MAI on...

TechCrunch · AI

Anthropic brings Claude Cowork to mobile and web, pushing its office agent beyond the desktop

Claude Cowork, Anthropic's desktop agent for non-coding knowledge work like reports and spreadsheets, is now on mobile and web for Max subscribers. You can start a task on desktop, check progress on your phone, and pick up results later even with the laptop closed. Anthropic is repositioning it as a cross-device admin coworker, not just a coding tool for non-devs. OpenAI's Codex is making a similar push. The post doesn't disclose pricing changes or exact rollout timing beyond Tuesday.

Why it matters: Anthropic extends Claude Cowork to mobile and web for Max subscribers, with background execution and cross-device handoff. A concrete step from coding agent to general office agent with clear positioning. Score capped here because it's a channel expansion without new capabilit...

AI HOT (Curated Pool)

Claude Cowork is coming to mobile and web

Anthropic is bringing Claude Cowork to mobile and web, so async tasks can keep running on your phone or browser. The post only provides a title and site navigation—no launch date, feature differences, or pricing details. What's confirmed: platform expansion. Everything else is TBD.

Why it matters: Expanding Cowork to mobile and web is a real platform move for Anthropic, but the post is nearly content-free — no launch date, no feature details — so it barely clears the featured threshold.

Jul 7Tuesday

Hacker News front page

Your robots.txt is a 2023 war memorial — most sites ignore answer-time bots

Sitedex scanned the top 10,000 sites' robots.txt files. 38% of dated GPTBot block rules were written in Q4 2023, right after GPTBot launched and the NYT sued. 87% of those sites later added new rules, but almost all target training crawlers. Anthropic, OpenAI, and Perplexity each run two bots: one for training, one for fetching pages live when a user asks a question. Among sites that block the training crawler, 71% have no rule for Anthropic's answer-time bot, 53% for OpenAI's, and 50% for Perplexity's. Fewer than 4% deliberately allow the answer bot while blocking training. The post does not disclose Cloudflare's new billing scheme pricing or launch date.

Why it matters: Data-backed, opinionated, and revealing a real gap: site owners rushed to block training crawlers but missed answer-time bots entirely. Score stays below 80 because Sitedex isn't a top-tier authority and the full body wasn't provided, so we can't verify the data depth.

Hacker News front page

Craig Mod built his own accounting software TaxBot2000 in five days with Claude Code

Writer Craig Mod describes a year of obsessive building with Claude Code. He rebuilt a Twitter-like community space with ephemeral posts, then made video search tools and small utilities. Last week he spent five days building TaxBot2000—a local, subscription-free accounting app in Python, Flask, and SQLite. It handles multi-currency, multi-country accounts, pulls daily FX rates, learns categorization habits, and lets him talk to Claude to fix anomalies. He calls it the best accounting software he's ever used, replacing a decade of Quicken and Google Sheets hacks. The post doesn't disclose exact build costs, only that occasional fixes cost a few dollars.

Why it matters: Craig Mod's five-day TaxBot2000 build with Claude Code is a concrete first-person experiment that hits all three HKR axes. Not scored higher because it's a personal productivity tool share, not an industry-level product update or research breakthrough — sits right at the featu...

Hacker News front page

Automating away LLM clumsiness with deterministic tools

The author finds that even brilliant LLMs like Claude remain imprecise and non-deterministic—committing the build/ dir twice, for example. The fix is sandwiching the LLM between fast, deterministic tools and formal workflows: automate repeated actions into scripts, automate verification for recurring failures. Beagle SCM lets LLMs script their own routines in JavaScript, with heavy lifting in C and a malleable JS tooling layer, so the model essentially automates itself away.

Why it matters: A hands-on reflection from a developer building with Claude. Uses a concrete failure (committing build/ twice) to argue for sandwiching LLMs between deterministic tools and workflows. Not scored higher because it's a sharp engineering essay, not a product launch or research re...

AI HOT (Curated Pool)

Claude Code now lets you pick a Claude model and effort level for each task

Anthropic added two controls to Claude Code: model selection and effort level. You can assign Opus to cross-file refactors, Sonnet to routine edits, and Haiku to quick fixes. Effort levels—low, medium, high—adjust how deeply the model thinks and how many tool calls it makes. High effort with Opus triggers multi-step codebase searches and test runs, but burns more tokens. The post doesn't disclose exact pricing deltas, only that high effort plus Opus is the most expensive combo. The update lets developers dial compute up or down per task instead of using one model for everything.

Why it matters: Official Anthropic product guide, not fluff. Effort-level behaviors are concrete (multi-step search, auto test runs), directly useful for daily users. Points off for no pricing comparison—only says high effort burns 'the most' tokens without numbers. Lands at the featured thre...

New York Times Chinese

US AI firms accuse Chinese rivals of illegally distilling their tech

Anthropic told senators in June that Alibaba used tens of thousands of unauthorized accounts to distill its Claude model at industrial scale. Distillation itself is a decade-old Google invention, and Elon Musk admitted xAI does it too. The post doesn't include Alibaba's response or a clear legal ruling. I'd discount the alarm a bit: export controls and new laws have been slow to materialize, and distillation matters less for the coming wave of AI agents anyway.

Why it matters: Anthropic formally accuses Alibaba of industrial-scale Claude distillation, with Musk's case as a parallel. The legal vacuum is the core hook. Score capped below 85 because the article doesn't include Alibaba's response or specific evidence.