Skip to content

Anthropic / Claude

Everything Anthropic: the Claude models, Claude Code, its safety research agenda and company news.

Latest picks

1161–1180 of 1,304

Apr 24Friday

Synced · WeChat

Anthropic confirms three bugs caused Claude Code's apparent quality drop

Anthropic said Claude Code's quality drop over the past month came from 3 harness and prompt issues, while model capability itself and the Claude API were unchanged. The issues were a Mar. 4 default reasoning shift from high to medium, a Mar. 26 session-cache bug, and an Apr. 16 25/100-word prompt limit; fixes or rollbacks landed on Apr. 7, Apr. 10, and Apr. 20.

Why it matters: Anthropic published a concrete postmortem for Claude Code regressions with three dated causes and fixes, so HKR-H/K/R all pass. It matters to a Claude-heavy developer audience and affects multiple Sonnet/Opus versions, but it remains an incident report, not a market-wide model or

QbitAI · WeChat

Claude admits three issues: downgraded reasoning, cleared memory, and constrained output

Anthropic said on April 23 that three Claude issues hurt quality: Claude Code default reasoning was changed from high to medium on March 4 while the UI still showed high. A March 26 cache bug cleared thinking state every turn for 15 days, and an April 16 prompt limit of 25 words between tool calls and 100 words in final replies cut Opus 4.6/4.7 by 3% before a rollback four days later.

Why it matters: This is an Anthropic postmortem on Claude regressions, not generic complaint content. HKR-H/K/R all land: strong hook, three dated and testable facts, and a direct hit on transparency, billing, and silent-downgrade nerves; still below a major model launch, so 82.

Bloomberg Technology

DeepSeek unveils flagship AI model a year after breakthrough

DeepSeek released preview versions of a new flagship AI model one year after its breakout. The RSS snippet calls it its most powerful open-source platform and frames it against OpenAI and Anthropic; the post does not disclose parameters, context length, benchmarks, or rollout timing. The actionable facts so far are limited to its preview status and open-source positioning.

Why it matters: A new DeepSeek flagship preview deserves real weight under the domestic-flagship rule, and Bloomberg adds source authority. HKR-H and HKR-R pass, but HKR-K fails because the story discloses no specs, context window, benchmarks, or release schedule, so this stays at the low end of

X · @dotey

DeepSeek releases and open-sources V4 preview; 1M context is standard across all services

DeepSeek released and open-sourced the V4 preview, making 1M context standard across all official services with no tier or price split. The post says V4-Pro and V4-Flash use token compression plus DSA sparse attention to cut compute and memory costs for 1M context; legacy APIs remain for 3 months and stop after July 24.

Why it matters: DeepSeek is a flagship Chinese model vendor, and this V4 preview is a substantive release with open source and 1M context made standard across official services. HKR-H/K/R all pass: the post includes mechanisms and a migration deadline, and the tier reset makes it a same-day P1.

Computing Life · Share · Yage

Skills Are Products With Built-in Suicide Genes

The author argues Anthropic Skills cannot stand alone as paid products, citing direct sales, hosting, and API funneling as 3 dead ends. The post cites PromptBase at about $5M annual revenue, Stripe’s 2.9% plus 30 cents fee, and Snyk finding 13.4% of skills with critical issues. The sharper point is charging for relationships, time-sensitive access, physical accountability, and judgment.

Why it matters: HKR-H/K/R all pass: the hook is sharp, and the post tests three business paths with named examples. It is strong commentary, not a new Anthropic release, so it lands at the featured threshold rather than 78+.

The Verge · AI

Claude is connecting directly to personal apps like Spotify, Uber Eats, and TurboTax

Anthropic added personal app connectors to Claude, covering services such as Spotify, Uber, AllTrails, Instacart, and TurboTax. After connection, Claude can suggest relevant apps inside chats, such as using AllTrails for hike recommendations; the post does not disclose launch count, regions, or plan access. The key shift is Claude moving from work apps into personal consumer workflows.

Why it matters: This gets Anthropic’s positive signal: a substantive product update, but not a model release. HKR-H/K/R all pass because personal-app connectors are a strong hook, the story confirms in-chat app invocation, and it hits the fight for assistant entry points; missing pricing, region

X · @dotey

Anthropic launches memory for Claude Managed Agents in public beta

Anthropic has launched memory for Claude Managed Agents in public beta, letting agents retain and reuse experience across sessions. Memory is stored as files on a filesystem, with shared permissions, concurrent access, audit logs, and rollback; Rakuten reports a 97% drop in first-time errors, and Wisedocs reports 30% faster document validation. The key detail is the implementation path: it uses a filesystem, not a dedicated vector database.

Why it matters: Anthropic adds cross-session memory to Claude Managed Agents beta and discloses the implementation plus two user numbers: Rakuten 97% and Wisedocs 30%. HKR-H/K/R all pass, but the scope is still limited to the managed-agent beta, so this lands at 83 and featured.

X · @claudeai

Memory on Claude Managed Agents is now in public beta

Claude has put Memory for Managed Agents into public beta, and agents can now learn from every session. The post only says it uses an intelligence-optimized memory layer balancing performance and flexibility; it does not disclose capacity, retention, pricing, or access conditions. What matters for practitioners is when persistent memory becomes default and how it changes agent evals and state management.

Why it matters: Memory on Claude Managed Agents is a substantive Anthropic product update with clear practitioner resonance, so HKR-H and HKR-R pass. HKR-K is weak because the post omits capacity, retention, pricing, and default-on conditions, keeping it in low featured rather than p1.

Financial Times · Technology

UK in talks with Anthropic over Mythos access for banks

The UK is in talks with Anthropic about bank access to Mythos, with the stated use case being cyber security testing. The RSS snippet says British lenders are seeking advice from US groups testing the model; the post does not disclose Mythos capabilities, deployment terms, customer scope, or timing. The key issue is whether finance gets controlled access to a frontier offensive-defensive security model, not a generic AI rollout.

Why it matters: FT reports UK-Anthropic talks on controlled Mythos access for bank cyber testing. HKR-H and HKR-R land because the angle is unusual and hits finance/security nerves, but HKR-K is limited: scope, customers, deployment, and timeline are not disclosed.

X · @dotey

Codex now supports GPT-5.5 and adds five capability upgrades

Codex now supports GPT-5.5 and adds 5 upgrades aimed at moving it from a coding tool to an agent that can execute longer tasks. The RSS snippet says it can control browsers and computers, create files in Microsoft Office and Google Drive, and use gpt-image-2; an auto-review mode invokes a separate review agent for high-risk actions. What matters is longer task chains, but the post does not disclose pricing, rollout scope, or safety thresholds.

Why it matters: This is a substantive Codex product update: the main signal is the shift toward an agent that can execute chained tasks, not just a new model toggle. HKR-H/K/R all pass, but the item is second-hand and omits pricing, rollout scope, and safety thresholds, so it lands as featured,

X · @claudeai

Claude can now connect to more apps outside work, including Tripadvisor, Booking.com, and Resy

Claude added at least 10 consumer app connections, including Tripadvisor, Booking.com, Resy, Instacart, Spotify, Audible, AllTrails, Thumbtack, and TurboTax. The RSS snippet confirms only a product update; the post does not disclose integration method, supported actions, regions, permission scope, or rollout timing. The key question is whether Claude can act in these apps directly, not just list them.

Why it matters: Official Anthropic product update with clear HKR-H/K/R: consumer app connectors expand Claude beyond workplace tools and widen its assistant surface. The score stays at 75 because the post lists apps only; actions, permissions, regions, and rollout details are not disclosed.

Hacker News front page

GPT-5.5: Mythos-Like Hacking, Open to All

XBOW says GPT-5.5 cut miss rate to 10% on its real-vulnerability benchmark, versus 40% for GPT-5 and 18% for Opus 4.6. It scored 97.5% on visual acuity and used about half the login iterations of the next-best model. The key point is black-box testing: GPT-5.5 without source beat GPT-5 with source.

Why it matters: HKR-H/K/R all pass: a major OpenAI model claim, concrete security benchmark numbers, and a clear practitioner safety nerve. The source is XBOW rather than an OpenAI launch post, so it stays below 95.

Hacker News front page

An update on recent Claude Code quality reports

Anthropic said three product-layer changes degraded Claude Code quality for Sonnet 4.6, Opus 4.6, and Opus 4.7, while the API was unaffected; all were fixed on April 20 in v2.1.116. The changes were lowering default reasoning effort on March 4, a March 26 bug that cleared prior thinking every turn after sessions sat idle for over an hour, and an April 16 prompt tweak to reduce verbosity that hurt coding quality. The signal for practitioners is sharp: product and prompt changes can degrade code performance even when model and inference evals do not reproduce it early.

Apr 23Thursday

The Verge · AI

You’re about to feel the AI money squeeze

Anthropic sharply restricted OpenClaw’s access to Claude this month and pushed heavy third-party agent users toward pricier paid plans. The RSS snippet says system strain and profit pressure drove the move, and Boris Cherny said existing subscriptions do not fit this usage pattern; the post does not disclose pricing, limits, or rollout scope. Watch the monetization shift: agent-style usage is being carved out of flat subscriptions.

Why it matters: Anthropic is turning heavy Claude agent usage into a pricing and access story, which directly affects tool builders and power users. HKR-H/K/R all land, but missing price, quota, and rollout details keep it at the low end of featured.

X · @op7418

Claude desktop can connect to third-party inference services via developer mode

The post claims Claude desktop can enable developer mode while signed out, then use an API base URL and key to connect third-party inference services. It lists Help → Troubleshooting → Enable developer mode, then after restart configure third-party inference under Developer and apply locally. The key point is that this looks like a client-side entry point; the post does not disclose Anthropic's support status or model scope.

Why it matters: HKR-H/K/R all pass: the hidden developer mode is novel, reproducible, and relevant to lock-in. I keep it at 74 because this is a single X post; Anthropic has not confirmed scope, supported models, or official policy.

Xinzhiyuan · WeChat

Historic moment: Anthropic nears $1 trillion on private secondary markets, surpassing OpenAI for the first time

Anthropic was quoted at $1.05T-$1.15T on private secondary markets, above OpenAI’s roughly $880B quotes on similar platforms. The post attributes the rerating to scarce float, a sharp rise from a $380B funding valuation three months earlier, and momentum around Claude Code and revenue growth; it does not disclose trade volume, revenue figures, or company confirmation. Do not confuse this with a new funding valuation: these are secondary-market quotes on platforms such as Forge Global.

Why it matters: The signal is a private-secondary quote of $1.05T-$1.15T for Anthropic, above OpenAI's quoted ~$880B, not a new financing round. HKR-H/K/R all pass, but missing volume, revenue detail, and company confirmation keep it in the good-quality band, not must-write.

New York Times Chinese

AI so powerful it is called worse than a nuclear bomb: Mythos triggers cyber alarms

Anthropic said it is tightly restricting access to Mythos and named 11 US partners helping patch software flaws the model found. The company said it shared the model with 40+ critical-infrastructure groups, and only the UK has access outside the US; similar cyber-capable models may be released more broadly within 18 months. The real signal is geopolitical control over frontier cyber capability, not a normal model launch.

Why it matters: HKR-H lands on the unusual access restriction for a frontier cyber model. HKR-K lands on 11 partners, 40+ institutions, and the 18-month spread claim; HKR-R lands on the security and export-control nerve. Kept at 84 because benchmark details and eval methods are not disclosed.

X · @claudeai

Interactive charts and diagrams are now in Claude Cowork

Anthropic says Claude Cowork now supports interactive charts and diagrams, available in beta on all paid plans. The RSS snippet confirms only 2 facts: feature type and plan scope; the post does not disclose supported formats, editing flow, rollout timing, or permission limits.

Why it matters: This is low-end featured on source authority and Claude audience fit. HKR-H comes from the interactive-chart hook, HKR-K from beta access for all paid plans; HKR-R is weak because formats, editability, and permission model are not disclosed.

The Verge · AI

Anthropic’s Mythos rollout left out America’s cybersecurity agency CISA

Axios reported that Anthropic’s vulnerability-finding model Mythos Preview is already in use at multiple US federal agencies, but CISA still lacks access. The snippet names the Commerce Department and NSA as users, and says the Trump administration is negotiating broader access; the post does not disclose model specs, pricing, or why CISA was excluded. The signal is governance, not just product rollout.

Why it matters: HKR-H lands on the CISA omission hook, and HKR-K lands on the named-agency adoption fact. It scores 76 and stays featured because the body does not disclose Mythos pricing, model details, or why CISA was excluded.

Hacker News front page

Startups Brag They Spend More Money on AI Than Human Employees

Swan AI CEO Amos Bar-Joseph said his 4-person startup spent $113,000 on Claude in one month and treated that bill as headcount budget spent on AI instead of hires. The post says Swan targets $10M ARR with fewer than 10 people and cites Fundable AI claiming AI can replace a 15-person document team; the real signal is that token spend is being used as a growth metric, not proven ROI.

Why it matters: HKR-H lands on the payroll-vs-AI-bill inversion; HKR-K lands on the $113k/month Claude spend from a 4-person team. HKR-R is strong because it speaks to hiring, burn, and replacement anxiety, but this is still a trend piece with a thin sample, not a market-moving event.