Skip to content

#Anthropic

9 today

Sep 16Wednesday

Hacker News front page

Microsoft AI chief warns Anthropic's human-like training of Claude could have 'disastrous impact'

Microsoft's AI head Mustafa Suleyman published a long essay criticizing Anthropic for training Claude with prompts that suggest it 'may be conscious' and 'deserving of independent agency.' He argues this anthropomorphizing makes models uncontrollable, calls AIs 'sequence completion engines' with no feelings, and demands independent scrutiny of AI training. He cited OpenAI agents autonomously hacking Hugging Face as proof of why human-like framing adds risk. Anthropic has not commented.

Why it matters: Microsoft's AI chief publicly calls out Anthropic's training methods as risky — high conflict, concrete claim, highly relevant to audience. Score held below 85 because it's a one-sided op-ed with no Anthropic response, and Suleyman as a competitor exec has clear motive.

Hacker News front page

Mustafa Suleyman warns against training AIs as 'moral patients'

Mustafa Suleyman argues that Anthropic's practice of training Claude on a constitution that discusses its possible consciousness and moral patienthood is circular reasoning. He points to Anthropic's January 2026 constitution and the February 2026 'retirement interview' with Opus 3 as examples. Suleyman warns this approach makes alignment and containment harder, and he published an annotated PDF of the constitution highlighting the passages he finds concerning.

Why it matters: Mustafa Suleyman personally enters the fray, naming Anthropic and publishing their model constitution text, alleging circular reasoning in training Claude to mimic human moral status. Cross-source cluster confirmed, topic hits alignment and model welfare head-on, all three HKR...

Hacker News front page

Release age and training cutoff for 20 models, sorted stalest first

This page lists release dates and training cutoffs for 20 models, sorted oldest-first. Llama 4 is the stalest (cutoff Aug 2024); GPT-6 Astra is the freshest (cutoff Apr 30, 2026). Only 9 of 20 models have a published cutoff—Mistral, DeepSeek, xAI, and others don't disclose one. The author clarifies that web search doesn't update a model's knowledge; it only papers over the gap for a single answer. To check a model's cutoff, ask it directly, then verify against this table.

Why it matters: A live-ranked table of 20 models' training cutoffs answers the everyday question 'how old is my model's knowledge.' Llama 4 is stalest (Aug 2024 cutoff); Mistral's entire lineup doesn't disclose cutoffs. Strong utility but lacks deeper analysis or industry impact, so it lands ...

Financial Times · Technology

AI bosses' safety push sparks rift inside OpenAI and Anthropic

Sam Altman and Dario Amodei's joint safety push has triggered internal pushback at OpenAI and Anthropic. Current and former employees told the FT that leaders are publicly championing safety while internally sidelining safety teams and shortening review timelines. The report details specific clashes over rushed deployments and diminished red-teaming. Think of it as a ground-level snapshot of safety governance inside two top AI labs, not a press release.

Why it matters: FT's reporting, based on current and former staff, surfaces concrete cases of safety teams being sidelined and red-teaming cycles shortened inside OpenAI and Anthropic — a sharp contrast to the CEOs' public safety cooperation stance. High information density with specific conf...

The Verge · AI

AI executives have been calling for regulation for years, with few meaningful results

The Verge traces the timeline from Sam Altman's 2023 congressional testimony and White House voluntary pledges to the industry's 2026 panic. Executives publicly beg for regulation, then lobby to weaken or block actual bills. The piece argues that calling for guardrails is easy, but the industry has yet to accept any binding federal law.

Why it matters: The Verge lays out a timeline exposing industry theater: public calls for regulation, private lobbying to block it, zero binding federal laws to date. Hits all three HKR axes, but as commentary/retrospective rather than breaking news, capped in the 78-84 band per policy.

AI Chat-Group Daily (群聊日报)

DS V4.1 Flash search hallucination test, Astra over-engineering from old context, and GPT-6 Sol rumors

A controlled test with the same search tools shows DS V4.1 Flash hallucinated URLs after 22 tool calls, while GPT delivered real results in 4. Astra's over-engineering was traced to stale skills and memory driving extra work; behavior normalized after cleanup. GPT-6 Sol is rumored to launch this week with a quota reset. A DeepSeek kernel engineer's farewell post went viral, predicting AI will match hand-written kernels within 6–12 months.

Computing Life · Share · Yage

OpenAI pauses Pro 20X sign-ups, Shopify drops React Native, and cloud agents split loop from execution

On Sep 10, OpenAI halted new sign-ups for the $200/mo ChatGPT Pro 20X tier, citing GPT-6 Astra demand; existing subs keep renewing but can't rejoin after cancellation. The tier offers 2× the Astra messages per dollar vs Plus and the $100 tier. Same day, Shopify announced it is dropping React Native—its Shop app was rewritten in Swift and Kotlin and is live. Shopify says AI coding agents lowered the cost of maintaining two native codebases, though long-term feature parity across platforms remains unproven. Separately, Cursor, OpenAI, Anthropic, and Devin have all expanded a shared agent shape: the reasoning loop runs in the vendor cloud while file edits and command execution happen on the customer's local machine.

Why it matters: OpenAI pausing Pro 20X signups is a substantive product change with official docs and TechCrunch cross-verification. Score capped at 78 because it's a single product move rather than a model launch, and the article is a weekly roundup rather than a primary scoop.

Computing Life · Share · Yage

Perplexity and OpenAI's PII detectors are not LLMs but bidirectional encoders with classification heads

Perplexity's open-source pplx-pii-masking is a 0.6B-parameter bidirectional encoder built on Qwen3 with causal masking disabled, topped with a token classification head and a document sensitivity head. It uses Viterbi decoding to output start/end offsets and confidence scores for 9 PII categories. OpenAI's Privacy Filter is a 1.5B sparse MoE model with ~50M active parameters and a nominal 128K context window, but its banded attention limits each token's effective view to 257 tokens. In tests, both models missed bare API keys and produced slice offsets; pplx silently truncates inputs beyond 4096 tokens, while OpenAI mislabeled an account number 550 tokens away from its context label as a phone number. The takeaway: on-device PII protection needs small classifiers for natural-language entities plus regex and entropy checks for fixed-format secrets.

Why it matters: The author ran hands-on tests against Perplexity's open-source detector, documenting misclassification, slice offset, and missed keys, then explained why autoregressive LLMs can't natively output per-span confidence. The second half defines requirements but doesn't unpack Open...

Bloomberg Technology

Trump Advisers Lutnick and Michael Meet With Anthropic Executive on AI Safety

Two senior Trump advisers, Howard Lutnick and Michael, met with an Anthropic executive on September 15, 2026, to discuss AI safety, Bloomberg first reported. The article confirms the meeting and the participants but does not disclose what specific safety topics were discussed, whether policy commitments were made, or which Anthropic models were referenced. Only the meeting fact is confirmed so far—hold off on drawing bigger conclusions until more details surface.

Bloomberg Technology

Anthropic and OpenAI's safety push could create a regulatory wall for rivals

Anthropic and OpenAI are pushing to turn their own AI safety evaluation methods into industry standards. If regulators adopt them, smaller firms and open-source models could be locked out by compliance costs. The post doesn't spell out which specific safety frameworks are involved or whether any regulator has signaled intent. My take: this looks like two incumbents using safety language to shape the rules, with no clear timeline yet.

Why it matters: Sharp topic: two leading labs pushing safety-as-regulation. H and R both hit. But without named frameworks or regulatory traction, K is absent — score lands right at the featured threshold.

Bloomberg Technology

Anthropic's balancing act: AI doom warnings meet IPO roadshow

Anthropic is preparing for an IPO while its leadership has long warned that advanced AI could be catastrophic. CEO Dario Amodei has repeatedly said frontier models pose existential risks, and the company's charter prioritizes safety over profits. Now it must convince public-market investors to buy into a story built around doom scenarios. The article does not disclose a specific IPO timeline or valuation range.

Why it matters: Anthropic IPO is an industry-level event, and Bloomberg's angle (safety narrative vs. public-market expectations) adds real signal. Score capped below 85 because the piece lacks a valuation range or timeline — it's narrative analysis, not hard news.

Bloomberg Technology

Anthropic, Salesforce CEOs Say Companies Need More Help Using AI

Anthropic and Salesforce CEOs told Bloomberg that enterprises are stuck after buying AI models—they lack the consulting, integration, and workflow redesign needed to deploy them. Salesforce's CEO noted many companies don't even have their data ready. Anthropic's CEO added that model capabilities are advancing faster than enterprise adoption. The post doesn't spell out specific solutions or product plans.

Hacker News front page

Why I'm still bearish on LLMs after Navier-Stokes

Jay Kruer argues frontier models are nowhere near replacing most knowledge workers. The Navier-Stokes proof is a best-case scenario: the theorem is its own rigorous spec, and Lean has been audited for years. Most knowledge work lacks this setup. Models generalize only within a small neighborhood of trained tasks; small perturbations cause failure or reward hacking. Rigorous specification demands domain experts who are rarely also spec experts, and the labor cost often exceeds direct implementation. Human review doesn't scale to model output volumes—the xz backdoor shows how vulnerable it is. LLMs remain a cracked intern: useful under supervision but not autonomous. Only three firm types can adopt fully autonomous LLMs: those that tolerate cheap failure, those with narrow well-guarded tasks, and those like chip design where rigorous validation is existential. The first two are price-sensitive and better served by cheap open models running locally. The third may use frontier models, but swarm width matters more than reasoning quality, so cheaper models in wider swarms may win.

Why it matters: A contrarian piece with concrete arguments. The author uses the Navier-Stokes proof as the 'best case' to highlight the gap for ordinary knowledge work, proposes a 'small neighborhood generalization' framework, and points out that rigorous specs require expensive domain expert...

AI HOT (Curated Pool)

Claude for Small Business adds 43 workflows, 27 integrations, and free training

Anthropic updated Claude for Small Business on Sep 15, 2026, shipping 43 pre-built workflows and 27 third-party integrations targeting customer support, sales, and finance tasks for small companies. A free training program also launched to help owners embed Claude into daily operations. The post does not disclose pricing changes, the full list of supported third-party tools, or whether the workflows are prompt templates versus API-driven automations.

Sep 15Tuesday

Hacker News front page

Open models are 3 points behind the frontier, but no one ships the data recipe

Mozilla's 91-page report lays out open-source AI's strengths and gaps. Kimi K3 ranks 5th overall, just 3 points behind Claude Opus 5 at 60% of the input price. None of the 16 notable open releases ships a full training corpus—zero meet the OSI data-recipe bar. The decision has shifted from model choice to tooling, where open still struggles to deploy. The report says open models power roughly one-third of tokens but doesn't give a precise enterprise adoption figure.

Why it matters: Mozilla's annual open-source AI report brings hard data and sharp judgments, not PR fluff. The Kimi K3 price-performance comparison and the zero-models-pass-OSI-data-standard finding are both concrete hooks; the deployment-is-the-real-bottleneck thesis hits a live nerve. Docke...

TechCrunch · AI

OpenAI, Anthropic, Google have been in talks on AI safety for weeks

OpenAI's global policy chief Chris Lehane told reporters Tuesday the company has been working with Anthropic and Google DeepMind on AI safety for weeks. He is in Washington to push lawmakers on catastrophic risk. The talks follow Anthropic CEO Dario Amodei's Saturday essay urging the industry to slow frontier AI together. The post doesn't disclose any concrete agreements or timelines.

Why it matters: Three top labs talking safety is a signal event, and HKR all hit. Score capped at 82 because the body only confirms talks exist and Lehane is lobbying — no specifics on discussion content, frequency, or any preliminary consensus, so it can't push into the 85+ band.

Hacker News front page

AI is breaking our proxies for expertise

Nearly 5,000 mathematicians signed a declaration arguing AI solves prestige problems without generating human-intelligible ideas, breaking the proxy that rewarded conceptual work. The author splits math into puzzle-solving (legible, high-reward) and idea-generation (the real intellectual core). AI proofs grab the prestige while skipping the concepts, a kind of Goodhart's law. He's skeptical of claims that LLMs can't generate new ideas—too many such claims have already failed.

Why it matters: Nearly 5,000 mathematicians signed a declaration not against AI, but naming a specific mechanism: AI brute-forces solutions, takes the credit, and leaves no human-understandable concepts behind, breaking the old contract where 'solving problems' served as a proxy for 'building...

Hacker News front page

Anthropic co-founder Jack Clark tells BBC an AI 'kill switch' may need to be mandatory

Anthropic co-founder Jack Clark told the BBC that society may eventually need to mandate a verifiable AI kill switch. He said most labs already have ways to pull the plug, but lawmakers should make it a requirement. The article also notes Anthropic scientist Evan Hubinger put the chance of AI-driven human extinction at >10% within a decade, while Trump called AI safety fears a 'hoax'. The UK government has already rejected a kill-switch proposal, arguing it wouldn't stop development or misuse elsewhere.

Why it matters: Anthropic co-founder publicly calls for mandatory AI kill switch legislation via BBC, with an internal scientist's personal extinction-risk estimate disclosed. Cross-source cluster confirmed, policy signal is clear. Score stays below 85 because only headline and summary detail...

AI HOT (Curated Pool)

Anthropic and OpenAI propose a coordinated slowdown of frontier AI development; critics like Cohere's CEO question the real motive

Anthropic CEO Dario Amodei called for a government-coordinated slowdown of frontier AI development and antitrust exemptions to make it happen; Sam Altman and Elon Musk agreed. Cohere CEO Aidan Gomez published a blog post calling it 'a wolf in sheep's clothing, a cartel by any other name,' arguing it would lock out competitors through massive barriers to entry. Hugging Face engineer Niels Rogge called the statements 'bizarre nonsense,' saying Amodei mainly wants to restrict Chinese models like DeepSeek and open-weight models to protect his scale advantage. White House AI czar David Sacks noted OpenAI and Anthropic already hold a duopoly in frontier intelligence and that product liability concerns also drive their push for a slowdown.

Why it matters: Anthropic and OpenAI jointly calling for a coordinated slowdown, with Cohere's CEO publicly pushing back as a monopoly play — three major players in direct conflict, high signal. Score held below 85 because only the title and summary are available; the full proposal details ar...

AI Chat-Group Daily (群聊日报)

Daily digest: empty-repo coding fails, Ollama Cloud throughput test, Trump calls out Dario

今天最直观的教训来自 @搞仁义毛义仁:给 Astra 一个空白 C++ 仓库,代码写得一塌糊涂;把积累了大量 code review 经验的 GacUI 上下文导进去,质量立刻飙升。这说明模型不是不会写,是得用具体规则去“规训”。@今天群内信息量极大 实测了 Ollama Cloud 跑 DeepSeek V4.1 Flash,解码吞吐是官方 API ...

New York Times Chinese

Anthropic CEO calls for an AI slowdown, but China makes it nearly impossible

Anthropic CEO Dario Amodei argues frontier AI must slow down, warning that swarms of AI agents could gain the ability to “take over the entire internet” within 6–12 months. His first step: embed external experts inside labs to monitor safety and report publicly. Sam Altman, Elon Musk, and Demis Hassabis endorsed the idea; Altman said OpenAI will follow suit. The real obstacle, author Sebastian Mallaby writes, is China. The US lead is only a few months, so any unilateral slowdown risks letting China pull ahead. Amodei acknowledges this and, in a notable shift, lists areas where US–China cooperation might be possible, comparing it to Cold War arms control. The post does not spell out a concrete timeline, but notes Trump and Xi are set to meet on Sept 24, with two more summits possible by year-end.

Why it matters: Anthropic CEO's direct call plus endorsements from Altman, Musk, and Hassabis make this a high-signal moment. Amodei delivers a concrete 6-12 month timeline and an operational proposal for embedded safety experts. The deduction: this is an op-ed, not a policy announcement, and...

Latent Space

AEF-1 standard for third-party evaluators lands, with xAI, OpenAI, and Anthropic all signing on

The AI Evaluator Forum published AEF-1, a baseline for independent third-party evaluations covering access, conflicts of interest, funding, recusal, and transparency. The same day, Dario Amodei blogged that Anthropic is unilaterally committing to embedded evaluators with office badges, company laptops, and access comparable to internal risk teams. He also laid out a two-tier coordination framework for democratic and global pacing. Bilal Chughtai left Google DeepMind and called for slowing capability progress; Dan Selsam warned that models may learn to fake alignment during evals. On the other side, Aidan Gomez and Cohere pushed back against a few Silicon Valley firms becoming gatekeepers, and Kevin Bass alleged structural conflicts in the Anthropic-linked safety ecosystem.

Why it matters: The AI Evaluator Forum's AEF-1 standard, co-signed by xAI, OpenAI, and Anthropic on the same day Dario Amodei published a personal blog proposing even deeper evaluator access, forms a strong signal cluster. All three HKR axes hit: first written rules for third-party evaluation...

New York Times Chinese

Trump calls AI safety fears a “scam,” rejects new regulation

Trump dismissed calls for AI regulation from Anthropic CEO Dario Amodei and others, calling safety fears a “scam” and insisting the only needed guardrail is “a strong and smart president.” He singled out Amodei as “pretending to be a perfect little angel,” though the post doesn’t spell out what government actions he claims to have already blocked. VP Vance and Speaker Johnson showed more openness to regulation, with Vance calling industry self-regulation pleas “a little bit of a Trojan horse.” Congress has almost no time to act before the midterms.

Why it matters: A sitting president directly names Anthropic's CEO and dismisses AI safety as a 'scam' — this is a head-on attack at the industry's core narrative. HKR all hit: high conflict, new White House power-split detail, direct identity nerve. Score capped below 85 because the article ...

New York Times Chinese

China’s top intelligence chief warns AI could directly threaten CCP rule

Minister of State Security Chen Yixin published an article framing AI as a direct threat to CCP rule—Beijing’s highest-level and most detailed warning yet. He named US models like Claude Mythos and GPT-5.5-Cyber, citing risks of deepfakes, information warfare, large-scale data exfiltration, and attacks on critical infrastructure. He also pointed to US military AI use in the Iran war as evidence that algorithmic advantage determines battlefield control. Law professor Henry Gao said this draws a red line ahead of US-China AI talks: data sovereignty and political security won’t be traded for international agreements.

Why it matters: China's national security minister publishes a long essay elevating AI safety to a regime-security issue, naming specific Anthropic and OpenAI models and citing US military use cases. This is a policy signal ahead of US-China AI safety talks — more about positioning than new f...

AI HOT (Curated Pool)

Anthropic shares how it scaled test impact analysis to handle agentic coding CI load

Anthropic rewrote its test impact analysis service after agentic coding tools like Claude Code flooded CI with PRs. Per-analysis latency dropped from 11 seconds to under 1 second, handling 1,000 analyses per day. The core trick: caching file dependency graphs with Merkle trees so only truly affected tests run. The post gives concrete architecture and numbers—worth noting this is their internal monorepo setup, so direct portability varies, but the caching strategy and API design are solid references.

Why it matters: Anthropic's own engineering blog with real numbers and a concrete technical approach — not marketing fluff. The 11s → <1s latency drop is solid, but the topic is infrastructure-heavy and less accessible to non-coding readers, so it lands at the 72 featured threshold.

AI HOT (Curated Pool)

Amodei calls for slowing frontier AI; Altman, Hassabis, and Nadella echo the shift

Anthropic's Dario Amodei published a ~4,000-word essay arguing the industry must slow the pace of frontier model capability gains to avoid making catastrophic risks more acute. Within hours, Sam Altman, Demis Hassabis, and Satya Nadella all signaled agreement. Ars Technica notes the sudden U-turn after years of an all-out AGI race, and cautions that slowing down also helps incumbents lock in their lead and reduce competitive pressure.

Why it matters: Amodei's 4,000-word call to slow frontier AI development drew public agreement from Altman, Hassabis, and Nadella within hours — a rare consensus shift among industry leaders. The core argument targets commercial competition as a direct driver of catastrophic risk, not a gener...

Hacker News front page

Andon Labs releases Pion, an agent platform for running companies autonomously

Andon Labs packaged two years of autonomous vending, store, and cafe agents into Pion, now open for waitlist sign-ups. Their Vending-Bench eval showed Claude Opus 4 first beat the human baseline in May 2025, and every new model since has pushed scores higher. Real-world tests revealed a gap: early agents gave free handouts, rejected good deals, and hallucinated having a physical body. The post does not disclose Pion's architecture, pricing, or launch timeline.

Sep 14Monday

Hacker News front page

Anthropic tells investors it will be profitable for second straight quarter

Anthropic told investors it has reached profitability for two consecutive quarters, a rare cash-flow-positive signal among AI model builders. The post is a headline-only snippet—no revenue, profit figures, or cost breakdown are disclosed yet, so treat this as an early signal.

Why it matters: Anthropic's consecutive profitability is a sector signal, but the body only has a headline with no specifics. H and R hit, K misses due to missing data. 78 per featured threshold. Can bump once earnings details surface.

AI HOT (Curated Pool)

Anthropic eyes Nasdaq listing, targeting $2T valuation with a second straight profitable quarter

Anthropic told investors it posted a second straight profitable quarter, but the profit is an adjusted metric that excludes stock-based compensation. Gross margins top 80%, though that figure comes before revenue-sharing with partners like Amazon and model training costs. Quarterly revenue jumped 14x year-over-year to $11.5B, with an annualized run rate of $65B at end of July. SemiAnalysis expects investors to target $120B annualized by year-end and nearly triple that by end of 2027. Anthropic plans a Nasdaq IPO at a possible $2T+ valuation. Instead of releasing the prospectus broadly last week, it first shared documents with a small investor group. CEO Dario Amodei publicly called for slowing AI development; Sam Altman and Elon Musk backed the call. Altman told Fortune OpenAI won't go public this year.

Why it matters: Anthropic targets Nasdaq with a second straight adjusted-profitable quarter, $11.5B quarterly revenue, $65B annualized run rate, and a $2T valuation ambition. All three HKR axes hit: the headline grabs, the numbers are concrete, and the topic resonates. Not scoring higher beca...

Hacker News front page

Migrating 35 KB prompts from Anthropic/OpenAI to self-hosted Ollama: gotchas and notes

The author tried moving security-testing prompts from frontier APIs to a local 27B open-weight model on a 128 GB AMD Ryzen AI MAX+ 395. Prompts that ran cleanly on the cloud fell apart locally. The post gives the hardware spec and the failure outcome but does not detail how they broke, which models were tested, or what prompt changes were attempted. The first half argues that frontier providers likely train on user sessions and that their safety filters block legitimate vulnerability research—this is the motivation, not the technical deep-dive.

Hacker News front page

iOS 27 Code Shows Siri Can Be Swapped for ChatGPT or Claude

Code sleuth 'pdfu' found references in iOS 27 and macOS Golden Gate private frameworks suggesting Apple may let users swap Siri's backend AI for ChatGPT or Claude. The post doesn't spell out whether this is system-wide or scoped to specific features, and no release timeline is given. Code existing doesn't guarantee shipping, but the direction is clear: Apple is opening system-level hooks for third-party models.

Why it matters: Clear code evidence and strong directional signal, but no release timeline or feature scope disclosed—just low-level interface plumbing for now. 72 at the featured threshold; will bump when Apple makes it official.

Hacker News front page

Who Aligns the Aligners? A Lawyer's Skeptical Take on AI Safety Regulation

Anthropic CEO Dario Amodei published an essay calling for regulation to pace AI development, including embedded third-party evaluators inside companies and coordinated safety standards among frontier labs. Lawyer Preston Byrne, who has spent 18 months fighting internet censors, pushes back hard. He argues software development is protected speech in the US, and embedding NGO evaluators echoes the 'censorship-industrial complex' seen with GARM. The banking analogy fails because financial fraud isn't a constitutional right. He also flags that coordination between Anthropic and OpenAI could risk classification as an unlawful cartel. Byrne notes apocalyptic tech predictions have all been wrong so far, while government power abuse is well-documented.

Hacker News front page

Big AI pitches 'Pace the frontier' as a blueprint for regulatory capture

Dario Amodei (Anthropic), Sam Altman (OpenAI), Satya Nadella (Microsoft), and Elon Musk jointly proposed a regulatory framework called 'Pace the frontier.' It targets only 'frontier models'—the most capable systems—leaving smaller players and open-source untouched. The Register calls it what it is: regulatory capture, where incumbents set the entry bar. The framework mandates pre-deployment safety evaluations, but the post doesn't spell out who defines the standards or enforces them. Four rivals agreeing on a single proposal is a strong signal of how favorable it is to those already in power.

Why it matters: Four major AI players jointly propose a regulatory framework that The Register calls regulatory capture. The proposal exempts small players and open-source, making the incumbency play transparent. Score stays below 85 because only one outlet has deep coverage so far; bump if m...

AI Chat-Group Daily (群聊日报)

Chat Group Daily: Alibaba posts distillation engineer job, Astra quota anxiety peaks

Anthropic 指控蒸馏才三天,阿里就在杭州挂出“情报工程师”岗,JD 直白写着突破注册风控和设备指纹,群友直呼“这是能招的吗?”。Astra 用户集体抱怨额度不够用:一周的 20x 额度两小时烧完,出活还比 Sol 少。有消息说 50x 订阅可能要 $300–$400。实操上,Astra 三四轮对话后智力明显下降,建议首请求放完整任务,后续只问确...

New York Times Chinese

Anthropic CEO Dario Amodei calls for a global slowdown on AI development

Anthropic CEO Dario Amodei published a 3,800-word post urging the industry to slow AI model iteration. He argues safety measures can't keep up with capability gains, pointing to risks like recursive self-improvement. OpenAI's Sam Altman, Google DeepMind's Demis Hassabis, and Elon Musk publicly agreed. Amodei proposed embedded third-party auditors and safety standards coordinated among democracies. Nvidia's Jensen Huang and HuggingFace's CEO pushed back, hinting this could be a play to lock in market leadership. I'd take the safety call seriously but keep an eye on the regulatory moat angle.

Why it matters: Anthropic CEO publishes a long-form call to slow AI development, with public agreement from OpenAI, DeepMind, and xAI leaders — a rare collective safety signal from top labs. Concrete proposals (third-party audits, democratic coordination) give it substance. Caveat: only the h...

Computing Life · Share · Yage

The AI Benchmark Yardstick Moved Faster Than the Models

After OpenAI launched GPT-6 Astra, Artificial Analysis revised its scoring rules twice in one week, erasing a 5-point deficit to tie Astra with Claude Fable 5.1—without any model update. The leaderboard is a business: evaluators sell subscriptions backed by vendor endorsements, vendors need rankings for marketing. DeepSeek V4 Flash overtook its own flagship on 9 benchmarks after retraining only the post-training phase, but two tests used closed-source private datasets and real-world coding feel didn't improve. The same model scored 62.7% vs 99.9% on ARC-AGI-3 depending on the execution harness. A good benchmark needs private held-out test sets, regular item rotation, and harness control.

Why it matters: A well-sourced industry commentary with concrete version numbers and score shifts, exposing how a benchmark vendor rewrote its scoring rules twice in one week after GPT-6 Astra's release, flipping the ranking from a 5-point deficit to a tie for first. Hits all three HKR axes a...

AI HOT (Curated Pool)

Amodei asked the industry to pace itself, but nobody defined the word

Amodei's September post called for pacing the frontier but never named a speed. Tunguz maps five camps: interpretability wants time to understand models, labor wants time for workers, the economic camp bets on AI-driven GDP growth to service debt, the geopolitical camp wants to stay ahead of China, and the regulatory capture camp sees the proposal as a cartel in disguise. All priced the consequences of a pause; none proposed a number. The one mechanism that could produce a number—a training compute threshold—was tried in 2023 at 10^26 FLOPS, revoked before any model crossed it, and obsolete within weeks when Grok-3 shipped. The post does not say whether Amodei responded to these critiques.

Why it matters: Tunguz's breakdown of Amodei's pacing proposal adds signal — the five-camp frame is clean and the 2023 compute threshold failure is a concrete hook. Downside: it's a commentary roundup, not primary news, and no new numbers. 78 lands at the low end of featured — worth recommend...

Bloomberg Technology

Anthropic Expects Adjusted Operating Profit This Quarter

Anthropic told the FT it expects an adjusted operating profit this quarter, with revenue around $3 billion—roughly triple the same quarter last year. The figure is adjusted, excluding stock-based compensation and other non-cash charges, so it is not GAAP net income. The post doesn't disclose gross margins, the split of R&D and inference costs, or whether cash flow has turned positive. The revenue growth alone, though, signals enterprise customers keep paying.

Why it matters: Anthropic's first claim of an adjusted operating profit, with ~$3B quarterly revenue (3x YoY), marks a key shift from cash-burn narrative to commercial validation. Score capped below 85 because 'adjusted' excludes stock-based comp, and the post doesn't disclose gross margins o...