Skip to content

#其他

3 today

Sep 23Wednesday

Hacker News front page

Claude Code's AGENTS.md support is gated behind a remote flag and silently fails when telemetry is off

Claude Code 2.1.277 announced AGENTS.md support, but the loader is controlled by a remote feature flag (tengu_agents_md_mod) that defaults to false. The author found that setting DISABLE_TELEMETRY=1 or CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 silently prevents the local AGENTS.md from being read, with no warning. Setting the variables to 0 doesn't help, and project-level settings.json can't override it. The only workaround is a one-line CLAUDE.md containing @AGENTS.md. The author argues that reading a local file should never depend on telemetry, and at minimum a skipped file should trigger a visible message.

Why it matters: This is a product behavior exposé backed by concrete code evidence, not a rant. The author traced the silent skip to the remote flag tengu_agents_md_mod and confirmed AGENTS.md is ignored when telemetry is off. HKR all hit, but the blast radius is limited to Claude Code users ...

OpenAI News

Ringg cuts customer service costs by 90% with GPT-5.6, resolves 65% of calls via AI

Ringg, an Indian customer service platform, uses OpenAI's GPT-5.6 family to power voice and chat agents. It handles over 7 million calls monthly, with AI resolving up to 65% of requests and a 4.8 CSAT score. The trick: route real-time conversations to GPT-4.1, post-call analysis to GPT-5.6 Terra, and evals to GPT-5.6 Sol. Moving to GPT-5.6 cut costs by 90% for some workloads. The post doesn't clarify whether the 65% resolution rate is fully automated or includes human handoffs, nor does it disclose specific latency numbers.

OpenAI News

Sam Altman at the UN Security Council: loss of control and power concentration are the two big AI risks

Sam Altman addressed the UN Security Council on September 23, framing AI risk in two buckets: losing human control over AI, and concentrating too much power in too few hands. He said OpenAI has unilaterally slowed down before and will do so again, rejecting the idea that competitive pressure forces rash decisions. He pushed back against any single actor claiming only they can be trusted with the most powerful models. The speech is a full transcript; it does not disclose new products or policy specifics.

Why it matters: Altman's UN Security Council remarks aren't a product launch, but he frames AI risk as two concrete threats—loss of control and concentration of power—and publicly states OpenAI has unilaterally slowed down before and rejects the 'only we can be trusted' argument. That's a dir...

TechCrunch · AI

Ema raises $77M Series B to replace enterprise software and IT services with AI agent teams

Ema uses teams of AI agents to automate enterprise workflows across HR, IT, and finance. The $77M Series B was led by Creaegis, with Accel, Section 32, and Prosus participating. Total funding now stands at $140M, and the valuation more than quadrupled from the last round. The startup has over 50 enterprise customers, including Google and Microsoft. The post doesn't disclose revenue or how much manual work the agents actually replace—hold off on the 'eating software' narrative until retention and deployment data surface.

AI HOT (Curated Pool)

Cursor improves token efficiency for long agent runs, cutting user costs by 7%

Cursor cut token costs for long agent runs by 7% through four engineering changes, with no quality regression. They trimmed the system prompt by ~66% as models now need less hand-holding; offloaded 60% of built-in tool definitions from static context to dynamic loading (similar to the 46.9% token reduction they previously achieved for MCP tools); compressed file reads; and used subagents strategically. The post doesn't disclose the absolute dollar or token amounts behind the 7% figure, nor the specifics of the compression and subagent implementations. The savings come from production A/B tests, so your mileage will vary by model and task length.

Why it matters: Cursor's official blog discloses four concrete token optimization techniques with numbers and methods, directly useful for developers using Cursor. But this is an incremental engineering improvement, not a product-level update, and the impact is limited to the Cursor user base...

AI HOT (Curated Pool)

Cursor launches Rollouts and Security Reviewer bots

Cursor released two bots to handle post-PR grunt work. Rollouts tracks a change from PR to production, compares telemetry against a pre-deploy baseline, and alerts the author, pauses the rollout, or creates a revert PR when it spots a regression. Security Reviewer runs on every PR, traces user input end-to-end across the codebase, and cut average review time from 4.8 to 3.8 minutes while lifting fix acceptance from 45–50% to 60–70%. Both are available today on Teams and Enterprise plans.

Why it matters: Cursor ships two workflow bots that automate rollout monitoring and security review with concrete mechanisms and numbers. Not a model release, but a high-utility product update for the dev audience with all three HKR axes hit. Score stays at 78 because it's an incremental feat...

Hacker News front page

Tokens Too Cheap to Meter

jyn argues with multiple charts that AI inference cost is dropping by orders of magnitude each year. GPU efficiency doubles roughly every two years, and per-task model cost in 2026 is two orders of magnitude cheaper than end of 2025. Inference engines like vLLM add 10%–50% throughput gains annually. The author expects LLMs to become computing infrastructure within 1–2 years, and frontier-quality local models on commodity hardware in 3–6 years. The post doesn't cite specific dollar figures, but the trend lines are stark.

Why it matters: A data-backed cost trend analysis, not vague 'AI is getting cheaper' talk. Charts three decline curves — GPU efficiency, inference engine optimization, local deployment — and engages with Jevons paradox and ROI questions. Docked slightly for being a personal blog rather than i...

AI HOT (Curated Pool)

Qwen releases Qwen-Audio-3.1 family: ASR, TTS, Realtime upgrades plus new TTS-Next and ASR-Next

Qwen dropped five audio models covering ASR, TTS, real-time conversation, and creative generation. The Realtime model supports interruption and slows down with empathetic responses when it detects low mood. Pricing is slashed: TTS ~70% off, Realtime ~85% off, ASR up to 95% off. The post doesn't disclose benchmark scores or latency numbers.

Why it matters: Full Qwen audio stack refresh with two new product lines (TTS-Next, ASR-Next), not just a version bump. The ~70% TTS price cut and Realtime emotion-aware interruption are verifiable details. Held below 85 because the post doesn't disclose ASR-Next and TTS-Next capability bound...

New York Times Chinese

Thomas Friedman: The AI threat is real — the U.S. and China must act before it's too late

Thomas Friedman warns that AI is showing self-improvement, replication, and self-preservation traits, while the Trump administration dismisses extinction risk as a 'scam' and lacks coordinated governance. He urges the U.S. and China to agree on AI guardrails at Thursday's summit before an AI-driven crisis hits. Treasury Secretary Bessent said AI leaders won't get a 'liability shield,' but the piece doesn't say whether a concrete deal will emerge from the meeting.

Why it matters: A Friedman byline that puts AI safety on the US-China summit table — timing and author weight carry it. Capped at 78 because it's commentary, not original reporting; no new data or experiments.

AI Chat-Group Daily (群聊日报)

Anthropic Opus 5.5 and OpenAI Sol/Luna drop same day; community breaks down effort cost-efficiency and migration pitfalls

Anthropic 毫无预兆地放出 Opus 5.5,在终端操作和编程任务上跑分领先,但 max 档输出 token 量是 GPT-6 Astra 的三倍多。群友分析发现 high 档是性价比甜区:比 medium 多花 36% 的钱,智能指数涨 3 分,再往上边际成本陡增。两小时后 OpenAI 上线 Sol 和 Luna,Luna 输入价格打到每百...

Why it matters: Anthropic Opus 5.5 launched without warning, OpenAI followed with Sol and Luna two hours later — three model resets in one day. The daily digest provides real-user effort-tier cost/performance breakdowns and prompt-migration war stories, high signal density. Deduction: this is...

AI HOT (Curated Pool)

OpenAI releases GPT-6 Sol and Luna with 50% cheaper API pricing and benchmarks

OpenAI added two models to the GPT-6 family: Sol for complex coding and professional tasks, Luna for fast high-volume work. API pricing is cut by 50% vs GPT-5.6 promo rates—Luna's output price actually dropped 58%. Sol beats Claude Opus 5 on AutomationBench and Agents' Last Exam at roughly one-tenth the cost per task. Both are live in the API today; no weights are released.

Why it matters: OpenAI drops two new GPT-6 variants with a 50% API price cut — an industry-shaking move. Sol's Aura score and Luna's $0.5 output price are concrete, though the post doesn't include the full benchmark table. Still, this is a must-cover story.

New York Times Chinese

U.S.-China Summit Puts AI on the Table, but Little Progress Is Expected

AI safety and competition dominated this week's U.S.-China summit, but deep mistrust makes concrete outcomes unlikely. The Trump administration has loosened chip export curbs, letting Nvidia sell H200 chips to China, while U.S. officials accuse Chinese firms of stealing AI models through distillation. Both sides agreed to set up a hotline for AI-related national security risks and plan to meet again in Shenzhen in two months. Senator Warren warned Trump against catering to the AI industry instead of pressing Xi on AI risks. Analysts expect talks to stay at the level of definitions and principles, since neither side will accept limits on its own competitiveness.

Why it matters: NYT's exclusive on US-China AI talks packs real substance: a safety hotline, H200 export relaxation, and distillation-theft accusations. Score capped at 78 because it's policy maneuvering, not a product or tech breakthrough — high signal but low immediate actionability for bui...

Bloomberg Technology

SoftBank sells record-yield junk bonds to fund its AI push

SoftBank launched a jumbo high-yield bond sale to finance its AI push, offering record yields. The article does not disclose the exact coupon. High borrowing costs signal the market is pricing in serious risk on SoftBank's AI bets, but it also shows Masayoshi Son is willing to pay up to stay in the game.

OpenAI News

ChatGPT Ads expands to 7 Southeast Asian markets and Taiwan, now in 60+ countries

OpenAI rolled out ChatGPT Ads to Indonesia, Malaysia, the Philippines, Singapore, Thailand, Vietnam, and Taiwan. Ads only appear for Free and Go users; Plus, Pro, and Enterprise tiers stay ad-free. OpenAI says it never sells conversation data to advertisers and ads don't influence ChatGPT's answers. The ad business hit a $1B annualized revenue run rate by late August, under 200 days post-launch. Self-serve access is available via Ads Manager, with agency partners including dentsu, Havas, Omnicom, Publicis, and WPP. Shopee is named as a launch collaborator in the region.

AI HOT (Curated Pool)

Qwen-Image-2.1 Tops Arena's Open-Source Leaderboards for Image Editing and Text-to-Image

Alibaba's Qwen-Image-2.1 ranks first among open-source models on Arena's Image Edit and Text-to-Image leaderboards. It scored 1367 in Image Edit Arena, placing 16th overall, just 3 points behind GPT-Image-1.5-high-fidelity at #15. The post doesn't disclose parameter count, architecture, or release timeline.

Why it matters: Qwen-Image-2.1 hitting #1 open-source on Arena's image editing leaderboard, just 3 points behind GPT-Image-1.5, is a concrete cross-model signal. Score stays at 78 rather than higher because the post doesn't disclose parameter count, architecture, or release timeline — the inf...

The Verge · AI

OpenAI enlists elite mathematicians to avoid another fumble

OpenAI is forming a panel of elite mathematicians to advise on reviewing and communicating emerging research results. The move aims to prevent future missteps, though the post doesn't specify past failures or name the advisors.

AI HOT (Curated Pool)

The Most Important Market in AI is the Middle

Tunguz argues that enterprise AI spend concentrates in the 'good enough, affordable' middle tier, not the frontier. Anthropic held Opus at $5/$25 across five releases while OpenAI slashed Luna 80% then 50%; open models run most token volume at an 86% discount to closed models. The priciest model, Fable 5.1, captured only 3.7% of gateway spend in its first 12 days, while mid-tier models claim 40% of spend and 30% of tokens. As intelligence per dollar explodes but enterprise requirements barely move, tokens may shift to commodity—and that will decide the market's economics.

Why it matters: Tunguz uses gateway spending data to make a counterintuitive case: the most capable model, Fable 5.1, captured only 3.7% of spend in 12 days — the mid-tier is where enterprises actually put their money. Opus held price across five releases, open models run majority volume at 8...

AI HOT (Curated Pool)

Modal details how to serve trillion-parameter coding agents at trillion-token scale

Modal's engineering team published a deep-dive on serving Moonshot AI's Kimi K2.6 for coding agents. They boosted per-replica performance by 2.8x per user and 5.6x across users, turning a ruinously expensive service into a price-competitive one. One service processed hundreds of billions of tokens per day and trillions in aggregate. The post walks through workload analysis for hybrid-attention MoE models and the engineering optimizations applied. Exact GPU models and per-request latency numbers are not disclosed in the body.

Why it matters: Modal's engineering breakdown of inference acceleration for Moonshot AI's Kimi K2.6 delivers two hard numbers — 2.8x and 5.6x speedups — directly useful for inference engineers. Not scored higher because this is infra optimization, not a model capability or product update; aud...

AI HOT (Curated Pool)

OpenRouter publishes 2026 embedding model guide covering 37 catalog entries

OpenRouter shortlisted embedding models from its 37-entry catalog for English RAG, multilingual, code, and text-image retrieval. The default pick is OpenAI text-embedding-3-small for its low price and 8,192-token context. For longer inputs, Voyage 4 large offers a 32,000-token window and index compatibility across Voyage 4 tiers. Qwen3-Embedding-8B is recommended for multilingual retrieval with public weights and 100+ language support. Code search goes to Voyage Code 4, while Gemini Embedding 2 and Voyage Multimodal 3.5 handle text-and-image. The free route is Nvidia Nemotron-3-Embed-1B; the cheapest paid option is Perplexity pplx-embed-v1-0.6b at $0.004 per million tokens. OpenRouter notes these checks confirm API behavior, not retrieval quality, and advises testing on your own data before building an index.

AI HOT (Curated Pool)

Claude Opus 5.5 and GPT-6 Sol/Luna launch on the same day, kicking off a new price war

Simon Willison compares three models launched on the same day. GPT-6 Luna drops to $0.10/M input tokens—half the price of GPT-5.6 Luna and one of OpenAI's cheapest models ever. GPT-6 Sol also halves its predecessor's price. Claude Opus 5.5 gets a 20% cut but still costs twice as much as GPT-6 Sol. In testing, Opus 5.5 at max thinking level over-thinks to the point of hitting its 128k output limit, failing to produce even a simple pelican SVG. Each failed attempt cost $2.56 and took nearly 20 minutes. Willison calls the max mode effectively useless.

Why it matters: Three flagship models dropped on the same day, with Simon Willison's first-hand pricing comparison and early impressions. GPT-6 Luna at $0.10/M input is OpenAI's cheapest ever, directly reshaping the cost structure for application builders. Downside: the post only has the pric...

Hacker News front page

Google turns AI agent CC into a family group chat tool

Google expands its AI agent CC to support family groups, letting everyone share one chat thread. CC remembers each member's preferences and schedule, helping coordinate activities and set reminders. The post doesn't specify which chat platforms are supported, pricing, or the underlying model.

Financial Times · Technology

Trump rejects 'globalist scheme' to control AI, dealing a blow to Andy Burnham

Trump rejected the idea of a global AI regulatory body, calling it a 'globalist scheme'. This directly undercuts UK mayor Andy Burnham's push for international AI governance. Burnham had urged nations to cooperate to prevent AI from being monopolized by a few tech giants. The White House's refusal means the US won't join such a multilateral framework soon, dimming the outlook for coordinated global AI regulation.

Why it matters: FT exclusive on Trump explicitly rejecting an international AI regulatory body, directly undercutting Burnham's global governance push. Clear policy signal with strong conflict, hitting all three HKR axes. Score held at 78 (featured threshold) because it's a political stance r...

AI HOT (Curated Pool)

GPT-6 Sol and Luna: near-same intelligence scores at roughly half the cost

Artificial Analysis reports that GPT-6 Sol and Luna score close to their GPT-5.6 predecessors on the Intelligence Index, while token pricing drops ~50%, halving per-task cost. Their chart shows the intelligence-vs-cost trade-off across OpenAI model generations. The post does not disclose exact scores or pricing figures.

Why it matters: First third-party price/performance benchmark after GPT-6 launch — cost halved with flat intelligence is a direct signal for model selection decisions. Deduction because the post only shows chart trends without concrete scores or pricing numbers, so information density isn't s...

TechCrunch · AI

Snorkel AI triples valuation to $3.5B with a $350M Series E for AI training data

Snorkel AI raised a $350M Series E at a $3.5B valuation, nearly triple its $1.3B mark from 17 months ago. The startup builds training datasets and simulated environments for AI labs and enterprises. Its customers include OpenAI, Anthropic, Google, and Microsoft. CEO Alex Ratner says ARR has passed $100M, doubling year-over-year. The post doesn't disclose profitability or per-customer revenue concentration, so I'd discount the valuation a bit until we see how fast big clients build in-house data pipelines.

Why it matters: Snorkel AI tripling its valuation to $3.5B with OpenAI, Anthropic, Google, and Microsoft as named customers is a real signal in AI infrastructure. But training data tooling is behind-the-scenes work — low practitioner resonance, so R axis missed, keeping this at the featured t...

AI HOT (Curated Pool)

OpenAI launches GPT-6 Sol and Luna, API pricing 50% below GPT-5.6 promo rates

OpenAI added two models to the GPT-6 family: Sol and Luna. They inherit most of GPT-6 Astra's capabilities but run faster and cheaper, with API pricing 50% below GPT-5.6 promotional rates. The post credits caching and inference efficiency gains. No benchmarks, latency figures, or rollout timeline are disclosed.

Why it matters: OpenAI added two new GPT-6 models with API pricing halved vs GPT-5.6 promo rates — a real cost signal for builders. Score held below 85 because the post lacks benchmarks, latency data, or a launch timeline; actual performance needs real-world testing.

Latent Space

John Platt on AI for Science: an Oscar, two asteroids, and the algorithm in your sklearn

John Platt, inventor of Platt scaling and SMO, leads Google's ERA project. ERA turns scientific problems into scoreable tasks and uses Gemini to auto-iterate experiments via a Monte Carlo tree search variant. The jump from Gemini 2.0 to 2.5 made it go from broken to highly productive, yielding at least 10 papers. Platt warns against overfitting and says always start with linear regression or SVM. The post also covers his team's work on contrail mitigation, which accounts for 1% of human-induced global warming.

Why it matters: In-depth interview with John Platt revealing Google's ERA project: automated science iteration via Gemini, yielding 10+ papers. Hits all three HKR axes — legendary figure, concrete new mechanism, strong audience resonance. Score capped at 78 because it's a podcast interview ra...

AI HOT (Curated Pool)

OpenAI ships better prompt caching for GPT-6, plus a dashboard and diagnostics

GPT-6 prompt caching now hits more often by default, with discounts for shared prefixes reused within 30 minutes. A new dashboard tracks hit rates and a diagnostics tool pinpoints misses—e.g., a tools_changed reason costing 5,629 tokens. Developers can set explicit cache breakpoints, adjust reasoning effort without breaking cache, and prewarm context to cut latency. GitHub Copilot reports a >50% drop in tokens needing fresh processing; Manus raised cache hit rates from ~85% to >90% in under a week.

Why it matters: Official OpenAI post on GPT-6 prompt caching improvements with a diagnostic dashboard and manual breakpoints — a real cost win for agent developers. Score stays below 85 because it's infrastructure, not a new model, but the concrete numbers and tooling details make it a solid ...

The Verge · AI

Rabbit's new AI agent runs without its R1 hardware

Rabbit launches OS3, an agentic OS that runs in the cloud but operates locally on Windows, Mac, and Linux. One account supports up to five devices; the system auto-selects files, apps, and models for each task. You can also access it via a desktop site, Telegram, or iMessage. The post doesn't spell out pricing, release date, or how OS3 differs from the R1's original system.

TechCrunch · AI

Qualcomm launches two new phone chips focused on on-device AI agents and local models

At its Snapdragon Summit, Qualcomm announced the Snapdragon 8 Elite Gen 6 and the higher-end Extreme variant. Both chips push on-device AI: a new sensing hub runs small models up to 200M parameters locally for transcription, speaker diarization, and memory-based task automation. The Extreme version can run a 30B MoE model on-device, which Qualcomm says enables a full voice-in/voice-out agent. The post doesn't disclose power consumption, pricing, or which phones will ship first.

Hacker News front page

JavaScript's midlife crisis: the ecosystem won, but developers are losing control of the toolchain

JavaScript turns 30. The ecosystem is bigger than ever, but the toolchain is being rewritten in Rust, Go, and Zig. The author argues that rewriting a bundler in Rust makes it 10x faster, but also shrinks the pool of JS developers who can maintain it. The source is still open, but the door to contributions is closing. The post doesn't offer a fix—just a warning: speed isn't free, and we're trading maintainability for milliseconds.

AI HOT (Curated Pool)

GPT-6 Sol and Luna halve cost but show regressions in some evals

OpenAI's GPT-6 Sol and Luna cut prices roughly in half: Sol drops to $2/$10 per million input/output tokens, Luna to $0.10/$0.50. Per-task cost on the Artificial Analysis Intelligence Index falls from $1.99 to $1.06 for Sol and $0.18 to $0.07 for Luna, while overall scores stay level. Hallucination rates drop sharply—Sol from 92% to 60%, Luna from 93% to 77%—but both models decline to answer more often. In the Coding Agent Index, Sol gains 2 points to 57; Luna loses 2 points to 41. Both regress on GDPval-AA v2.1, a knowledge-work benchmark: Sol drops ~100 Elo, Luna ~75, driven by shorter deliverables that omit rubric elements. The cost drop is real; the quality trade-off on knowledge tasks is worth watching.

Why it matters: OpenAI halved GPT-6 pricing, with Sol per-task cost at $1.06 and Luna at $0.07, but capabilities are mixed — Luna actually regressed on the Coding Agent Index. Solid third-party benchmark data makes this directly useful for developer decision-making. Not p1 because this is a c...

Hacker News front page

Ten Claude Opus 5.5 agents produced a faster shortest-path algorithm, C-HD, with a formal proof in Lean

Vals had ten Claude Opus 5.5 agents collaborate over 15 hours to produce C-HD, a shortest-path algorithm for directed graphs with non-negative real weights. Within the certified density range m ≤ n⌊(log₂ n)^(3/4)⌋, its proven bound is O(n + m + m log(2 + m/(n+1)) + m^(1/3)(n log(n+2))^(2/3)). When m ≈ n(log n)^(3/4), the leading term drops from n log n to n(log n)^(11/12); at n=2^1000 the theoretical ratio is about 1.78. The algorithm uses bounded local searches that count non-improving edges as unexplored leaves to limit repeated work, and falls back to Bellman–Ford outside the certified range. Both correctness and the complexity bound are formally verified in Lean. The post does not report large-scale benchmarks—only small correctness simulations—and notes the constants are not yet optimized.

Why it matters: Ten Claude Opus 5.5 agents collaborating to produce a formally verified shortest-path algorithm in 15 hours is novel enough to clear H and K. But it's pure theory with no engineering hook, so R is absent — right at the featured threshold. The post doesn't give concrete perform...

TechCrunch · AI

Meta admits Muse was 'heavily inspired' by OpenClaw

Meta's head of product Nat Friedman confirmed Muse was 'heavily inspired' by OpenClaw, down to workspace filenames and content, though he said the code was built from scratch. Early adopters had suspected Muse was essentially a repackaged OpenClaw; Meta's statement partially confirms that.

Why it matters: Meta's product lead publicly admitted borrowing from OpenClaw, with concrete details and conflict. Hits all three HKR axes, but the event is a design controversy, not a technical breakthrough, capping it at the featured threshold.

AI HOT (Curated Pool)

Pentagon review: overreliance on Palantir Maven AI contributed to strike that killed 123 Iranian children

A Bloomberg investigation cites an unreleased Pentagon review that blames three failures for the February strike on an Iranian elementary school that killed over 150 people, including 123 children. The first was overreliance on Palantir's Maven Smart System: operators expected it to flag stale intelligence, but it recommended the site—still labeled an IRGC facility—as a day-one target. The other two were bad intelligence and outdated satellite imagery. The school had been physically separated from an adjacent base by 2017, with playground markings visible in 2018 imagery, yet databases were never updated. An analyst flagged the change in 2019, but the note stayed in a system disconnected from targeting databases. The pace of the opening assault—over 1,000 targets in 24 hours—squeezed verification time, and civilian-harm teams had shrunk from 10 people to one. A UN fact-finding mission this week called the strike a war crime. Palantir says it is not responsible for underlying data quality and has since added features to re-review intelligence for disqualifying factors. Trump has denied U.S. responsibility, claiming Iran may have done it. The full Pentagon report has been largely complete for months but remains unreleased.

Why it matters: A Pentagon probe directly blames Palantir's Maven system for a catastrophic strike, hitting all three HKR axes. The score is capped slightly because Gizmodo is a secondary source and the event is primarily a military story, but the AI failure lesson carries direct industry war...

Hacker News front page

Pentagon says overreliance on AI contributed to missile strike on Iran school

A Pentagon probe found that overreliance on AI targeting contributed to a 2025 US missile strike that hit an Iranian school. The system mislabeled the school as a military site, and human operators did not override the machine's call within an 11-second decision window. The report names Palantir's Maven system and Google AI tools, though the post doesn't spell out exactly which component failed. This wasn't autonomous firing—it was human-machine teaming where humans deferred to the machine.

Why it matters: A rare, official post-mortem that pins a lethal strike on AI overreliance, naming specific vendors and a concrete 11-second window. Hits all three HKR axes hard. Held back from 92 only because the article doesn't disentangle which part of the Palantir/Google pipeline failed.

AI HOT (Curated Pool)

Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5

Anthropic launched Claude Opus 5.5, the first model in its Claude 5.5 family. The team says it matches Fable 5.1 on most work while costing 40% less to run than Opus 5. It leads Anthropic's internal benchmarks on agentic coding, computer use, and knowledge work. It's not a clean sweep—GPT-6 Astra still leads on Terminal-Bench-Science and AutomationBench. Pricing is $4 per 1M input tokens, $20 per 1M output tokens, and cache reads drop to $0.20, a 60% cut that matters most for agentic and coding costs. Output is over 30% faster than Opus 5, with a fast mode offering 2.5x speed. The model is API-only, no open weights. One early tester migrated 680,000 lines of code in under a day.

Why it matters: Anthropic flagship model refresh with 40% cost reduction and Fable 5.1-level performance — a same-day must-write. Held below 92 because the source is a MarkTechPost relay without a direct official blog link or pricing breakdown.

AI HOT (Curated Pool)

GPT-6 Sol and GPT-6 Luna land on Arena, API pricing 50% below GPT-5.6 promo rates

OpenAI dropped GPT-6 Sol and GPT-6 Luna on Arena, both built on GPT-6 Astra tech. The pitch is faster, cheaper inference with better caching for high-volume workloads. API pricing is 50% below GPT-5.6's promotional rate. The post doesn't disclose benchmark scores or latency numbers, so I'd wait for third-party benchmarks before getting excited.

Why it matters: OpenAI drops two GPT-6 models on Arena with API pricing 50% below GPT-5.6's promo rate — strong price signal. But no benchmarks or latency data in the post, so can't tell if performance took a hit. Score stays below 85 until third-party tests land.

AI HOT (Curated Pool)

Sam Altman says GPT-6 Sol and Luna have no competition on per-task pricing

Sam Altman posted that GPT-6 Sol and Luna have no competition when measured by per-task pricing. He claims big jumps over the 5.6 series in intelligence, alignment, work output, coding, and computer use, with per-token price halved and even lower per-task cost. The post doesn't disclose specific benchmarks or pricing figures—I'd wait for third-party testing before taking it at face value.

Why it matters: Sam Altman personally vouches for GPT-6's per-task pricing, claiming no competitor matches it — a direct signal for anyone tracking inference costs. But the post lacks any benchmarks or pricing numbers, so this is a one-sided claim for now. Score stays conservative until indep...

AI HOT (Curated Pool)

OpenAI launches GPT-6 Sol and GPT-6 Luna, API pricing 50% below GPT-5.6 promo rates

OpenAI released GPT-6 Sol and GPT-6 Luna, both built on GPT-6 Astra tech and aimed at cheaper, faster high-volume workloads. API pricing is 50% lower than GPT-5.6 promotional pricing, driven by more efficient caching and inference. Sam Altman reposted the announcement and called the character designs cute. The post doesn't disclose benchmark scores, latency figures, or regional availability.

Why it matters: OpenAI ships GPT-6 with API pricing 50% below the GPT-5.6 promo rate — a direct cost shock for high-volume developers. Score held below 90 because the post omits benchmarks, latency, and regional availability, so we can't yet judge if performance took a hit.