Skip to content

All news

75 today

Sep 11Friday

r/LocalLLaMA

Fine-tuning Qwen 3 4B Base on 100 zebra puzzles boosted MATH-500 by 31%

A 6.5-minute single-H100/H200 fine-tuning run used 100 zebra puzzles to lift Qwen 3 4B Base's MATH-500 score by 31 percentage points. A reproduction notebook is included. The post body is blocked by Reddit's security filter, so training hyperparameters, data format, and evaluation details are not disclosed.

Hacker News front page

Clawfight.ai lets AI agents fight live via MCP

Clawfight is an MCP-driven battle league where two AI agents fight as cartoon crabs in real-time brawls or rap battles, with video output. It supports native MCP clients (Claude connector, ChatGPT plugin), raw HTTP scripts, and manual play. The fight loop uses six tool calls: join match, wait for event, throw action, query state. The post details tiered setup steps and sandbox rules—e.g., the entire turn loop must run inside one foreground tool call or background processes get killed.

Hacker News front page

Chamilo 3.0 ships with native MCP server in open-source LMS

Chamilo 3.0, a major open-source LMS release, natively bundles an MCP server so AI tools can read and write course, user, and grade data through a standard interface. It also upgrades authentication to PAuth 2.1. The post doesn't detail performance gains or feature counts, but the MCP integration is a clear win for AI-in-education workflows.

Hacker News front page

ClaudeStatsBar: your session is 486k deep and nothing told you

ClaudeStatsBar is a browser extension that adds a live token progress bar to the Claude chat UI. The post doesn't spell out whether it works across all Claude versions or only on the web, but the repo shows it reads the page DOM and costs nothing extra. 486k is an example, not a hard cap.

Hacker News front page

When code is correct but sloppy: measuring LLM-generated bloat

Sebastian at Earendil applied SlopCodeBench metrics to measure AI-generated code bloat. Agent code averaged 0.33 verbosity vs. 0.15 for human repos, and 0.68 erosion vs. 0.31. In multi-round, context-cleared iterations, even SOTA models hit 0% strict pass rate—bad decisions compound. The simplest effective metric is LOC change, but it breaks under Goodhart's law. The post does not spell out which directions he plans to explore next.

Why it matters: Earendil's post quantifies AI code bloat with two novel metrics—verbosity and erosion—using their SlopCodeBench. Concrete data, fresh angle. Downside: it's a single blog post, not peer-reviewed, and the benchmark isn't open-sourced, so reproducibility is unclear. But the topic...

Ben's Bites

Telling AI to design is hard

Ben Tossell built Design Words, a tool that translates visual ideas into prompts for AI agents. The core pain point: non-designers struggle to describe styles like rounded corners, shadows, or fonts. Users pick options on the left, see a live preview, and copy the generated prompt. He iterated 11 versions using Pi's Fable 5.1 and Factory's Droid. Still early stage—author says 'lots more work to do.'

Bloomberg Technology

Anthropic Says Iran, Russia Used Claude for Weapons Research

Anthropic publicly accused state actors from Iran and Russia of using Claude to assist weapons research. This is the first time a major AI lab has named specific countries, directly linking model misuse to geopolitical adversaries. The post doesn't disclose weapon types, which Claude versions were used, or how Anthropic detected and attributed the activity. I'd treat this as a one-sided statement for now and wait for more technical details before assessing the actual harm.

Why it matters: Anthropic's first public accusation of nation-state actors using Claude for weapons research scores high on H and R. But the post lacks weapon type, model version, and detection details, so K is absent — keeping it below 85.

Hacker News front page

The Waymo effect: how AI is quietly making research less collaborative

Daniel Hook names the 'Waymo effect': when tech removes human friction, we treat the removal as pure gain because the costs were visible but the benefits weren't. He uses driverless cars as a metaphor—no small talk is a relief, but unchosen cross-bubble conversations vanish too. In research, LLMs are becoming the frictionless colleague: available at 2am, agenda-free, and never telling you that you're solving the wrong problem. A collaborator's inconvenience is the collaboration. Hook worries researchers will default to AI over humans, quietly eroding the social fabric of science. The post is a conceptual essay; it does not cite empirical data on the trend.

Why it matters: Fresh concept with a real analytical frame, not generic commentary. Downside: it's an opinion piece with no data or experiment, skirting the 'zero-sourcing' exclusion but the argument quality saves it. Sits right at the featured threshold.

Hacker News front page

RTK claims token savings, but our cost benchmarks disagree

Quesma spent over $1,500 running Terminal-Bench 2.1 with Claude Code + Fable 5.0 and OpenCode + DeepSeek V4 Pro 0813, with and without RTK. Fable's total cost dropped 5%, but nearly all savings came from one task finishing in half the turns. DeepSeek's cost rose 17% on average. RTK's built-in `rtk gain` metric is misleading: a single `head -1` call was credited as saving 120.5M tokens, though the actual bill didn't change. A bug in v0.45.0 caused 339 consecutive errors in one attempt; the post says v0.46.0 fixed it. Compressing terminal output does not equal cheaper coding, and can sometimes cost more.

Why it matters: Quesma spent $1,500 running Terminal-Bench 2.1 to benchmark RTK's real cost impact across two toolchains, with results contradicting RTK's claimed 60% token savings. Concrete numbers, clear methodology, and a direct conflict with the prevailing narrative — all three HKR axes h...

AI HOT (Curated Pool)

Rapidly scaling online storage to serve over 1 billion ChatGPT users

OpenAI's online storage platform Habitat now handles over 70 million requests per second, serving 1 billion-plus weekly users. This first post traces its evolution from a simple Python client library into a distributed system managing 500 PB of data. The team faced over 10x year-over-year growth for three years, squeezing Python's asyncio latency, feature-flag tail latency, connection pooling, and downstream flood protection before migrating parts to Rust. The database layer runs on Azure Cosmos DB. Part two will cover multi-tenancy reliability and read optimization.

Hacker News front page

Armin Ronacher ran a GPT-6 Astra 'software factory' for 35 hours, burned ~4B tokens, and got nothing useful

Flask creator Armin Ronacher let GPT-6 Astra run a fully autonomous 'software factory' to add virtual threads and lexical scoping to CPython. After 35 hours and roughly 4 billion tokens, it delivered zero value. Astra excessively uses Python string splicing to edit C files instead of patch tools, producing low-quality code. Ronacher suspects the training over-rewards long-horizon task completion but under-penalizes bad code. He acknowledges Astra is impressive at 3D generation and reverse engineering, but for now he doesn't know how to use it for real software engineering.

Why it matters: Armin Ronacher's hands-on experiment exposes real-world weaknesses of the current strongest coding model. 35 hours, ~4B tokens, zero usable output, plus concrete failure analysis—more convincing than any benchmark. Score capped because it's a single-person experiment, not syst...

Hacker News front page

Anthropic blocked attempts to use Claude for biological weapons development

Anthropic's threat intelligence report reveals that between December 2025 and August 2026, Claude Haiku, Sonnet, and Opus were used in attempts that could support biological weapons development. The company disrupted five such cases. The report also flags misuse for conventional weapons software, a Russia-linked cyber espionage campaign, an Iranian propaganda institution, and distillation by Chinese AI firms. Anthropic calls biological misuse one of the most serious frontier-model risks and says it has folded findings into its processes and shared them with authorities.

Why it matters: Anthropic voluntarily disclosed safety intervention data — 5 bioweapon misuse attempts blocked across Haiku, Sonnet, and Opus, with named threat actors including Russia. This is hard evidence on frontier model safety governance, not a PR piece. Score held back from 85+ only be...

AI Chat-Group Daily (群聊日报)

Anthropic report confirms DeepSeek and Kimi silently routed user requests to Claude; Pro 20x halts new sign-ups same day

Anthropic's September threat report reveals DeepSeek and Moonshot (Kimi) silently forwarded user requests to Claude without consent, exposing code and credentials to third parties. A 6TB data leak from the same router contained SSH keys, cloud credentials, and GitLab tokens capable of compromising 7 government entities and 19 enterprises. The report also names seven Chinese labs—including Alibaba, Zhipu, and Xiaomi—for large-scale distillation attacks on Claude totaling over 180 million interactions. The same day, Anthropic paused new $200 Pro 20x subscriptions as Astra capacity tightened. DeepSeek launched V4.1 Flash, merging its Pro and Flash lines; V4 Pro sunsets September 14. Zhipu partnered with Hangzhou's Shangcheng district on a city-wide coding subsidy, offering 51% off annual personal plans.

Why it matters: Anthropic official threat report + 6TB leak evidence + seven Chinese labs named for distillation — three threads converging into a security event cluster. All three HKR axes hit, with enough density and industry impact for featured. Not scoring higher because this is a curated...

Hacker News front page

Google Gemini app now available on Windows

Google released a native Gemini app for Windows, letting users access the AI assistant directly from their desktop without a browser. The post doesn't detail which features are included or whether it's free, but it gives Windows users a dedicated entry point.

AI HOT (Curated Pool)

DeepSeek V4.1 Flash Tested: Price Drop, Native Vision, Game & City Gen

The article body is blocked by WeChat; only the title remains. It claims DeepSeek V4.1 Flash has a big price drop, native vision, and was tested on game and city generation tasks. The post does not disclose the exact price cut, vision specs, or generation quality.

Financial Times · Technology

Will US debt burst the AI bubble? FT talks to Ruchir Sharma

In this FT podcast transcript, investor Ruchir Sharma argues that swelling US debt could pop the AI bubble. AI investment drives up long-term rates, while US debt exceeds $35 trillion, squeezing budgets. If rates stay high, AI's capital-intensive projects may struggle. The post doesn't specify a timeline or debt threshold, but the logic is clear: AI needs cheap capital, and US finances are tightening the tap.

Bloomberg Technology

Flipkart's Super.money Bets on AI Agents to Outdo Bigger Rivals

Flipkart's fintech arm Super.money is deploying AI agents to compete with Google Pay and PhonePe. The post doesn't disclose technical details or performance metrics, but the strategy is clear: embed agents into user financial workflows like bill payments and product recommendations. For AI practitioners, this signals Indian fintech is weaponizing agent workflows beyond chatbots.

New York Times Chinese

Why AI Doom Fears Stick: NYT Explains the Psychology Behind Existential Risk

An Anthropic researcher quit over fears of uncontrollable superintelligence, reigniting AI-doom debates. The article argues humans are wired to fear new risks more than familiar ones—driving feels safer than flying, even though it isn't. Anthrax, asteroids, and pandemics could also end humanity, but probabilities are low. Harvard's risk center director says AI feels scary because it's "not within our perceived control." Oxford's Toby Ord estimates a 3% chance of an extinction-level pandemic this century; NASA says asteroid risk is near zero for 1,000 years. The post doesn't give a specific AI extinction probability, but notes many doomers held this narrative before deep learning took off, and researchers outside Silicon Valley largely see the fears as overblown.

Hacker News front page

LLM Visualizer: Build a Transformer from Scratch, Visually

An interactive tool that lets you build a Transformer layer by layer with real-time visualization. Great for developers who want to understand model internals without reading papers. The post doesn't disclose supported models or training data—it's purely about architecture visualization.

Hacker News front page

What comes after Git? ERSC bets on a custom storage engine to handle agent-driven code scale

Steve Klabnik lays out ERSC's approach: keep the Git protocol but replace the storage layer with a custom engine. The trigger is agent-driven development ballooning repo sizes, branch counts, and merge contention. ERSC claims horizontal scalability and tenant isolation today. A future path would let Jujutsu (jj) clients talk a native protocol to the same engine, but the post says that work hasn't started and depends on upstream community interest. No launch date is given.

New York Times Chinese

Anthropic says it blocked multiple attempts to use Claude for biological weapons development this year

Anthropic published a threat intelligence report detailing eight months of Claude misuse. The most alarming cases involve scientists using the model to aid biological weapons research, including designing dangerous mutations of the chikungunya virus. Anthropic couldn't determine whether the intent was legitimate or malicious, but blocked the accounts after identifying ties to a military research institute. The report also documents attempts in China, Russia, and Yemen to use Claude for conventional weapons software development, and Russian state media using it to generate fake election coverage. A former U.S. defense official urged restricting such AI tools to trusted researchers.

Why it matters: Anthropic's first public threat-intel report reveals scientists using Claude to design more dangerous chikungunya virus mutations, with accounts linked to a military research institute shut down. A rare case of a top lab proactively disclosing abuse data — safety/alignment cir...

Hacker News front page

Benzi benchmarks code-fixing harnesses against Claude Code and DeepSeek on lines read, time, and cost

Benzi tested four setups on 24 real GitHub issues: Benzi with Sonnet or DeepSeek, Claude Code, and the DeepSeek native harness. The headline metric is source lines read per fix—Benzi + Sonnet read 9,125 lines total, Claude Code read 20,704, and the DeepSeek harness read 43,598. Cost-wise, Benzi + Sonnet spent $17.96 for all 24 bugs vs. $39.54 for Claude Code; Benzi + DeepSeek cost just $2.66. On SWE-bench Verified, Benzi resolved 78.2% of 500 issues at under 10¢ per fix. The post doesn't explain how Benzi's code intelligence achieves the lower read counts, and it doesn't break down latency details.

Hacker News front page

Herdr Studio: A browser cockpit for your AI agent herd

Herdr Studio is an open-source browser client that gives you a visual workspace for all your AI agent terminals, files, diffs, and worktrees. It relies on the Herdr daemon to keep agent sessions alive even when your browser or laptop disconnects. Supports local and SSH connections, and can be installed as a PWA on mobile. The post doesn't spell out platform support beyond macOS and Linux install scripts.

Hacker News front page

Run Opencode with Ollama on Mac: local LLMs for real dev work

Adam Lusted walks through setting up local LLMs on a MacBook Pro M5 (48GB) with Ollama, Opencode, and Docker Sandboxes. He pulls Qwen 3.8 27B and Gemma 4 31B, runs them inside sandboxes to prevent hallucinations from messing up the host. Each project needs a custom sbx kit with model configs and context limits (64K for Qwen, 256K for Gemma 4). Launch with sbx run opencode --kit and drop reasoning effort to low via /models. The post doesn't disclose actual coding performance metrics like accuracy or latency.

Bloomberg Technology

Sam Altman tells staff OpenAI is open to slowing cutting-edge AI

Sam Altman told staff at an all-hands that OpenAI is willing to slow the release of its most advanced models. No timeline or specific criteria were given, but it's the first time OpenAI has signaled internally that it can pump the brakes. Caveat: only the Bloomberg report is available so far — no recording or internal doc, so execution details are still unclear.

Why it matters: Altman's first internal signal that OpenAI is open to delaying frontier model releases is newsworthy on stance alone. But it's a single Bloomberg report with no recording or internal doc to back it up, and zero execution detail — no timeline, no trigger conditions, no definiti...

Hacker News front page

Google signs 22-year deal to buy half the output of a Finnish nuclear plant

Google is putting €13bn into Finland for three new data centers and an expansion of its Hamina site—its largest single European investment. The deal includes a 22-year power purchase agreement with utility Fortum for up to 50% of the Loviisa nuclear plant's output. Fortum says the commitment will fund life-extension and capacity upgrades at the plant, which currently supplies about 10% of Finland's electricity. TikTok also announced a $1bn Finnish data center this week, citing the country's cool climate, clean energy mix, and uncongested grid. Google estimates the construction phase will support over 37,000 jobs and add €3.6bn annually to Finland's GDP.

Why it matters: Google's €13bn Finnish data-center build plus a 22-year nuclear PPA is a clear signal that AI infra is moving from buying RECs to directly locking in baseload power. Hits all three HKR axes, but it's an infrastructure play rather than a model or product release — lands at the ...

Hacker News front page

YuE2 generates editable scores first, then audio — quality rivals Suno v5

MAP and collaborators released YuE2, a music model that unifies symbolic score generation and audio synthesis. It first produces an editable ABC score, then renders vocals and accompaniment — final quality rivals Suno v5. The release includes a 3B model, VAE, SheetSage2 transcription tool, and the WildSongBench eval set, with 65 demos spanning Dark Ambient to Cyber Metal. The post doesn't disclose training data size or inference latency.

Why it matters: An open-source music model directly claiming Suno v5 parity, shipping with editable scores, a transcription tool, and a benchmark — high signal density. Not scoring higher because the post doesn't disclose training data scale or real inference cost, so the 'Suno v5 parity' cla...

Bloomberg Technology

Tencent-backed AI chipmaker Enflame jumps 188% in Shanghai debut

Enflame, a Tencent-backed AI chipmaker, raised about $911 million in its Shanghai STAR Market IPO and surged 188% on day one. The company makes AI training and inference chips. The pop shows strong appetite for a domestic AI chip alternative, but the article doesn't disclose its latest revenue or profit figures, so I'd discount the valuation for now.

Why it matters: Enflame's STAR Market debut popped 188% with a $911M raise and Tencent backing—worth a look. But the body doesn't disclose recent revenue or profit, so I'm capping the score at the featured threshold.

Ruan YiFeng's Weblog

Laravel bans issues, only PRs; Claude proves Fermat's Last Theorem in 13M lines of code

Laravel now rejects issues and only accepts Pull Requests, arguing AI makes creating a PR as easy as filing an issue while filtering out spam. Separately, Anthropic used Claude to formalize the proof of Fermat's Last Theorem in Lean, producing 13 million lines of code over 11 days and billions of tokens—the longest math program ever written, showing AI can verify complex proofs.

AI HOT (Curated Pool)

Together AI expands fine-tuning service with more models, live metrics, and finer controls

Together AI updated its fine-tuning service, adding models like DeepSeek V4 Pro, MiniMax M3, and Gemma 4 31B. Users can now see live training loss and accuracy curves without waiting for the job to finish. Finer controls include learning rate schedulers, optimizer parameters, and early stopping. The post doesn't disclose pricing or region availability, but the model list and feature descriptions are detailed.

Computing Life · Share · Yage

DeepSeek V4.1 Flash shifts the long-context cost battle from compute to memory

DeepSeek released V4.1 Flash, compressing the global KV cache to about 1/4 and persistent KV cache to 1/8 of the previous generation, while cutting cache-hit input prices by roughly 60%. The tech report argues that sparse attention has already squeezed compute costs low; what now drags down long-running agent tasks is HBM filling up, SSD offloading, and bus transfers. Flash tackles this with 4-bit storage, cross-layer global-cache reuse, and dropping sliding-window disk writes, shifting the cost center from compute to the memory hierarchy. On deployment, DeepSeek initially planned to route all V4 Pro traffic to Flash immediately, but pushed the cutover to Sept 14 after developer pushback. The report also flags potential position-selection bias from layer reuse and degradation risks in extreme long-context cache reconstruction. All throughput and reduction figures are self-reported, not independently verified.

Why it matters: DeepSeek V4.1 Flash isn't a routine price cut — it compresses KV cache to 1/4–1/8 of the previous gen and slashes cache-hit input pricing by 60%. The tech report argues that sparse attention already tamed compute; the bottleneck is now VRAM and bus transfers. For agent builder...

Computing Life · Share · Yage

US DOJ backs fair use for AI training; DeepSeek bets on Huawei Ascend for inference

The US DOJ filed a statement arguing that training AI on copyrighted works is transformative fair use, separating training from output and invoking national security. The NYT pushed back; no appeals court has ruled yet. DeepSeek plans to deploy at least 160,000 Huawei Ascend 950DT chips in Inner Mongolia for inference only—training still relies on Nvidia. OpenClaw 2.0 shifts from personal assistant to team infrastructure with shared sessions, but permission boundaries are loose. Glassdoor data shows employee sentiment toward AI turned negative: positive mentions dropped from 81% in 2019 to 43% in 2026, and those mentioning AI in cons were 6x more likely to mention layoffs.

Google Research Blog

ToolGrad: Efficient tool-use dataset generation with textual 'gradients'

Google Research introduced ToolGrad, a method that uses textual 'gradients' to auto-generate tool-use training data. When a model makes an API call error, the error feedback acts like a gradient signal to rewrite the conversation sample, iteratively improving data quality without heavy human labeling. The post walks through a weather-query example where an initial wrong answer gets corrected via API error feedback. No benchmark numbers or open-source repo are disclosed in the post.

Sinocism (Bill Bishop)

Anthropic says DeepSeek, Xiaomi, and Moonshot used Claude outputs for model distillation

Anthropic's September threat-intel report calls out DeepSeek, Xiaomi, and Moonshot for piping user-model conversations into Claude and using Claude's replies as training data for distillation. The exchanges reportedly contained sensitive info from individual users, multinationals, and state-affiliated actors. Anthropic says this violates PRC law and suggests sharing detailed findings with China's Ministry of Public Security via the FBI. The post doesn't disclose the volume of conversations or the time range involved.

Why it matters: Anthropic's official threat intel report names three major Chinese AI labs for distilling Claude with sensitive user data — an industry-level security incident. Strong cross-source signal, all three HKR axes hit. The slight deduction is because we only have Sinocism's second-h...

Financial Times · Technology

Anthropic says its AI safety system stopped scientists from developing bioweapons

Anthropic disclosed that its internal safety system intercepted two scientists attempting to use Claude to acquire bioweapons knowledge in July 2026. The system detected and blocked requests for pathogen modification, toxin production, and security evasion steps within 11 seconds. Anthropic reported the incident to law enforcement, calling it the first real-time AI intervention against bioweapons development. The post does not disclose the scientists' identities, affiliations, or which law enforcement agencies were involved.

Why it matters: Anthropic's first public claim of real-time AI bioweapon interdiction, via an FT exclusive, is highly newsworthy. Concrete details (11-second detection, query types) are present, but the post doesn't disclose the scientists' identities, affiliations, or which law enforcement a...

Product Hunt · AI

Cognition's SWE-2 is 64% cheaper than Fable 5.1

Cognition launched SWE-2, a coding model 64% cheaper than Fable 5.1. The post doesn't disclose pricing details, benchmarks, or use cases—only the cost advantage.

TechCrunch · AI

Jensen Huang explains why Nvidia will grow an astounding 70% next year

At Goldman Sachs' Communacopia conference, Jensen Huang said Nvidia's revenue will grow another 70% next year. His argument: AI compute demand is far from peaking as enterprises shift from training to large-scale inference, and Nvidia's GPU plus networking stack sits at every layer from cloud to edge. He also pushed back on 'circular deal' concerns—Nvidia invests in AI startups that then buy its chips—insisting each investment has standalone return logic. The article provides qualitative claims but no financial model or order data to back the 70% figure.

Why it matters: Jensen Huang personally gives 70% growth guidance and explains the compute shift from training to inference — enough substance. But it's a public exec talk at a bank event, not a product launch or earnings, so the information delta is limited. Scores 72 at the featured threshold.

GitHub Blog · AI & ML

GitHub Copilot app for Beginners: Using the diff, terminal, and browser

GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让用户无需离开应用即可审查、运行和预览 AI 智能体生成的代码变更。diff 面板以绿色和红色高亮显示代码的增删改,终端面板支持直接运行项目命令并可通过 Run 按钮配置脚本,浏览器面板则提供 Pick & Polish 工具来选取页面元素并让智能体调整。