Skip to content

All news

67 today

Sep 3Thursday

TechCrunch · AI

India's richest man wants to turn aging computers into AI-ready PCs

Reliance Jio opens JioPC cloud PC service to all internet users in India, not just its broadband subscribers. Old computers (up to 8 years) can get 8 vCPUs, 16GB RAM, 1TB storage from the cloud to run AI apps without hardware upgrades. Plans start at ~$11 for 2 months, $42–$53 for 12 months. India had 65M+ PCs in 2025, ~34M of which are aging. The post doesn't spell out latency or bandwidth requirements, so temper expectations.

AI HOT (Curated Pool)

NVIDIA to acquire Hugging Face for $12.93B? Body is just a CAPTCHA wall

The headline claims NVIDIA is acquiring Hugging Face for $12.93B, but the article body is a CAPTCHA wall with zero details. No deal terms, timeline, or official confirmation are disclosed. It's impossible to verify whether this acquisition is real, agreed, or just a rumor.

Sep 2Wednesday

Hacker News front page

Google Introduces Gemini 3.8 Flash and 3.8 Flash Cyber

Google announced Gemini 3.8 Flash and 3.8 Flash Cyber today. The Flash model targets low-latency inference for real-time apps, while the Cyber variant is fine-tuned for cybersecurity tasks. The post does not disclose benchmarks, pricing, or regional availability.

AI HOT (Curated Pool)

Google explains harness engineering: building deterministic guardrails so coding agents can self-repair

Shir Meir Lador from Google AI breaks down harness engineering: wrapping a coding agent in deterministic guardrails—sandboxing, repair loops, and progressive context discovery—so it can self-correct. She cites an OpenAI experiment where 3 engineers shipped an internal beta with zero manually-written lines, and shows a code snippet using Google ADK 2.0 and Antigravity SDK to bound the agent to a workspace and persist its trajectory memory.

TechCrunch · AI

HiddenLayer nabs $100M as enterprises rush to secure their AI deployments

AI security startup HiddenLayer raised $100M, three years after its $50M Series A. Back then, real-world AI attacks at scale were hard to find. Now security vendors are racing to build products that monitor AI agents and the tools they use. The post doesn't disclose valuation or lead investor, but notes the risk of agents going haywire in production is real even if public exploits remain rare.

Why it matters: HiddenLayer's $100M round is sizable for AI security, and the article shifts the threat narrative from models to agents and toolchains — a real knowledge gain. Missing valuation and lead investor details keep it from the 78+ band.

TechCrunch · AI

Amazon's shopping AI can now tell you if that message is a scam

Amazon added a scam-detection feature to Alexa for Shopping. Forward a suspicious email or text, and the AI checks it against billions of official Amazon messages—analyzing sender, content, timing, and metadata. Roughly 360,000 customers per year ask support if a message is real; now they can ask the AI directly. The system improves as users report more scams. The post doesn't specify regional availability or a launch date.

The Verge · AI

NYC bans AI use for students through 8th grade for one year

Mayor Zohran Mamdani announced a one-year moratorium starting in the 2026-2027 school year, barring roughly 600,000 NYC public school students from 2-K through 8th grade from using AI in class. Teachers are also banned from using AI to grade assignments. A small pilot for AI education tools will run alongside the ban. The post doesn't spell out penalties or the pilot's scope.

Why it matters: NYC's AI ban for 600K students is a strong policy signal. Score held back because the article doesn't spell out enforcement or what the pilot actually tests.

Hacker News front page

WebLLM: Run LLMs directly in your browser, no server needed

MLC-AI's open-source WebLLM runs LLMs directly in your browser via WebGPU acceleration. It supports Llama, Gemma, and other popular models, achieving near-native speed on consumer GPUs. The catch: first load requires downloading several GB of weights, and memory usage is high. Great for offline assistants and privacy-sensitive use cases, but don't expect it to replace cloud inference.

Hacker News front page

Three sites made 215,128 "best software" pages for AI. Perplexity cites them

Trellner Research tested Perplexity's sonar and sonar-pro across 380 software categories. 59.8% of citations came from domains ranked worse than #100,000, and 23.4% from domains not in the top million at all. Three sites—wifitalents.com, worldmetrics.org, and gitnux.org—published 215,128 machine-generated "best software" pages between them, all registered after December 2023. Two of them set their homepage HTML title to "Facts & Grounding Page" with a meta description calling it a "machine-readable record"—clearly written for retrieval models, not people. Separately, guideflow.com, a vendor's marketing blog, became the third most-cited source, ahead of Gartner. The test only covered Perplexity, not Google, so the finding is specific to one engine, but it shows how easily retrieval layers treat bulk-generated SEO content as trustworthy evidence.

Why it matters: Trellner Research tested two Perplexity models across 380 software categories and found nearly 60% of citations come from low-authority domains, with three machine-generated sites producing 215K pages and getting cited heavily. Solid methodology, transparent data, and it hits ...

Hacker News front page

A third of Perplexity's citations don't contain the number they're cited for

Haus Research audited Perplexity's two search models across 310 factual questions about tech companies. Of 1,826 citations attached to sentences with figures, 34.7% pointed to pages that either wouldn't open or contained none of the numbers from that sentence. Scored per claim, 14.4% fail. Dead links are only 1.3%; the bigger problems are paywalled pages (16.1%) and open pages that don't carry the cited figure. Example: asked for Vercel's cheapest paid plan, the model answered $20/month and cited Vercel's pricing page, which contains no '$20'. For headquarters, models produced exact street addresses and cited Wikipedia articles that contain none of those addresses. The report also flags a pattern of SEO-spam pages generated per query template, later taken down, that the models cited as sources. The post does not include a response from Perplexity.

Why it matters: Haus Research tested Perplexity's two search models with 310 questions, manually verified 1,826 citations attached to numerical claims, and found 34.7% of links either inaccessible or mismatched. Methodology is transparent and sample size is solid—this isn't a casual dunk. Hel...

Hacker News front page

LLM Intelligence vs. Cost: Why the Log Scale Misleads You

OpenTeams engineer Guido Imperiale argues that ArtificialAnalysis's intelligence-vs-cost plot is misleading. The log scale hides a 250x real price gap between cheap and expensive models while exaggerating trivial differences among cheap ones. He redrew three linear-scale charts, replacing official API prices with OpenRouter's cheapest third-party rates and calculating local-model cost by electricity. The charts show sharply diminishing intelligence returns per dollar above a score of 50. The post does not disclose the exact electricity cost formula.

Hacker News front page

Humanoid robots are nowhere near replacing human workers, despite flashy demos

Humanoid robots look impressive lately: Unitree did backflips at a gala, and a Tiangong robot ran 100m in 8.86 seconds. But author Kai Williams talked to experts and says don't panic. Tesla's 2024 Optimus bartending was teleoperated by humans, not autonomous. Dance and sprint demos are far easier than real manipulation like picking up a Coke can. Robots still can't generalize across tasks, learn on the job, or handle multi-hour projects. Practical hurdles—reliability, cost, safety around people—remain unsolved. The post estimates it will take many years, maybe decades, to overcome these.

Why it matters: A well-sourced, concrete analysis cooling down humanoid robot hype. The author uses recent viral moments — Tesla's teleoperated bartending, Unitree's backflips — to unpack three technical gaps: manipulation, generalization, and sustained work. High information density, not emo...

Hacker News front page

Mistral now trains on user input by default, except on enterprise tier

Mistral now uses free and Pro user input/output data for model training by default. Enterprise tier is excluded. Users can opt out in settings. The post doesn't specify whether historical data is retroactively excluded after opt-out, nor the exact retention period for training data.

Hacker News front page

Anthropic launches a Claude content checker that reads C2PA credentials to tell if a file was made or edited with Claude

Anthropic released a browser-based tool at claude.com/check-content that checks uploaded images, video, or audio for a C2PA content credential tied to Claude. The tool only reads the embedded credential, not the file itself, and the file never leaves your device. A positive result means Claude processed the file; it says nothing about the content's truthfulness. A missing signal doesn't rule out Claude—the credential could have been stripped, or the model/platform may not support marking. Supported formats include JPG, PNG, MP4, MP3, up to 100 MB.

最佳拍档 (BestPartners)

Anthropic releases MHS, a hardware standard for models to control physical devices

The post only has a title with no body. Anthropic announced MHS (Model Hardware Standard), described as a physical-world counterpart to MCP, aimed at letting models like Claude control lab equipment or robots. The title mentions 'physical MCP', 'lab automation', and 'embodied AI', but does not disclose protocol details, supported devices, or release timeline.

AI HOT (Curated Pool)

Cursor launches Self-Hosted Machines so cloud agents run on your own infrastructure

Cursor cloud agents can now execute tool calls on machines inside your network while inference and planning stay in Cursor's cloud. Teams register their own machines via a worker that maintains an outbound HTTPS connection, giving agents direct access to internal repos, private services, and custom hardware like GPUs or Macs. Cursor says over 60% of its internal PRs are already created by cloud agents, and this targets enterprises that need network isolation or specialized infrastructure.

Why it matters: Cursor decouples cloud agent execution from its own infra, letting enterprises keep code and GPUs on-prem while still using the cloud brain. It's a real architectural shift, not a minor tweak. Score stays at 78 rather than higher because it's launch-day with no user validation...

Hacker News front page

LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes

LLM judges fail to detect omissions in AI-generated clinical notes. A new benchmark of 500 note pairs shows detection accuracy for added/altered content at 0.79-0.94, but for omissions only 0.50-0.63—barely above chance. Restructuring the task helps: first list all facts from the transcript, then check each against the note. A two-step pipeline achieves 2.7% false alarms; a single-prompt method catches 12% more omissions at 6.2% false alarms and one-tenth the cost. Two physicians validated the pipeline as more reliable. Both methods miss omissions when the fact is restated elsewhere in the note. Dataset and code are open-sourced.

Hacker News front page

Multiverse Computing releases Quasar 438B, the highest-scoring European model on Artificial Analysis

Multiverse Computing launched Quasar 438B, its first large model, a reasoning model for enterprise agents and coding that supports English and Spanish. It scores 43 on the Artificial Analysis Intelligence Index, the highest among European models, ahead of Mistral Medium 3.5 at 30 and NVIDIA Nemotron 3 Ultra at 38. It outputs 500 tokens in 15.3 seconds, faster and smarter than Mistral Medium 3.5. Long-context reasoning hits 75.0, close to Claude Opus 5 at 75.7. Terminal-Bench v2.1 scores 69.3, leading Mistral by 18.7 points but trailing Claude Opus 5 at 89.1. The model is available via the CompactifAI API. The post does not disclose training data, detailed parameter count, or pricing.

Why it matters: First 438B reasoning model from Europe with concrete benchmark numbers and competitor comparisons — enough signal. But the source is the company's own blog, no third-party testing yet, so score stays at the featured threshold.

Latent Space

Anthropic drops Claude Fable/Mythos 5.1: new SOTA for coding, but 70% more output tokens

Anthropic launched Claude Fable 5.1 and Mythos 5.1 on Sep 1, claiming SOTA on coding and knowledge work. Fable 5.1 hits 55.8% on Terminal-Bench 4.0 and is pitched for autonomous multi-step tasks. Cache read price dropped 75% to $0.25/MTok, but Artificial Analysis found output tokens rose 1.7x, netting a ~20% per-task cost increase. Community speculation suggests Fable and Mythos may share weights with different safety routing—the post doesn't confirm this. Early praise for coding ability is offset by complaints about rate limits, false safeguard triggers, and subscription UX.

Why it matters: Anthropic dropped Claude Fable/Mythos 5.1 with a 55.8% Terminal-Bench 4.0 score, a 75% cache read price cut to $0.25/M tokens, and a 70% increase in output tokens. A capability upgrade plus major pricing shift makes this a same-day must-write. Not a 95 because we only have Lat...

AI Chat-Group Daily (群聊日报)

DeepSeek V4 Flash beats Sol in real-world use; Anthropic drops Fable 5.1

Community members ran two-month SBS comparisons and a week-long 5.1B-token workload on DSH + DeepSeek V4 Flash, concluding it feels better than GPT-5.6 Sol in real tasks. Sol overthinks and produces bloated output; V4 Flash is fast (2.3s first token) and cost ¥362.84 total. A 'subscription gym paradox' theory argues subscription-based harnesses quietly throttle usage while pay-per-token models don't. Anthropic launched Fable 5.1 with 75% cheaper cache reads, but Fable 5 scored below Opus 5. Also: Astra hits Critical cybersecurity tier, Anthropic's $35B compute deal, Qwen 3.8-Max-0902 benchmark run, Microsoft AI secretary setup, and Grok Bot hands-on.

Why it matters: The side-by-side data is solid — 5.1B tokens, ¥362.84 total spend, 2.3s first-token latency — but the source is an anonymized chat log, not an official release or reproducible benchmark. That caps the authority. HKR all hit, so featured is the right tier.

New York Times Chinese

John Ternus Takes Over as Apple CEO, Facing a Balancing Act

John Ternus officially became Apple CEO on Tuesday, with Tim Cook moving to executive chairman. Ternus immediately faces a wave of executive departures: over 400 former Apple employees now work at OpenAI, and several hardware leads have left. He must balance Apple's financial discipline with pushing new products like a foldable iPhone and deeper AI integration. The real test, the article notes, will come if Ternus and Cook clash on strategy around AI investment, the canceled car project, or Vision Pro.

Why it matters: Apple's CEO transition is a major industry event, and the NYT deep-dive provides concrete details on talent drain and internal strategic tensions, hitting all three HKR axes. Not scored higher because this is more a personnel analysis than a product or technical breakthrough—d...

AI HOT (Curated Pool)

Meituan LongCat-2.0 Launches Free Trial on Cline

Meituan LongCat-2.0 is now available for free trial on Cline. The post does not disclose model specs, capabilities, or trial duration—only the title is confirmed.

AI HOT (Curated Pool)

Qwen3.8-Max-0902 tops Code Arena and leads the Pareto frontier at $5/MToken

Alibaba Qwen's new Qwen3.8-Max-0902 scored 1,691 on Code Arena's WebDev leaderboard, ranking first overall. At a blended price of $5/MToken, it's the highest-scoring model on the Pareto frontier. Available now on QwenCloud. The post doesn't disclose further technical details or comparison data.

Why it matters: Qwen3.8-Max-0902 tops Code Arena's overall leaderboard with a $5/MToken blended price and a Pareto-frontier claim — a substantive domestic flagship model update that earns the positive-signal bump. HKR all hit: topping the chart creates suspense, concrete score and pricing add...

AI HOT (Curated Pool)

Nvidia Nears $12.9B Deal to Acquire Hugging Face

Bloomberg reports Nvidia is close to buying Hugging Face for about $12.9B, with the total deal potentially reaching $14B. That's 2.9x its 2023 valuation and roughly 86x annualized revenue of $150M. Nvidia also discussed a $1B employee retention package. No final agreement yet, and details could still shift.

Why it matters: Bloomberg-sourced: $12.9B price, $1B retention, 86x revenue — three hard numbers make this a major story. Hugging Face is the de facto distribution layer for open-source models; Nvidia absorbing it reshapes the inference and training toolchain landscape. Not scoring higher bec...

AI HOT (Curated Pool)

Qwen releases Qwen3.8-Max-0902: 2.4T parameters, 1M token context window

Qwen3.8-Max bumps to the 0902 version with 2.4T parameters and a 1M-token context window. Post-training focuses on coding and cowork, targeting complex enterprise tasks, scientific research, and long workflows. The post doesn't include benchmark comparisons or pricing.

Why it matters: Alibaba Qwen drops a new flagship: 2.4T params, 1M context, post-training aimed at coding and long-chain collaboration. Domestic flagship release gets featured-tier treatment per policy. No benchmarks or pricing disclosed, so real competitiveness is unclear — score held at the...

Hacker News front page

Simon Willison tests Claude Fable 5.1's pelican benchmark across five reasoning levels

Simon Willison ran his classic 'SVG of a pelican riding a bicycle' prompt against Claude Fable 5.1 at five reasoning levels. Low and medium produced near-identical outputs with no visible reasoning, taking ~23 seconds and ~10 cents. At xhigh the model spent 7m51s and $1.83, adding real detail. Max ran for 13m54s and $3.30, delivering his best Anthropic pelican yet—blue hat, basket with a fish, feet on pedals—though he still says it lacks the flair of Gemini 3.7 Flash. Separately, Fable 5.1 hit 52.6% on the new Terminal-Bench-Science 0.1 benchmark, up from 24.7% for Fable 5.

Why it matters: Simon Willison ran a controlled five-tier reasoning comparison on Claude Fable 5.1 with concrete latency and cost numbers, making it more useful than the official announcement. Score stays below 85 because this is a personal evaluation rather than a major capability breakthrou...

Hacker News front page

Anthropic banned a paying user for "suspicious signals" — no warning, no human appeal

A long-time Claude Max subscriber at $200/month got banned overnight with a template email citing "suspicious signals" — no clause, no example, no human appeal path. A colleague in the Philippines was banned right after paying $100 for Max. GitHub issues show similar cases in Brazil, Singapore, and Malaysia, some hitting multiple linked accounts within hours of upgrading. The author argues Anthropic's enforcement feels like 2010s Google account bans: automated, opaque, and disproportionately painful for individuals who depend on the product. Enterprise customers get account managers; Max users get a no-reply address and a reference ID. The author filed an appeal but no longer trusts a single frontier lab with their entire workflow, and plans to diversify across other providers and open-weight models.

Why it matters: Multiple paid users across countries report sudden bans after payment — not an isolated glitch. HKR all hit, but this is a user complaint, not an official statement, so capped at 78 due to information asymmetry.

AI HOT (Curated Pool)

Anthropic releases Claude Fable 5.1 and Mythos 5.1, with hands-on tips from a tester

Anthropic dropped two new models, pitched as its most capable for coding and knowledge work. Tester Thariq says they're solid and a full review is coming. Two practical notes: use low effort for tasks that need less verification or have fewer edge cases, and switching effort no longer breaks the prompt cache.

Hacker News front page

Scott Aaronson: LLMs didn't need built-in self-reference—intelligence just emerged

Scott Aaronson argues that models like GPT 5.6 Pro and Fable discuss Gödel and themselves fluently, yet no one baked self-reference or strange loops into the stack. Those abilities emerged as a free byproduct of pretraining on everything. He says the GEB view that self-reference is the secret of intelligence should be buried alongside geocentrism and phlogiston. The post is a personal essay; it doesn't include benchmarks or quantitative evidence.

Why it matters: Aaronson uses 2026 model behavior to push back hard on GEB and Penrose—sharp take with concrete model references. Downside: it's a personal blog essay with no experimental data, more a high-quality opinion piece than a research output. Featured because the topic sparks real di...

Bloomberg Technology

Nvidia nears $14 billion deal to acquire Hugging Face, possibly this week

Bloomberg reports Nvidia is close to acquiring AI model and dataset platform Hugging Face for roughly $14 billion, with a deal possible this week. The article body is behind a paywall, so deal terms, regulatory approvals, and integration plans are not disclosed.

Why it matters: Nvidia's acquisition of Hugging Face is one of the biggest AI M&A deals this year — the $14B price and timing make it a must-watch. Bloomberg broke the story, so the source is solid, but the paywall blocks details on terms and integration plans.

Computing Life · Share · Yage

Real-time video generation cost drops below playback time, reshaping interactive live streaming economics

fal's H3 Max Live achieves faster-than-playback generation for 5-second clips, letting viewers alter scenes via chat. 15-second clips still take 16 seconds, and cross-clip consistency is unsolved. At $144–288/hour for 768p, a single stream needs 1,400–2,900 concurrent viewers to break even. The real bottleneck is platform access rules across markets, not moderation tech.

Why it matters: A cost breakdown piece that dissects fal's real-time video generation speed, pricing, and consistency gaps. The 5s clip is indeed faster than playback, but 15s lags, and cross-clip consistency is unsolved. The $144-288/hr cost and 1,400-2,900 concurrent viewer breakeven point ...

Computing Life · Share · Yage

Nvidia's $12.9B Hugging Face deal can't dodge antitrust this time

Nvidia agreed to buy open-source model platform Hugging Face for $12.9B, its largest acquisition ever. Hugging Face's annual recurring revenue is about $150M, putting the deal at 86x ARR. Nvidia is buying the default entry point for global developers and the demand signals that come with it. Over the past two years, Nvidia and Microsoft repeatedly dodged antitrust reviews by licensing tech and hiring teams, but Hugging Face's core asset—13M users and platform traffic—can't be moved that way. A full equity purchase triggers mandatory review. The post notes that losing neutrality could cost the 41% of downloads coming from Chinese open-source models, eroding the trust that underpins the valuation.

Why it matters: Nvidia's largest-ever acquisition targets a $150M-revenue platform for $12.9B — 86x ARR says this is about owning the default entry point for 13M developers, not the P&L. Microsoft's exit, mutual silence, and antitrust exposure make this the week's top story. Score capped belo...

The Verge · AI

Google needs Hollywood more than the studios need AI

Google is reportedly approaching major Hollywood studios, offering large sums for licenses to train AI models on copyrighted material. The deals carry little downside for Google, but studios risk trading long-term content leverage for short-term cash. The RSS snippet doesn't disclose offer amounts or negotiation status.

Hacker News front page

Local LLM setup on M4 Pro Mac Mini: Qwen3.6 and Gemma-4 on 48GB RAM

Kevin Lewis shares his local LLM setup on an M4 Pro Mac Mini with 48GB RAM. His main model is Qwen3.6-35B-A3B-OptiQ-4bit (MoE, 3B active params per token, ~20GB RAM), and a lightweight Gemma-4-E4B-it-OptiQ-4bit (~2.4GB). He uses oMLX as the inference server, Tailscale to connect iPhone and MacBook, and runs Hermes agent backend, Apollo chat, Pi coding, and Raycast queries. His point: local isn't meant to replace cloud APIs entirely, but to handle 80% of daily requests, avoiding price hikes, rate limits, model degradation, and data privacy risks. The post doesn't disclose specific inference speed or power consumption figures.

TechCrunch · AI

AfterQuery reportedly becomes Y Combinator's fastest unicorn, now valued at $3.2B

AI training-data startup AfterQuery raised a round at a $3.2B valuation, just five months after its $30M Series A at $300M. YC partner Gustaf Alströmer says it's the fastest launch-to-unicorn in the accelerator's history. The founders, now 22 and 23, were in YC's Winter 2025 batch 18 months ago. In April the company claimed a $100M annualized revenue run rate and named Nvidia, Legora, and Motif Technologies as customers. Unlike Mercor or Scale, AfterQuery trains models on how professionals work—encoding their decisions and reasoning—rather than just verifying answer accuracy. The post doesn't disclose the round size or lead investor; AfterQuery didn't respond to comment requests.

Why it matters: Fastest YC unicorn, 10x valuation jump in 5 months, $100M ARR — three hard numbers. The ding: this is TechCrunch relaying a report, not a primary announcement, and the post doesn't disclose the round size or investors. I'm discounting slightly for that.