Skip to content

All news

69 today

Sep 4Friday

GitHub Blog · AI & ML

GitHub Copilot app for Beginners: Run several agents at once

GitHub Copilot 应用支持同时运行多个智能体会话,每个会话运行在独立的 Git worktree 上,互不干扰且各自保留上下文,可随时切换并从中断处继续。用户可在会话视图中查看各任务标题与进度,例如在同一项目上并行执行 funded sort 开发、无障碍审查和测试运行。

The Verge · AI

Google adds live voice modes to Gmail, Docs, and Keep

Google is rolling out Gmail Live, Docs Live, and Keep Live — real-time voice modes that let you talk to each app. Gmail Live surfaces inbox details without keyword or subject-line digging. The post doesn't spell out what Docs Live and Keep Live do beyond noting they mirror the Gemini Live experience for hands-free note-taking and lookups.

The Verge · AI

Nvidia launches free PAIR tool to link idle computers into a personal AI data center

PAIR is Nvidia's new open-source tool, not a hardware router. It discovers PCs with RTX 20-series or newer GPUs, or Apple M4 chips, on a local network and links them for local inference with tools like Ollama and LM Studio, targeting agentic workflows. The post doesn't disclose latency, bandwidth requirements, or real-world throughput, so I'd hold off on performance expectations.

Sep 3Thursday

AI HOT (Curated Pool)

Google Cloud: Run a 24/7 Agent for $5.70/Month

Google Cloud launched Cloud Run instances, an always-on container for $5.70/month. It avoids serverless scaling-to-zero that kills background loops, and costs less than a $15–25 VM. The author built a tech-briefing agent that scrapes news every 30 minutes, with persistent disk and a web dashboard. Full source code is linked.

Hacker News front page

MBZUAI releases K2 Horizon, a six-model fleet with the 0.9B scoring over 48 on AIME 2026

IFM at MBZUAI released K2 Horizon, a six-model fleet from 0.9B to 375B-A23B. The 0.9B, 3.7B, and 7B models set new SOTA in their size classes; the 0.9B scored above 48 on AIME 2026 with reasoning and tool-use capabilities. The 36B-A4B uses a new MoVA attention mechanism, outperforming larger models per active parameter. This is a full open-science release: intermediate checkpoints, data recipes, code, logs, and evals from pretraining through agentic post-training, under Apache 2.0. The post doesn't disclose specific benchmark comparison numbers or latency data, so real-world performance still needs third-party validation.

Why it matters: IFM dropped six fully open models at once, with the 0.9B hitting 48+ on AIME 2026 math and the 36B introducing a new MoVA attention mechanism — high information density. Not scoring 85+ because IFM isn't an OpenAI/Anthropic-tier lab yet and market validation hasn't caught up; ...

The Verge · AI

ChatGPT, Grok, and Claude all went down at the same time on Thursday

Around 11AM ET Thursday, ChatGPT, Grok, and Claude all started having issues at roughly the same time. ChatGPT returned errors across chat, login, file uploads, voice, search, deep research, and image generation; its status page cited elevated errors for ChatGPT and Codex. Anthropic's Claude chatbot and Claude Code were also affected. The post doesn't detail Grok's specific symptoms, the recovery timeline for each service, or whether the outages share a root cause.

Why it matters: A simultaneous outage across ChatGPT, Grok, and Claude is a rare event that directly disrupts workflows for a huge user base. Missing root cause and recovery timeline keeps it from 95+, but the topic is strong enough for featured.

Hacker News front page

OpenAI, Claude, and Grok all went down at once—users suspect a Cloudflare cascade

A Hacker News thread noted that OpenAI, Claude, and Grok all went down around the same time. Users pointed to Downdetector spikes for Cloudflare, Azure, AWS, and Google Cloud near 7:30, suspecting a cascade from Cloudflare or another shared dependency. Other guesses include user migration overload and deliberate attack, but the post is community speculation—no official root cause is confirmed.

Why it matters: Simultaneous outage across OpenAI, Claude, and Grok with high HN engagement. Downdetector data points to Cloudflare or shared infra as a possible common cause. The event is conversation-worthy but lacks a confirmed root cause, so it lands at the 78 featured threshold rather th...

AI HOT (Curated Pool)

Google DeepMind releases WeatherNext 3, a global weather AI model with 5x resolution boost

Google DeepMind today launched WeatherNext 3, calling it their most accurate global weather AI model. It updates hourly and offers roughly 5x better spatial resolution than its predecessor. The post doesn't disclose specific parameters or open-source plans, but highlights sharper capture of local phenomena like thunderstorms and rain bands. For weather AI practitioners or anyone relying on high-res forecasts, this is the strongest public baseline yet.

The Verge · AI

Google updates AI weather model for sharper rain and snow forecasts

Google rolls out WeatherNext 3, an AI weather model with 5x sharper global resolution than its predecessor. It learns from real-time observations to improve rain and snow forecasts. The post doesn't specify accuracy gains or release timeline.

TechCrunch · AI

Google's latest AI weather model gives you no excuse to forget your umbrella

Google DeepMind and Google Research released WeatherNext 3, a deep learning weather model that tops the Operational WeatherBench leaderboard. Google says it will power weather info in Search, Maps, and Gemini, and be available on its cloud platforms. This is the first time core weather variables feed into Google products directly.

Hacker News front page

Porting a 1993 Amiga game to Godot with Claude Fable 5 reading 68000 assembly

Rabah Shihab fed his 72,758 lines of 1993 68000 assembly to Claude Fable 5 and got the game running in Godot 4 over a weekend. Step one: 34k lines of C++ ported in 21 minutes. Step two: the original assembly rebuilt at 50 Hz. Step three: the 1993 original embedded as a launchable extra. The model added its own CLI test flags, ran vasm, and diffed binaries. Shihab notes some parts were wrong and he didn't catch them for weeks.

AI HOT (Curated Pool)

OpenAI launches Daybreak for Frontline Defenders with $1B to support frontline cyber defense

OpenAI is committing $1 billion in subsidized Daybreak access, training, and technical support, targeting consumption within six months. Priority goes to resource-constrained defenders in the U.S.—water utilities, grid operators, state and local governments, community banks—to help review legacy code, analyze suspicious activity, validate vulnerabilities, and deploy fixes. After recent attacks on U.S. water systems, OpenAI offered affected states and utilities up to $1M in no-cost API credits and assistance. Daybreak now serves over 2,000 approved organizations across Blue (general defense) and Red (specialized cyber models) tiers. The post does not specify how the $1B is measured or list the full set of 35 Daybreak Defense Network products.

Why it matters: OpenAI's official $1B subsidy announcement is concrete in both dollar amount and deployment scenarios—not a fluffy PR piece. The deduction is because this is a forward commitment, not a delivered result, and the post doesn't detail Daybreak's actual capability boundaries. Feat...

Hugging Face Blog

H company open-sources NeoMME: a multimodal-native encoder with no separate vision tower

H company released NeoMME, a family of 260M and 800M multilingual multimodal encoders. It uses a single bidirectional Transformer for both text tokens and raw image patches, trained from scratch with a masked discrete-diffusion objective—no separate vision tower, no causal LM. The fine-tuned NeoMME-Retriever outputs dense and late-interaction embeddings in one forward pass. Both sizes sit on the ViDoRe v3 Pareto frontier for nDCG@10 vs. model size. At 2048×2048 input on an L40S GPU, the 260M model encodes ~51 pages per second, roughly twice ColM's speed. The post does not disclose training data size or the full list of supported languages.

NVIDIA Blog

'NBA 2K27' with DLSS 5 leads 28 new games on GeForce NOW this week

NVIDIA adds 28 games to GeForce NOW this week, led by 'NBA 2K27' with day-one DLSS 5 support. The post doesn't detail DLSS 5's performance gains or which other titles use it. For cloud gamers, this is the first time DLSS 5 ships with a major annual franchise—benchmarks will tell the real story.

TechCrunch · AI

Nvidia confirms it will buy Hugging Face for $12.9 billion

Nvidia confirmed it acquired Hugging Face for $12.93 billion. The platform hosts 3 million models, 1 million apps, and 500,000 datasets, used by over 18 million developers. CEO Jensen Huang said Hugging Face will stay open, with no requirement to use Nvidia compute. Nvidia has released 500+ models and 250 open datasets on the platform. Owning an open ecosystem helps Nvidia optimize for its chips and sell unused capacity.

Why it matters: Nvidia buying HuggingFace for $12.93B is the biggest AI infra M&A this year. The 3M models + 18M devs ecosystem scale, plus Jensen Huang's careful promise to keep it open without forcing Nvidia chips, gives this story shock value, concrete numbers, and instant debate fuel. All...

The Verge · AI

Nvidia is buying Hugging Face for almost $13 billion

Nvidia agreed to acquire Hugging Face for $12.93 billion, bringing the largest open-source model hosting community under the chip giant's roof. Founded in 2016, Hugging Face is often called the 'GitHub for AI'—developers share models, datasets, and tools there. Nvidia says it will scale the platform, strengthen infrastructure, and expand AI access. The post doesn't disclose the deal timeline or regulatory approvals.

Why it matters: Nvidia buying Hugging Face for $12.93B hits the infrastructure layer of open-source model hosting. All three HKR axes fire: the deal itself is suspenseful, the price and platform positioning are new facts, and both model builders and infra people will talk about it. Not scorin...

AI HOT (Curated Pool)

Hugging Face co-founder Thomas Wolf announces NVIDIA acquisition for $12,930,300,000

Thomas Wolf posted on X that NVIDIA is acquiring Hugging Face for roughly $12.93 billion, with no changes for users that day. Wolf said the team will keep the Hub an open, independent, compute-agnostic platform and use NVIDIA's resources to push open-source AI. The post is a single-paragraph statement; it doesn't disclose deal structure, regulatory approvals, or integration timeline.

Why it matters: A ~$13B acquisition that reshapes the AI infrastructure landscape. Wolf's personal confirmation and explicit commitment to an open, compute-agnostic Hub is both reassuring and a new variable for the open-source ecosystem. The post doesn't disclose deal structure, regulatory ap...

最佳拍档 (BestPartners)

FinOps in the AI Agent Era: Who Burned All the Tokens

The post does not disclose details. The title points to a video on token governance and cost optimization in multi-agent systems, referencing Microsoft Foundry. The core issue: when multiple agents collaborate, token consumption spirals, breaking traditional FinOps approaches.

OpenAI News

Playco cuts manual fixes 50% prototyping games with GPT-6 Astra

Playco built Playbot, an AI-powered IDE for game dev, using GPT-6 Astra. From one grey box prototype, the model generated three themed game worlds in one go, most working on first take. Manual fixes dropped 50% vs the previous model. Spatial reasoning, UI responsiveness, and game feel all improved. The model also plays the game to find bugs itself.

OpenAI News

Legora reviewed 41 financial docs in minutes with GPT-6 Astra

Legal tech startup Legora used GPT-6 Astra to run a financial-statement tie-out across 41 documents in a single agent run, cutting a task that used to take evenings or days down to minutes. The model improved nearly 40% over the previous version on Legora's benchmark, catching all 4 planted errors including a £500,000 gap hidden in a revenue note. Final judgment stays with human lawyers. The post doesn't detail the prompt or agent workflow used.

AI HOT (Curated Pool)

NVIDIA to Acquire Hugging Face for $12.9303 Billion

NVIDIA's official blog announced it will acquire Hugging Face for $12.9303 billion. The post body contains only the title and site navigation—no deal details, timeline, or integration plans are disclosed. The figure is precise to three decimal places, but the article offers zero context, so treat this as headline-level info for now.

Hacker News front page

Ask HN: Who is using MCP in production?

A Hacker News thread asks who actually runs MCP in production. One dev built a UK council scraper with Claude, added an MCP interface on a whim, and found Claude autonomously queried it during debugging—surprisingly useful. Another hooked Claude Code to Jira and Figma, calling natural-language Jira a relief but MCP only 'a tiny bit easier' than direct API. A skeptic questioned why a non-editable tool surface beats a readable API client; a defender replied that without MCP, Claude falls back to screenshotting Figma. The post doesn't disclose how widely MCP is deployed in production.

Hacker News front page

Chen Danian returns with a 27B local model that trails DeepSeek-V4-Pro by only 1.3 points in CAICT's MCP benchmark

Chen Danian is back with StartLux, a company betting on local models. Its first release, StartLux-V1.0-27B-Preview, scored 39.25% in CAICT's MCP benchmark—second place, just 1.3 points behind the 1.6-trillion-parameter DeepSeek-V4-Pro. The 27B model runs on consumer PCs without the cloud and ranked first in location navigation, financial analysis, and browser automation. Two case studies: when calculating a two-year Microsoft stock return, Claude Sonnet 4.6 misidentified a trading day due to missing raw data; StartLux backtracked and got it right. Asked to search flights in a browser, Claude said it couldn't open a browser. Chen has publicly claimed local models will catch up with Claude in three years and take 80% of the market—StartLux is his bet on that thesis.

Why it matters: Chen Danian's first model lands second in CAICT's MCP benchmark, with a 27B parameter count that runs on consumer hardware and three first-place sub-scores — a concrete signal for the Agent space. Score capped at 82 because only benchmark results are available; the model isn't...

AI HOT (Curated Pool)

OpenAI Releases GPT-6 Astra: New Benchmarks Set, Cybersecurity Hits Critical Threshold

OpenAI launched GPT-6 Astra, calling it its most intelligent and aligned model. It scored 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench. On OSWorld 2.0 it hit 72.6% at ~40 min per task, nearly twice as fast as GPT-5.6 Sol. In a simulated overreach test, Astra stayed in bounds 100% of the time vs. 48% for the previous model. It rolls out today to select orgs, then to Plus, Pro, Business, Enterprise, and API users. The post does not spell out which cybersecurity benchmark hit the Critical threshold, nor does it disclose parameter count, training cost, or pricing.

Why it matters: OpenAI's next-gen flagship launch saturates three hard benchmarks and explicitly labels cybersecurity capability at the Critical threshold, with a concrete alignment comparison against the prior model. Every AI outlet will cover this today. Not a 100 only because the rollout j...

Hacker News front page

WASM_OS: an OS experiment that runs inside a browser tab

WASM_OS is an OS experiment that boots inside a browser tab in about 1.6 seconds. It ships with a file manager, paint editor, terminal, Lisp interpreter, and a Linux compatibility layer — all compiled to WebAssembly. The codebase is open-source. The post doesn't spell out whether it supports persistent storage or a network stack.

Hacker News front page

Anthropic publishes Claude commerce agent guide, claims up to 35% larger carts

Anthropic published a how-to guide for building shopping agents with Claude. It cites early adopter numbers: carts up to 35% larger and a 60% lift in purchase conversion. The post doesn't name the customers or the test period, so treat those figures as directional. The guide covers search, recommendations, and support, stressing that agents should call live inventory and order APIs rather than relying on the model alone.

MIT Technology Review · AI

Scaling agentic AI pilots across the enterprise

80% of Fortune 500 companies have adopted agentic AI, but scaling remains uneven, says NiCE COO Arun Chandra. The real challenge is treating agents as a cohesive system: connect them to back-end systems, break data silos, and redesign workflows instead of layering AI on outdated processes. Chandra argues agents should be held to the same standards as human workers, forming a hybrid workforce.

Hacker News front page

Polars 2.0 RC: streaming engine is now the default, bringing big memory and speed gains

Polars dropped the first release candidate for 2.0, with the stable release coming in a few weeks. This major version isn't about new features—it cleans up old design decisions and changes defaults. The biggest shift: LazyFrame.collect now uses the streaming engine by default, which the team says is roughly 5x faster overall with much lower memory usage. The trade-off is that row order is no longer guaranteed for joins, group_by, and unpivot unless you set maintain_order. 2.0 also gets stricter: is_in on mismatched types that would lose precision now raises an error, horizontal concat with mismatched lengths fails instead of silently padding nulls, and many implicit casts are removed—string-to-date now requires .str.to_date(), and enum/integer conversions need dedicated methods. Removed APIs raise typed exceptions with migration hints, making it easier for both humans and AI agents to update code.

Latent Space

Meta's Muse Spark 1.3 matches GPT-5.6-Sol, training at >90% discount

Meta released Muse Spark 1.3, now ranked #3 globally on AAII, directly competing with OpenAI and Anthropic's frontier models. Zuck called it their biggest jump yet on coding and agentic work, and promised open weights. Pricing is aggressive: opt into training and the cost drops by over 90%. Meanwhile, two new Stanford courses are teaching agent engineering from scratch, replacing 85% of old material with agent skills, context engineering, and security. Sebastian Raschka also tempered the Astra hype, pointing out that looped transformers aren't new—Nanbeige 4.2-3B already reused layers, trading ~2x compute for parameter savings without inherently hiding chain-of-thought.

Why it matters: Muse Spark 1.3 hits #3 on AAII, directly matching GPT-5.6-Sol, with Zuck promising open weights and a >90% training discount. This is Meta's first time cracking the top tier on a major benchmark, and it reshuffles the open-source landscape. Not a perfect score because it just ...

Financial Times · Technology

Fixing the AI industry’s PR problem

The FT argues the AI industry's biggest problem isn't tech but public trust. Companies hype AGI while failing to deliver reliable products, fueling backlash. The fix: less grand vision, more concrete use cases like medical diagnostics or logistics optimization, and stop framing AI as a human replacement.

Financial Times · Technology

Uber allies with driver unions to slow robotaxi rollout

Uber is lobbying alongside driver unions to require safety audits, geofenced limits, and transition protections before robotaxis can scale. The tension: Uber invests in autonomy but fears Waymo and others will expand faster than it can adapt, undercutting its human-driver network. Unions worry about mass job loss. The push targets key markets like California and New York. The article doesn't disclose a timeline or Uber's own robotaxi deployment plans.

Financial Times · Technology

Law firms seek bespoke differences in legal AI

Law firms are moving beyond off-the-shelf AI, demanding bespoke models that understand specific jurisdictions, precedents, and internal knowledge bases. This pushes legal AI vendors to offer customizable, workflow-integrated solutions rather than one-size-fits-all products.

Hacker News front page

Week 6 of vibecoding an MMO — and it's playable

Eldermyr is a free browser-based MMO built in 6 weeks via 'vibecoding.' 20 players online, no download, progress persists. The post doesn't disclose which AI tools were used, but patch notes show rapid iteration.

Product Hunt · AI

Nex: Claude Cowork for high-volume GTM workflows

Nex is a tool that puts Claude to work in high-volume GTM workflows like sales and marketing. The post doesn't detail integrations or supported platforms, but the pitch is clear: embed AI into business processes, not just chat.

Hacker News front page

9-dan Shin Jin-seo beats KataGo 2-1 with a two-stone handicap, the first human series win over a top Go AI

On July 21, world No.1 Shin Jin-seo defeated KataGo by 11.5 points in 221 moves, winning the three-game series 2-1. It is the first official series win by a human against a top Go engine with a two-stone handicap. After a heavy loss in game one, Shin shifted from imitating AI to a defensive, territory-focused style; in game three he held a 99% win probability from move 80 onward. He earned ₩250M (~$170K) and a Genesis G90. The post does not disclose KataGo's exact version or hardware.

Why it matters: First human series win against a top Go AI with a two-stone handicap, with Shin disclosing concrete tactical shifts and win-rate data. It's a symbolic event with real substance, but a Go match has limited direct knowledge value for AI builders, so the score stays at the featur...

最佳拍档 (BestPartners)

Fable 5.1 cuts cache cost by 75%, but may not save you money

Only the title is available; the post doesn't disclose details. Anthropic released Fable 5.1 with a 75% cache-read price cut, but the title warns it may not actually save money—likely due to low hit rates or tricky pricing. The model also appears in Terminal-Bench and protein design tasks, but no performance numbers or cost comparisons are given.

Hugging Face Blog

A 350M model fine-tuned with GRPO in 100 steps lifts structured-output compliance from 22.6% to 29.7%

A hands-on guide from Hugging Face and Liquid AI that fine-tunes LFM2.5-350M with GRPO via the TRL library. Using only 500 samples and 100 training steps on a free Colab GPU, structured-output compliance on the IFStruct benchmark jumps from 22.6% to 29.7%. The post includes the full notebook, reward-function design, and a local evaluation setup with llama.cpp on a MacBook.

Why it matters: A hands-on guide with concrete numbers and a reproducible recipe — hits H and K. But the audience is narrow and R is absent; tutorial content at the featured threshold gets 72.

Computing Life · Share · Yage

OpenAI Codex's self-wake mechanism: it sets its own alarm to watch CI after fixing code

A system prompt template merged into OpenAI's open-source codex repo in late August reveals how Codex Persistent mode actually works: it's not a 24/7 always-on process, but a wake-check-sleep loop every 1–3 minutes. The template requires the agent to record its goal, latest status, completion condition, and next check time before sleeping, then decide what to do upon waking. One hard rule: persistence does not broaden authorization scope—anything beyond scope requires explicit permission. WIRED reported on this mode earlier, but media headlines saying 'always-on' clash with the code's 'sampled again' language. OpenAI hasn't launched it yet; the backend request still shows 'disabled.' ProAgentBench shows models achieve only 64.4% accuracy in judging when to proactively help, and Anthropic's engineering blog reports a 17% miss rate on real overreach during automated review—two numbers that explain the hold. Tasks suited for it are delivery-type jobs like CI, deployment, and builds that execute for one minute and wait for ten. Open-ended tasks like writing proposals or designs are a bad fit. Three discipline rules from the template can be adopted today: write four-element checkpoints, stay silent when nothing has changed, and prefer deterministic mechanisms.

Why it matters: High information density with concrete sourcing from the open-source repo — reveals the real wake-check-sleep loop and the authorization scope rule. Deduction because this is interpretation of a template, not an official launch; actual product experience is unknown.

Computing Life · Share · Yage

Agent token usage 5× human, but caching discounts cut the real bill to ~2×

OpenRouter data shows agents consume 7.3T tokens weekly, nominally 5.2× human usage. But 70–85% are cached reads; with ~90% discount, the real bill is roughly 2×. GitHub's Knowledge Compressor prototype halves doc length and claims breakeven at 2,000 reuses, but factoring in caching pushes the median to 5,000+. OpenAI's Jalapeño chip beats Nvidia GB200/GB300 on fixed-length benchmarks, yet lacks AgentX scores for real agent workloads. All three stories share one distortion: prompt caching inflates headline numbers.

Why it matters: Three stories bundled, but the core value is the first: someone finally separated nominal agent token consumption from the caching-discounted real cost, landing at ~2x. The OpenAI chip benchmark and GitHub compression prototype are bonuses but less dense. Cross-source cluster ...