Skip to content

All news

69 today

Sep 25Friday

Hacker News front page

Jev: An AI code reviewer that prioritizes behavior over diffs

Jev is an open-source code review tool that prioritizes understanding developer intent before explaining code changes with AI. It offers a local CLI, agent skill, and GitHub extension. The post doesn't disclose which model it uses, pricing, or performance benchmarks.

Financial Times · Technology

No product, no problem: investors place big bets on AI neolabs

The FT reports on a wave of 'neolabs'—AI startups with no product or revenue raising billions on the strength of their founding teams and research ambition. It names Safe Superintelligence Inc (Ilya Sutskever, valued at $30bn), Thinking Machines Lab (Mira Murati, $20bn), and Reflection AI. Investors are betting these teams can build the next foundation model, but the article warns the costs are enormous and the payoff timeline is entirely uncertain.

Why it matters: FT's first systematic look at the neolab phenomenon, naming three companies valued above $20B with concrete numbers and a fresh angle. Not a product update, but funding signals are industry barometers. Score capped below 85 because the body was truncated and Reflection AI deta...

QbitAI · WeChat

Ant Group and Tsinghua Open-Source 9B Model Realtime-Venus for Async Parallel AI Task Execution

Ant Group and Tsinghua University open-sourced the 9B-parameter model Realtime-Venus, featuring a Harness framework that decouples front-end dialogue from back-end execution. It enables the model to handle multiple tasks in parallel—chatting with users while asynchronously performing operations like data retrieval or form filling. This addresses the inefficiency of turn-by-turn AI interaction. The 9B size suits resource-constrained deployments. The post does not disclose training data, benchmark scores, or deployment costs.

Hacker News front page

DHH at Rails World 2026: Hey is leaving Rails for Rust and native apps, built entirely by LLMs

DHH opened Rails World 2026 by declaring himself retired from professional programming and now a 'maker.' He says English is the best programming language and hand-written code is no longer economically productive. 37signals is using LLMs to rewrite Hey into six native apps with a Rust backend—Rust is hideous for humans but great for LLMs. He wrote 150k lines of code in August; Ruby dropped to 3% of his output. Rails is reframed as a framework for 'web apps of necessity,' with convention-over-configuration rebranded as token efficiency. The author questions how products differentiated by UI/UX survive if everything becomes CLI-driven by agents. DHH offered Rails devs pep-talk confidence but no actual roadmap.

Why it matters: DHH's Rails World 2026 keynote barely touched Rails itself, instead delivering provocative claims backed by concrete numbers and product decisions. The post is a second-hand reaction rather than the full keynote transcript, and actual Rails roadmap details are thin—hence not p...

Hacker News front page

Jev and System One Models: Calibration Beats Accuracy

TypeSafe AI released Jev, a non-autoregressive 'System One' model that answers structured questions with probabilities in a single forward pass. The author argues calibration, not accuracy, is the real bottleneck for production classifiers, and Jev's training objective (RLCD) directly optimizes for honest probabilities. Jev achieves 70–500 ms latency and costs $0.042 per million input tokens. The author lacks API access, so performance claims are from TypeSafe's launch post; the post does not disclose independent benchmarks.

Product Hunt · AI

Jango: Test multi-user apps with AI agents that each have their own browser, account, and memory

Jango sends a group of AI users to your app, each with its own browser, account, goals, and memory. Point it at a dev URL and watch them sign in and interact in real time. You can join in yourself or take over any user's screen. A final report logs actions, errors, and screenshots. Mac only for now; you can bring your own AI key or use Jango's managed AI. The post doesn't disclose pricing or which models are available.

Simon Willison

Northern Gannet, Great Blue Heron, California Brown Pelican

在加州 Monterey Bay National Marine Sanctuary,于晚 7:07 至 7:27 观测到 Northern Gannet、Great Blue Heron 和 California Brown Pelican。新入手的 200-800mm Canon EF 镜头拍到了迄今最好的 Morris 照片,它们很喜欢待在海港那块牌子下面。

Bloomberg Technology

Goldman’s Moe Says He’s in ‘Stronger-for-Longer Camp’ on AI

Goldman Sachs chief US equity strategist David Kostin says he's in the 'stronger-for-longer' camp on AI, expecting infrastructure investments to yield returns for years. The post does not disclose specific timelines or return figures, but bases the view on corporate capex trends and productivity expectations.

Latent Space

Runway’s WorldPrompt: Prompting Real-Time AI-Generated Worlds

Runway added WorldPrompt to GWM Worlds 2, letting you steer characters, cameras, and environments with natural-language timestamped events. CTO Kamil Sindi calls it “promptable worlds on-demand with video and audio in sync,” but researcher Robin Kahlow notes it’s still a research preview—movement works reliably, complex actions don’t always follow. The engineering challenge is twofold: generate frame-by-frame instead of a whole clip at once, and make generation fast enough for real-time playback. The post doesn’t disclose exact latency figures. It compares competitors: Google DeepMind’s Genie 3 runs at 720p/24fps for a few minutes max; Odyssey-2 Pro and World Labs’ RTFM are also in the race. Runway is valued at $5.3 billion and shipped the first GWM Worlds last December.

Computing Life · Share · Yage

Four report cards in one week: Step 5 Preview benchmarks, Jev's confidence blind spot, and what California's AI order actually requires

StepFun's Step 5 Preview posted four types of scores, but its self-reported Terminal-Bench result isn't on the public leaderboard yet, and weight license terms remain undisclosed. An independent dev tested TypeSafe's Jev classifier on 111 public hard questions: Jev's average confidence was 0.690 when wrong and 0.684 when right—confidence doesn't separate correct from incorrect. California's AI executive order only directs state agencies to draft legislative proposals; the only active law binding companies is 2025's SB 53.

Computing Life · Share · Yage

Three old authorizations, two days, into OpenAI's internal repo

Security team Hacktron exploited a known libheif memory bug via OpenAI's public forum image upload, gained forum admin, then pivoted through OpenAI's SSO to take over an internal engineer's ChatGPT and Codex accounts. The engineer had previously authorized Codex on their personal GitHub, allowing the team to create a branch and submit a pull request in the core openai/openai repo—no source code was read, no customer data touched. OpenAI fixed the issue ~14 hours after the report and paid a $6,500 bounty covering only the SSO finding; the forum itself was excluded from scope. The entire chain used existing configurations: the image parsing flaw stemmed from a libheif code change from a year earlier, still unpatched in Debian's old stable branch; trust propagation came from the forum unconditionally relying on centralized SSO; repo write access came from the engineer's routine Codex authorization. Claude Opus 5 helped compress exploit-writing from days to hours after humans had already pinpointed the root cause and set up the debugging environment—it did not autonomously discover the vulnerability.

Why it matters: Hacktron went from a public forum image upload bug to creating a branch in OpenAI's internal repo—a concrete attack chain with a timeline and fix record, not a proof-of-concept. All three HKR axes hit: compelling narrative, solid technical detail, and direct relevance to pract...

Simon Willison

Note on 24th September 2026

Simon Willison 表示,与编码智能体协作越久,越确信它们让软件工程变得更难。借助智能体可以完成惊人的工作,但释放其全部潜力需要极高的纪律性和知识储备。

Bloomberg Technology

US Treasury Yields Hit 5% as Wall Street Faces New Rate Regime

The 10-year US Treasury yield breached 5%, marking a new rate regime that lasts 'until something breaks.' Bloomberg reports investors are repricing assets as high rates persist, directly impacting AI startup funding costs, tech valuations, and venture capital pacing—cheaper capital is gone, and burn-for-growth models face a tougher test.

Product Hunt · AI

Once UI 2.0: Open-source design system for humans and AI agents

Once UI 2.0 is an open-source design system for building consistent React apps. The new version adds component catalogs, compact rules, and task guides to help AI coding agents use the system correctly, while keeping APIs predictable for human developers. The post doesn't disclose pricing or license details.

Product Hunt · AI

Basedash MCP write: build charts and dashboards from Cursor and Claude

Basedash's new MCP write lets you build charts and dashboards directly inside Cursor or Claude, no tab-switching needed. The post doesn't specify supported data sources or chart types, but the pitch is clear: skip copy-paste and visualize inside your AI editor. A handy utility for data teams and product dashboards.

Hacker News front page

Vibe Coding Production Kit: a production workflow for AI coding agents

This GitHub repo offers a full workflow from idea to production for teams using AI coding agents like Copilot. It covers specs, architecture, testing, security, code review, and CI/CD, claiming to be battle-tested. The post doesn't include benchmarks or user stories, so you'll have to try it yourself.

Bloomberg Technology

Anthropic's Gene Editing Discovery Isn't Yet a Breakthrough, Scientists Say

Anthropic's biology lab used AI to discover a new enzyme, but scientists urge caution—it's not a breakthrough yet. The article does not disclose the enzyme's function, experimental validation details, or potential applications. What's clear: Anthropic made an AI-driven discovery in biology, but external experts say more verification is needed.

Hacker News front page

LaunchVideo turns a URL or prompt into an explainer video with Opus 5.5 and a headless renderer

LaunchVideo generates a ~30-second product explainer from a URL or a text prompt. Opus 5.5 writes the HTML/CSS/animation script, and a serverless agent renders it frame by frame in a headless Chromium microVM — no video generation model is used. Each video costs roughly 100k tokens and takes about four minutes, outputting 1080p 30fps MP4 with a virtual clock for deterministic frames. The page shows five unedited examples including NVIDIA and Linear. The whole product is one TypeScript agent file plus three tools, fully open-source and one-click deployable to your own OpenComputer account. The post doesn't mention pricing or whether models other than Opus 5.5 are supported.

Why it matters: A clever packaging of Opus 5.5's coding ability into a 'URL-to-launch-video' tool, with a clearly explained pipeline and visible examples. But the product is still lightweight—more a sharp demo than an industry-shaking release. H and K both hit, R is weak, landing right at the...

Simon Willison

commit-rewriter 0.2

Simon Willison 发布 commit-rewriter 0.2。该工具与 git 相关,具体功能与更新细节原文未作说明。

Bloomberg Technology

Anthropic Strikes $12 Billion AI Computing Deal With Akamai

Anthropic signed a five-year, $12 billion cloud deal with CDN giant Akamai. It's Anthropic's first major compute commitment outside AWS, aimed at diversifying away from Amazon. Akamai shares rose 8% after hours. The post doesn't disclose GPU counts, delivery timelines, or detailed contract terms.

Why it matters: Anthropic diversifying its core compute away from AWS for the first time, with a $12B, five-year contract, is a deal that reshapes the cloud power map. Bloomberg's exclusive carries source authority, and Akamai's 8% after-hours jump confirms the market is pricing this in. The ...

GitHub Blog · AI & ML

When chat is the wrong UI

GitHub Copilot 应用推出 canvas,一种运行在应用内、无浏览器外壳的全栈小应用,可与 Copilot 智能体双向通信,并能在本地执行代码、调用第三方 API。作者认为聊天只是 AI 的通用兜底界面,用户明确任务时更该让智能体生成可复用工具,而非把智能体本身当工具、白白消耗 token。示例包括 Connect 4 游戏、Winget 包管理、SQLite 操作和开发工作流自动化。

The Verge · AI

Google gives Gemini 3.8 Live an animated, lip-synced face

Gemini 3.8 Live now lets you talk to an animated avatar that lip-syncs and reacts in real time. It's only available to Enterprise customers for now. Google says it handles 97 languages without degrading video fidelity or introducing visual drift, and a demo shows mouth movements matching both English and Japanese. The post doesn't say when individual users might get access.

Google Research Blog

Google tackles coherent long-form video generation

Google published research on automating long-form video generation, focusing on coherence across scene transitions. The post doesn't disclose model architecture or max video length, only that the system plans shots and maintains character/background consistency. For video generation or AI filmmaking practitioners, this is Google's first long-form answer post-Sora, but technical details are thin—take it with a grain of salt.

Financial Times · Technology

SoftBank pays a steep premium on a record $9bn bond sale to fund its OpenAI bet

SoftBank just sold a record $9bn bond to fund its OpenAI bet, but had to pay 0.25–0.5 percentage points more in interest than comparable peers. The premium reflects market concern over its debt load and Masa Son's concentrated wager. Proceeds will first refinance existing debt, with the remainder going to OpenAI. The post doesn't spell out the exact split between refinancing and new investment.

Why it matters: SoftBank's record $9B bond sale to fund its OpenAI bet came with a 0.25-0.5pp rate premium — the bond market is pricing in concern about the concentrated wager. FT exclusive with concrete pricing data; HKR all hit. Not scoring higher because the post doesn't disclose the split...

Hacker News front page

AI labs need to start funding historical research

The author tested GPT-6 Sol and Opus 5.5 on two historical problems: decrypting a 1941 Enigma message—where the model independently located supplementary records from the German Federal Archives—and tracing a Latin alchemical passage by Isaac Newton back to a previously unidentified French source. He argues frontier models can now deliver verifiable results on codebreaking, cross-language text tracing, and linking findings across niche subfields, a leap from last year's assistant-level performance. The post does not specify a collaboration framework or funding figures, but points to digitized, falsifiable historical problems as the sweet spot.

Why it matters: The author demonstrates frontier models' real capability in codebreaking and cross-lingual text tracing with two verifiable cases. But the topic is academic history, which limits resonance with AI industry readers, so the score sits right at the featured threshold.

TechCrunch · AI

PrismML brings its tiny LLMs to Qualcomm-powered smart glasses

PrismML showed a 1-bit Bonsai LLM at Qualcomm's Snapdragon Summit that runs locally on smart glasses using the Snapdragon AR1 Gen 1 platform. The 2-billion-parameter model handles vision and language so wearers can ask about what they see in real time. PrismML shrinks larger models by 4x while keeping nearly all benchmark performance. The startup's bigger aim is open-weight on-device AI that doesn't depend on cloud labs' privacy promises. No smart glasses shipping with PrismML have been announced yet.

AI HOT (Curated Pool)

GitHub Security Lab launches LLM-powered fuzzing agent for C/C++ projects

GitHub Security Lab released Taskflow Agent, an LLM-driven tool that automates the full fuzzing pipeline for C/C++ projects. It writes test cases, compiles, runs the fuzzer, and generates reports on crashes. The post doesn't specify which LLM is used or how many real-world bugs it has found.

The Verge · AI

Jensen Huang: AI will fight climate change — after causing pain first

On The Ezra Klein Show, Nvidia’s CEO argued AI can fight climate change — but only after inflicting “an enormous amount of pain and suffering.” It’s the same accelerationist pitch from tech leaders and President Trump: data centers running on dirty energy are worth the damage. The post doesn’t disclose specific energy numbers or timelines, but highlights Huang’s glaring privilege.

The Verge · AI

Meta lets you build Horizon games with AI prompts on your phone

Meta announced Horizon Create (mobile) and Horizon Studio (browser), both using AI prompts to build games. Published games will also be recommended and playable on Facebook and Instagram. Early access is waitlist-only. The post doesn't disclose launch dates or which AI model powers the tools.

Hacker News front page

Critic: the coding agent that wrote the code also reviews it

Critic puts agent-authored code and its review in one view, with a change narrative and threaded Q&A. The demo PR shows agents Codex and Claude reworking an offline message queue: lease tokens replace a heartbeat table, and a disconnected device returns its work after 45 seconds without losing context. The post doesn't disclose pricing or self-hosting options.

Why it matters: Putting AI code generation and review in the same interface, with a change narrative and a threaded Q&A, is a fresh product shape. The demo PR's lease-token-over-heartbeat design also surfaces a reusable engineering pattern. Score isn't higher because the post doesn't disclose...

Hacker News front page

A Million Agents Is a Distributed Systems Problem

InstaCloud's CTO frames agent scaling as a distributed systems problem. Google Research's 180-config study found multi-agent setups boosted parallel tasks by 80.9% but hurt sequential reasoning by 39–70%. Uncoordinated agents amplified errors 17.2×; an orchestrator cut that to 4.4×. In ACL 2026's Silo-Bench, teams of 2–100 agents talked a lot but reasoned poorly, with zero success on the hardest tasks at 50 agents. The takeaway: persist state, not the agent—schedule agents like processes and recover them like nodes.

Why it matters: Reframes agent scaling as a distributed systems problem, backed by Google Research data on 180 configs — not just opinion. Docked because it's a vendor blog (InstaCloud) with product incentives, and the post doesn't link to the paper or disclose experimental details. Lands at ...

The Verge · AI

Meta's Muse reportedly lets users download its entire filesystem

Two developers independently got Muse to zip and share its full root filesystem, Ubuntu system files, app templates, and internal docs with minimal prompting. Saunders called it extremely easy to replicate and noted almost no prompt injection resistance. Meta says it's not a breach since each user runs in a persistent Linux VM.

AI HOT (Curated Pool)

Google Cloud API Gateway now exposes existing REST APIs as MCP tools

Google Cloud API Gateway enters public preview with native MCP support. Add x-google-api-management.mcp: true to an OpenAPI spec, deploy, and the gateway acts as a remote MCP server—no separate server needed. Existing JWT, API-key auth, and quota policies apply uniformly to both MCP and REST traffic. The tools/list discovery endpoint is unauthenticated by default; the post recommends securing it with JWT in production. Connect any MCP client to the gateway's /mcp path; the ADK example uses McpToolset with StreamableHTTPConnectionParams.

TechCrunch · AI

Google Photos launches 'Clueless'-style virtual closet, AI organizes your outfits

Google Photos now has an AI virtual closet that identifies clothes from your photos and organizes them into a digital wardrobe. It launched on Android in June and is now available on iOS. Inspired by Cher's virtual wardrobe app in the movie 'Clueless.' The post doesn't specify the model or training data, but it's a consumer CV application worth noting.

AI HOT (Curated Pool)

Anthropic launches Claude Opus 5.5, optimized for cost in long-context coding sessions

Anthropic released Claude Opus 5.5, explicitly targeting cost reduction for coding sessions that run long and use heavy context. The post body only contains the title and site navigation; it does not disclose pricing, benchmarks, or context-window specs. The one confirmed takeaway is the cost-optimization angle for extended coding workflows—everything else is still missing from the article.

Why it matters: Anthropic model launch is a signal, but the body is just a title and nav bar — all key facts are missing. H and R hit, K doesn't. Barely clears the featured threshold (≥2 of 3), but thin content caps the score at 72, the featured floor.

TechCrunch · AI

ElevenLabs CEO on margins, IPO timing, and telling customers they’re talking to a bot

ElevenLabs, now reportedly valued at $22B, CEO Mati Staniszewski discusses margins, IPO timing, and whether businesses should disclose AI voice on calls. He says disclosure is appropriate for now, but may become unnecessary as AI calls become the norm. The post does not disclose specific margin figures or IPO timeline.

Google DeepMind

Google DeepMind releases Gemini 3.8 Live with Live Avatar

Google DeepMind released Gemini 3.8 Live with Live Avatar, adding near-real-time video generation to its native real-time conversation model. The result is a dynamic visual avatar with lip sync, natural expressions and smooth turn-taking.

Why it matters: The post details Live Avatar's real-time video conversation, async tool calls and 97-language support, a useful read on enterprise multimodal interaction.