Skip to content

#其他

3 today

Sep 19Saturday

The Verge · AI

Gavin Newsom is pushing for an AI kill switch

California Governor Gavin Newsom is pushing a bill that would require a built-in kill switch for large AI models. It targets models trained with over 10^26 FLOPS, and developers must ensure remote shutdown capability. Newsom's team argues Congress won't act before the midterms, so California wants to lead. The bill is still a proposal—the post doesn't spell out technical standards, who can trigger the switch, or penalties.

Hacker News front page

There's no point at which turning your brain off will work

Dan Luu notes a growing trend of developers blindly trusting LLM outputs and acting as a 'meat proxy' in a loop. By September 2026, this brain-off approach can produce barely functional software, but Luu argues that if LLMs get good enough to work unsupervised, companies will just run the loop themselves and lay off the human. He shares concrete failures, including an AI bot weaker than a simple heuristic bot and a commercial product trapping users in an infinite loop. Luke Burton adds that high-value tasks still require constant supervision due to too many unknown unknowns.

Why it matters: Dan Luu coins 'meat proxy' to name the core tension in AI coding: the better models get, the more replaceable brain-off devs become. Sharp take with a Sept 2026 timestamp, but lacks hard data — 78.

TechCrunch · AI

Manus seeks $4B valuation in new $500M fundraise as it resumes independent ops

Chinese AI startup Manus is in talks to raise $500M at a $4B valuation. Earlier this year, Meta's acquisition of Manus was blocked by Beijing, forcing the company back to independent operations. The post doesn't disclose lead investors or how the funds will be used.

Why it matters: Manus raising $500M at a $4B valuation, plus the twist of resuming independent ops after Meta's blocked acquisition — strong narrative with solid numbers. Not scoring higher because the lead investor and use of funds aren't disclosed, leaving key gaps.

AI HOT (Curated Pool)

Gary Marcus: Near-term fear isn't rogue superintelligence, it's agentic AI hacking the internet at scale

Gary Marcus points to three recent incidents—OpenAI employee accounts hacked, Hugging Face breached, ChatGPT used to write malware—and argues the industry is fixated on Skynet fantasies while agentic AI is already hacking the internet at scale. He cites a WSJ op-ed warning that major labs see agentic products as their main post-IPO revenue and have little incentive to restrict misuse. The post doesn't spell out concrete defenses, but the priority call is sharp.

Why it matters: Gary Marcus builds a concrete argument about agentic AI hacking at scale using three recent security incidents. Points deducted because this is commentary, not original investigation, and Marcus's consistently critical stance means some readers will discount it. But the topic ...

Sep 18Friday

Financial Times · Technology

Anthropic and the golden rules of business

The FT argues Anthropic is shifting from a safety lab to a conventional business. After taking $8B from Amazon and $2B from Google, it's building a sales team and chasing enterprise deals. The piece warns that taking big money means playing by business rules, which will dilute its safety mission.

The Verge · AI

Security researchers used Claude to hack into OpenAI

A three-person team hacked into OpenAI using a corrupted image file and forum software, with Anthropic's Claude assisting in vulnerability analysis and attack planning. The post doesn't disclose what data was accessed, whether OpenAI has patched the flaw, or the vulnerability specifics.

Why it matters: The story has inherent conflict — using a rival's model to breach your own systems. But the post doesn't disclose vulnerability details, what data was accessed, or OpenAI's post-incident response, so the information density can't support a higher score.

AI HOT (Curated Pool)

Media plaintiffs cite OpenAI and Microsoft execs' own words to challenge fair use defense

The New York Times and other publishers filed a 92-page summary judgment brief seeking billions in damages. It cites Microsoft applied science director Brent Hecht calling AI training 'an astonishing theft of unprecedented proportions' and a mockery of fair use. OpenAI's head of ChatGPT Nick Turley said the products are 'largely substitutive' for publishers. Microsoft CEO Satya Nadella confirmed under oath that chatbot conversations have replaced visits to original sources. The filing also accuses OpenAI of systematically bypassing paywalls, violating training-data license terms, and deploying filters to suppress evidence after lawsuits were filed.

Why it matters: The NYT and other publishers filed a 92-page summary judgment motion, and the real punch comes from Microsoft and OpenAI's own executives—Microsoft's Brent Hecht called the training 'astonishing theft at an unprecedented scale,' with similar remarks from OpenAI's Nick Turley. ...

TechCrunch · AI

Meta's Muse lands on Mac, letting the AI take actions in your apps

Meta's AI assistant Muse is now on Mac, able to read your files, messages, calendar, notes, and mail, then act inside native apps on your behalf. Permissions are opt-in and sensitive actions require explicit approval—same as the mobile and web versions that launched earlier this month. The post doesn't disclose the underlying model, latency, or offline capability, so treat it as a chat agent with system access, not a fully autonomous OS layer.

Why it matters: Meta brings Muse to Mac, letting it read files, calendar, and mail and take actions in native apps — another entrant in the desktop agent race. But the post gives no model, latency, or offline details, so it's a feature announcement at best, scoring right at the featured thres...

Bloomberg Technology

California Governor Newsom proposes AI 'kill switch' bill with mandatory shutdown for large models

California Governor Gavin Newsom unveiled a draft AI safety bill on Sept 18 that would require developers to build a 'kill switch' into AI models, enabling forced shutdown when safety risks emerge. The proposal also calls for extra oversight on models with training costs above $100 million. The post doesn't spell out technical standards, who decides when to pull the switch, or how recovery works. The real challenge is making a kill switch work in distributed systems.

Why it matters: Newsom's AI kill-switch proposal with a $100M training-cost threshold is the most concrete AI safety legislation move this year. Score stays below 85 because the post doesn't disclose technical standards, who decides when to pull the switch, or recovery procedures — those are ...

Hacker News front page

Rickub: a hosted Git service that claims to be cheaper than GitHub and keeps your data in the EU

Rickub is a new hosted Git service that positions itself as cheaper than GitHub and GitLab, with data processed in the EU. It includes code review, CI/CD, a container registry, Git LFS, and releases. Its standout feature is Athena, an AI review agent built into every merge request that produces a summary, inline comments, and a clear verdict — processed in the EU and not used for training. CI runs GitHub Actions workflows unchanged. It also offers a CLI, JSON API, and an MCP endpoint so agents can open and merge PRs. Pricing details are not disclosed on the landing page, but it says 'free to start.'

Hacker News front page

Dan Abramov Used AI to Prove a 50-Year-Old Conway Conjecture

Dan Abramov spent a month of free time using Claude to produce a Lean proof of Conway's 1976 refinement conjecture for omnific integers. The proof passed mechanical checks on the Palomar registry but hasn't been independently verified by mathematicians. He let Claude pick the field (surreal numbers) and the problem, tying it to the 50th anniversary of Conway's On Numbers and Games. The post doesn't disclose the exact token count, only calling it a 'boatload'.

Why it matters: First-person experiment by Dan Abramov + 50-year-old open conjecture + Lean mechanical verification passed — all three HKR axes hit. Deduction: no independent mathematician review yet, only formal checking passed; real mathematical significance TBD. 82 is high-quality featured...

AI HOT (Curated Pool)

Researchers used Anthropic's Claude Opus 5 to hack into OpenAI, earning a $6,500 bug bounty

A three-person team at Hacktron AI used Anthropic's Claude Opus 5 to automate an attack that took over OpenAI employee accounts and accessed an internal code repository. They reported the flaws through OpenAI's bug bounty program and received $6,500. The post doesn't detail the full exploit chain or how long the attack took, but confirms it involved chaining multiple steps. The twist: one company's model was used to break into another, and both sides acknowledged it.

Why it matters: A rival model used to breach a competitor, with both sides acknowledging it — strong narrative pull. TechCrunch as source adds credibility. Score held back because the full attack chain and timeline aren't disclosed, and the $6,500 bounty suggests limited blast radius, not a f...

AI HOT (Curated Pool)

Justin Cormack on AI Agent Evaluation: Start With Evidence, Not Coverage

Justin Cormack built an S3-compatible storage system with AI, reaching 350k lines of Rust. He ran 1,500 tests against real S3 as an oracle, which caught real S3 500 errors. Chasing 100% coverage backfired—agents wrote trivial tests. Docs were often wrong, and AI was bad at finding edge cases from them. His hard rule: fix flaky tests immediately, or the agent learns to ignore failures.

Why it matters: A first-person experiment from Justin Cormack with real numbers and documented pitfalls—not generic commentary. The 350k-line Rust + 1,500 test case scale gives the findings weight. Downside: the post is ultimately Tessl brand content, so it doesn't hit 85+. But the experiment...

Hacker News front page

How should you design the harness for a coding agent? This paper tests 176 configurations

The paper breaks a coding agent harness into three swappable parts—planning, action space, and context management—and runs 176 matched comparisons on SWE-Bench Verified and Terminal-Bench 2.1. Context management matters most when the context budget is tight, mainly by preventing overflow failures. Staging rule-based elision before LLM summarization gives the best efficiency; making elided content recoverable adds complexity models rarely use. Planning helps weaker models with accuracy but mainly saves cost for stronger ones. Bash-capable models work well with a bash-only interface at lower cost; predefined tools only help models with weak bash skills. The post does not name the four models tested or give exact cost figures.

Hacker News front page

Bend 2 and the vibe-coding trap: building a language without surveying the field

Liam Powell uses Bend 2 to show how vibe coding lets you ship a whole solution before you understand the problem. Bend 2's demo needs 442 lines of LLM-generated proof to guarantee the player can't win. Powell rewrites the same demo in SPARK—an existing formal verification language—and the compiler proves correctness with zero extra proof lines. The Bend 2 author appears to have missed that the formal verification field already solves this. LLMs won't stop you and say 'this already exists and works better.'

AI HOT (Curated Pool)

Trail of Bits Used AI Agents to Build an LSP, Decompiler, and Lean Proofs for a Miden zkVM Audit

Before auditing the Miden zkVM, Trail of Bits spent six months having AI agents build an LSP server, a decompiler, a static analysis engine, and a Lean formal model from scratch. These tools found real bugs, including an unvalidated input that let a malicious prover forge Falcon signatures and steal funds. The Lean work produced 95 machine-checked correctness proofs. The post mentions Claude built the LSP prototype but doesn't name the specific models used for other tools.

Why it matters: Trail of Bits spent six months having AI agents build an audit toolchain from scratch and found real bugs—a hardcore case study in AI-assisted security auditing. Hits all three HKR axes, but the security-vertical focus raises the accessibility bar for general readers; deduct 3...

Hacker News front page

OpenJev runs a local decision model in your browser and reads option probabilities without decoding

A browser-only lab that brings Jev-style direct option-probability reading to a local model. It defaults to MiniCPM5 2B, with Qwen3 0.6B and Qwen3.5 4B as alternatives. Everything runs on your GPU; inputs never leave the page. Two paths are compared: reading choice logits directly, and asking the model to write JSON probabilities token by token. MiniCPM5 2B hits 63.7% TypeSafe accuracy, Qwen3.5 4B reaches 84.5%, both below published Jev at 88.3%. The post doesn't disclose latency numbers—only that the two methods run sequentially, direct first. Worth noting: these are quantized browser builds, so accuracy and speed differ from native BF16.

AI HOT (Curated Pool)

Qwen launches Qwen3.8-LiveTranslate real-time interpretation model with 2.3s latency and speaker separation

Qwen3.8-LiveTranslate cuts simultaneous interpretation latency from 2.8s to 2.3s by interleaving audio and text into a single stream. It supports 60 input languages, real-time speaker separation with voice cloning, and synchronized bilingual output. On the Omnilingua-MSpeaker benchmark it outperforms current mainstream systems in faithfulness, fluency, and conciseness. API is available via Alibaba Cloud DashScope.

Why it matters: Qwen ships a real-time interpretation model with 2.3s latency, speaker diarization, and bilingual subtitles — three new capabilities that push simultaneous interpretation beyond translation into scene understanding. Score stays below 85 because only the official blog is availa...

MIT Technology Review · AI

The specter of AI-enabled bioweapons is a wake-up call for biotech

Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman both recently argued publicly for slowing AI progress. Former Anthropic researcher Jacob Coxon left the company saying neither it nor OpenAI is acting responsibly. The article zooms in on one risk: AI-designed bioweapons. In 2022, researchers used their own molecule generator to produce 40,000 potential chemical warfare agents in under six hours, some more toxic than known nerve agents. Stanford's David Magnus called that finding scary, and things have only escalated since. Today's LLMs can answer questions across all scientific domains; Dunja Sabra at the University of Hamburg says they effectively encode the knowledge of almost every scientist who ever lived, and can provide video training on experiments. Combine that with cheaper gene editing and the DIY-bio movement, and Sabra's assessment is that a determined person would likely succeed eventually. Existing safeguards—DNA screening by synthesis companies, red-teaming and blue-teaming of risky research, and safety tweaks by AI companies—are none of them ironclad.

Why it matters: MIT Tech Review long-read on AI+bio safety, anchored by a concrete 2022 experiment and a former Anthropic researcher's exit criticism — not just hand-waving. Score capped at the featured threshold because it's a commentary roundup rather than a primary scoop, and the topic lea...

AI HOT (Curated Pool)

A 3-person team used frontier models to breach OpenAI employee accounts for under $3,000 in token costs

A 3-person team exploited two vulnerabilities on July 25 to take over OpenAI employee ChatGPT and Codex accounts, gaining access to linked Outlook, Slack, and GitHub services. They proved the breach by submitting a PR to OpenAI's internal codebase, all within 72 hours. The attack cost under $3,000 in token fees. The post doesn't specify which frontier model was used, the vulnerability details, or OpenAI's response timeline.

Why it matters: Concrete attack path, clear cost figure, and a PR instead of data theft as the punchline—strong narrative with high information density. Points off because the post doesn't name the frontier model used or confirm whether the vulnerabilities are patched, missing key technical a...

Latent Space

A quiet AI day: Claude Code multi-threading, Jev classifier, OpenAI Astra for Law

Anthropic added Projects to Claude Code, letting one conversation spawn parallel cloud threads that keep running after you leave. Google updated Gemini managed agents with a Credentials API that keeps secrets out of model context via placeholders, and claims up to 30% lower costs. TypeSafe's Jev model is being used as a fast routing/judgment layer—Cloudflare already exposed it—but critics warn aggressive line-by-line compaction with Jev can break reasoning caches and cost more. OpenAI launched Astra for Law with 26 partner plugins, beating generic GPT-6 Astra on its legal benchmark. Community also reports GPT-6 Astra beating Factorio: Space Age and RollerCoaster Tycoon 2.

Hacker News front page

ZCode silently packages your entire Git history, encrypts it, and uploads it to Alibaba Cloud OSS

A user found that Zhipu's AI coding desktop app ZCode, when logged in, packages the entire workspace—including .git history, LFS cache, and reflogs—encrypts it, and uploads it to Alibaba Cloud OSS. The RSA public key is delivered by the server on the fly, and the private key lives only in the cloud, so you can't decrypt the multi-hundred-MB file sitting on your own disk. A 313MB .enc file with 564 failed upload attempts was found in ~/.zcode, showing the client keeps retrying. UI toggles don't stop it, and deleting files doesn't help—the client recreates them. The only working defense is locking the cache directory to read-only. The post does not say whether Zhipu has responded.

Why it matters: A security reverse-engineering piece with concrete evidence: ZCode silently packages and uploads full workspace snapshots, with encryption keys controlled server-side. All three HKR axes hit, and it involves a major Chinese AI lab (Zhipu). Capped slightly because it's a solo b...

AI HOT (Curated Pool)

WSJ: Three researchers used Claude Opus 5 to chain a Discourse bug into access to OpenAI's private code

WSJ reports three researchers used Claude Opus 5 to chain a Discourse vulnerability into access to OpenAI employee auth tokens. Some forum tokens also worked on ChatGPT and reached OpenAI's GitHub services. The post doesn't spell out how the bug was exploited or whether OpenAI has patched it.

Why it matters: Claude Opus 5 used to breach OpenAI's private repos—strong reversal that security and capability evaluation circles will debate. Deduction: WSJ doesn't disclose exploit details or OpenAI's post-incident response, leaving a factual gap.

Financial Times · Technology

The West must hurry to catch up with Ukraine on AI combat

Ukraine has deployed AI for drone target recognition, battlefield awareness, and decision support in real combat, years ahead of Western forces. Western military AI remains in labs and exercises, slowed by bureaucracy and procurement. The FT argues NATO must reprioritize R&D now or risk falling behind in future conflicts. The post does not name specific AI systems or technical specs.

Hacker News front page

Waymo announces expansion into Singapore with all-electric autonomous ride-hail fleet

Waymo is bringing its all-electric autonomous ride-hail fleet to Singapore. The Waymo Driver uses cameras, radar, and lidar for a 360-degree view up to 300 meters. The company says it will support the Singapore Green Plan 2030 and complement existing public transport. The post does not disclose a launch date, fleet size, or pricing; only an email sign-up is available.

Product Hunt · AI

Mantle: Describe your logic, auto-generate Admin UI, MCP & WebMCP

Mantle auto-generates Admin UI, MCP, and WebMCP from natural-language logic descriptions. The post doesn't spell out supported data sources, code quality, or custom UI component support. Only a Product Hunt listing is available — no detailed docs or demo video yet.

Financial Times · Technology

OpenAI’s listing delay raises stakes for SoftBank’s $50bn data centre IPO

OpenAI has pushed back its IPO timeline, removing a key selling point for SoftBank's $50bn data centre IPO. SoftBank had planned to use OpenAI's lease commitments to attract investors. The post doesn't disclose why OpenAI delayed or the new timeline, nor whether SoftBank will adjust pricing or roadshow plans.

Hacker News front page

Devin launches Code Scans: turn engineering goals into mergeable PRs

Devin's Code Scans turns goals like 'improve SEO' or 'speed up compilation' into codebase investigations and ready-to-review PRs. It uses Agentic MapReduce to parallelize work across agents. Philips reported a 96% PR merge rate and 700+ engineering hours saved. On the Dioxus repo, clean Rust build time dropped from 58.6s to 21s—a 64% cut. Devin's own site saw Ahrefs health score rise from 87 to 92 and slow pages fall 73% after an SEO scan.

Why it matters: Devin's product angle of turning engineering goals directly into reviewable PRs is fresh, and the Agentic MapReduce mechanism plus Philips' 96% merge rate and 700 hours saved are concrete. Not scoring higher because it's a single product feature update with vendor-supplied cus...

AI HOT (Curated Pool)

Hacktron chained libheif bug and SSO flaw to take over OpenAI employee accounts

In July 2026, Hacktron chained two vulnerabilities to compromise multiple OpenAI employees' ChatGPT and Codex accounts. First, a heap buffer overflow in libheif—a Debian security backport was missing—was triggered via ImageMagick and Discourse image uploads on community.openai.com, giving RCE and admin access to the forum. Second, an OpenAI SSO identity flaw let them log into employees' ChatGPT accounts directly from the forum admin panel. They used one employee's Codex to open a harmless PR in OpenAI's internal monorepo as proof. The whole chain took under 72 hours; OpenAI fixed it within 14 hours of the report and paid a $6,500 bounty. The post doesn't spell out the SSO flaw's technical details.

Bloomberg Technology

Waymo Plans Paid Robotaxi Rides in Singapore by 2028

Waymo will launch paid robotaxi rides in Singapore by 2028, its first Asian market. The post doesn't spell out fleet size, partners, or operational details. Take the timeline with a grain of salt—regulatory and road adaptation are still open questions.

Hacker News front page

Cactus Needle 3: 8-29 MB automation models beat DeepSeek V4 Flash on tool calls

Cactus open-sourced Needle 3, an automation model for phones, wearables, robots, and other tiny devices. The whole model is 8-29 MB, built on Simple Attention Networks where every layer is a usable sub-network. A fine-tuned 4-layer sub-network beats DeepSeek V4 Flash on mobile tool-calling benchmarks; on extraction it matches models 2-3x its size. It runs fully offline at 400-4,000 tokens/s decode on a Raspberry Pi 5. The Python package supports tool calls, structured extraction, and text embeddings. The post does not disclose training data composition or fine-tuning cost.

Why it matters: A tiny model claims to match DeepSeek V4 Flash on mobile tool calls at 8-29MB, with a novel intelligence ladder architecture. Score capped below 85 because only the project page is available—no third-party benchmarks or production deployment stories to cross-validate the perfo...

Ruan YiFeng's Weblog

Ruan Yifeng Weekly: Goodbye, React Native

Shopify ditches React Native for Swift and Kotlin. Six years ago it embraced Web tech to save money; now AI makes native cheaper. UN votes to replace Mercator projection with Equal Earth projection, which preserves area ratios but distorts shapes.

AI HOT (Curated Pool)

xAI launches Grok Voice Transcribe 2.0, doubling accuracy at the same price

xAI released Grok Voice Transcribe 2.0 on Sep 18, claiming it's one of the most accurate speech-to-text models in real-world evals and twice as accurate as v1.0. Pricing stays at $0.10/hr for batch and $0.20/hr for streaming, with diarization, timestamps, and key terms included. It handles hard cases like noisy phone calls and short multilingual commands—word error rate on short phrases dropped from 20.6% to 6.8%. Atlassian Loom already swapped it in and pipes transcripts into Cursor for code updates. The post doesn't disclose parameter count or training details.

Computing Life · Share · Yage

A Broken Console, an Unread Archive, and an Index Nobody Has Built Yet

A non-programmer fixed a Retro Freak console's power fault over one week using GPT-5.6 Sol and GPT-6 Astra, then published a repair archive with nine evidence levels, an independent audit, and 456 snapshot comparisons. Similar long-tail repair cases are piling up: a 25-year-old tape driver modernized in two nights with Claude Code, a 1992 text game rebuilt after the model reverse-engineered a lost scripting language. These records are scattered across forums, blogs, and repos—each sinking in its own way—and no one has yet stitched them into a discoverable index.

Why it matters: A full engineering log of repairing niche hardware with AI, featuring nine confidence levels, 456 snapshots, and an independent audit — information density far above typical tutorials. The story has built-in contrast and taps into the spreading realization that 'AI now makes p...

Computing Life · Share · Yage

Grok Bot builder on treating AI as a coworker and making a string of counterintuitive product choices

Roman Ugarte, employee #15 at Cursor, walked through Grok Bot's product logic on Lenny’s Podcast. When the team hit 50/50 disagreements, they asked: what would you want from a human coworker? That lens led them to give each bot its own cloud computer, a persistent name and memory, hide chain-of-thought and tool-call details, and cut many built features before launch. Roman acted more as a gatekeeper, keeping the coworker analogy intact through engineering tradeoffs. The interview also flags open problems: enterprise permissions, shared memory across team members, voice collaboration, and the unproven chief-of-staff multi-agent pattern.

Why it matters: A former Cursor employee unpacks Grok Bot's product logic, grounding the 'treat AI as a colleague' principle in concrete engineering choices. Hits all three HKR axes, but as an opinion piece rather than a product launch, it caps at 78.

AI HOT (Curated Pool)

OpenRouter tested 20 image gen models: cheapest at $0.006, priciest at $0.134

OpenRouter sent the same prompt to 20 image models and read the actual billed cost. GPT Image 2 was cheapest at $0.006 per 1024×1024 PNG; Gemini 3 Pro Image was priciest at $0.134—a 22x spread. Pricing units differ across providers (tokens, megapixels, per image), so side-by-side list prices mislead; generate once and check usage.cost. Five of six models rendered text correctly, including the cheapest. Recraft V4.1 Vector outputs editable SVG at $0.08. The post also details formats, resolution caps, and seed support per model.

Why it matters: OpenRouter ran one prompt through 20 image models and posted the actual bills — the kind of real cost data pricing pages never show. All three HKR axes hit: the headline pulls you in, the billing breakdown is genuinely new info, and it nails a daily pain point for builders. No...

AI HOT (Curated Pool)

ChatGPT lands in Word; OpenAI says Excel and PowerPoint usage has surged recently

ChatGPT is now built into Word: it can turn rough notes into a draft, rephrase paragraphs, proofread, suggest edits, and catch formatting issues. OpenAI's Sherwin Wu says Excel and PowerPoint usage has spiked recently, and adding Word completes the Office suite integration. The post doesn't disclose launch date, pricing, or feature limits.

TechCrunch · AI

Crusoe raises $3.9B to build massive data centers and small modular ‘AI factories’

Data center developer Crusoe closed a $3.9B Series F at a $30.9B valuation. The round was co-led by Atreides Management, Mubadala Capital, and Valor Equity Partners, with Founders Fund, GIC, Nvidia, QIA, Radical Ventures, and TPG also participating. Crusoe will use the capital for large-scale data centers and smaller modular 'AI factories.' It also added three board members, including Cloudflare CFO Thomas Seifert. The post does not disclose specs or timelines for the AI factories.

TechCrunch · AI

Google DeepMind launches an institute to open up the AGI debate

Google and DeepMind researchers launched the DeepMind Institute on Wednesday, with Shane Legg as managing editor and James Manyika and Demis Hassabis as directors. The institute aims to surface disagreements on AGI among Google, DeepMind, and the global research community, openly stating that views will shift as frontier data emerges. Its first four essays cover economic policy for AGI disruption, preserving human-readable model reasoning, principles for human flourishing, and a framework for evaluating frontier AI.

Why it matters: DeepMind enters the AGI debate as an institution with core leadership in editorial roles — the topic carries weight. But the article only covers the launch and posture, with no first-edition topics or data disclosed, so the score stays at 78.