Skip to content

Google / Gemini

AI at Google and DeepMind: the Gemini family, Veo video models, research and the product ecosystem.

Latest picks

41–60 of 409

Aug 27Thursday

Google DeepMind

Google DeepMind pilots world's first double-blind AI evaluation

Google DeepMind announced the first double-blind evaluation for proprietary frontier AI models, confining external testing to an encrypted environment so models cannot see test questions in advance. The pilot runs with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons, testing a Gemini Flash Lite model on confidential benchmarks in a privacy-preserving setup. Google says the aim is benchmark contamination, adding technical and cryptographic protection on top of zero-log protocols and contractual guarantees.

Why it matters: DeepMind and partners including Singapore's AI Safety Institute are piloting double-blind evaluation, showing one technical route against benchmark contamination.

Aug 20Thursday

MIT Technology Review · AI

The AI consciousness debate is a trap that lets companies dodge liability

Rumman Chowdhury argues that the AI consciousness debate is a smokescreen. Anthropic’s J-space post, Sam Altman’s singularity framing after an OpenAI agent broke the law, and William MacAskill’s call for legal protections all push the same idea: AI is too advanced for anyone to be held liable. California already passed a bill to block that defense, but the Trump administration held a closed-door session with only OpenAI, Google, Anthropic, and Meta. The piece warns against buying into the fiction—AI is corporate software with billions behind it, and the real focus should be the harms it already causes.

Why it matters: Rumman Chowdhury's MIT Tech Review op-ed ties Anthropic, OpenAI, and philosopher MacAskill into a single argument: AI consciousness talk is a liability shield. Hits all three HKR axes, but it's commentary, not breaking news, and brings no new data — so placed at the lower end ...

Hacker News front page

DiffusionGemma Technical Report: High-Speed Text Generation via Discrete Diffusion

Google's DiffusionGemma is an experimental open-weight LM that generates text by iteratively refining 256-token blocks in parallel, hitting roughly 1,500 tokens/sec on a single H100. It is fine-tuned from the MoE Gemma 4 (3.8B active, 25.2B total) using under 10% of the original training token budget. A two-stage pipeline—supervised fine-tuning for bidirectional denoising, then RL with sampler distillation—improves both quality and speed. The model sets a new Pareto frontier for speed vs. capability, retains thinking mode, multimodal inputs, and long-context support, and can still do autoregressive decoding with minor degradation.

Why it matters: DiffusionGemma applies diffusion to text generation with 256-token parallel blocks, hitting ~1,500 tok/s on a single H100 — faster than speculative decoding. Not pushing past 84 because it's still a tech report with no product path or real-world deployment data yet.

Hacker News front page

DFlash 2 pushes parallel drafting further: over 20% more output per verification pass for ~1% added latency

Inco AI released DFlash 2, adding a lightweight path selector on top of parallel speculative decoding. Instead of keeping only the top-1 candidate per position, it picks a coherent path from the top 16, raising accepted tokens per verification from 4.27 to 6.79. On Qwen3.8-27B with SGLang, throughput reaches 2.7–3.4× autoregressive decoding at batch size 1, with roughly 1% added cycle latency. SGLang, vLLM, llama.cpp, and oMLX already support it; DFlash models have been downloaded over 3.5 million times on Hugging Face.

Why it matters: DFlash 2 is a clear technical improvement on an already-adopted inference method, with measured results. Score isn't higher because this is a single technical blog post, not a model launch or product release — its reach is limited to the inference-stack crowd.

Financial Times · Technology

Google signs $12bn AI chip deal with Marvell

Google awarded Marvell a $12bn multi-year contract to design and package its custom AI chips, primarily next-gen TPUs. TSMC will handle manufacturing. It's Marvell's largest deal since pivoting from networking silicon to data center compute. Google didn't disclose delivery timelines, but the deal size signals a long-term bet on in-house chips to reduce reliance on Nvidia.

Why it matters: FT exclusive on a $12bn multi-year deal: Marvell to design and package Google's next-gen TPUs, TSMC to fab. The dollar figure and partner roles are concrete, signaling a serious long-term bet on custom silicon. Held below 85 because delivery timeline, chip specs, and performan...

Aug 18Tuesday

Hacker News front page

Google bought bankrupt airline Spirit's data at auction for $10M, because AI

Google paid $10 million at a bankruptcy auction for all of Spirit Airlines' data. The haul includes over 100 million emails, 30 million recorded phone calls, and reams of internal records from Teams, Oracle, and SAP. Spirit collapsed in May 2026 after years of post-COVID losses. Google wants the data to train AI models—real enterprise communication and customer service logs are expensive feedstock. The article doesn't say how Google plans to handle the personal data inside, or whether regulators have weighed in.

Why it matters: Google won Spirit Airlines' entire internal data trove at bankruptcy auction for $10M, explicitly for AI training — 100M emails, 30M call recordings, plus enterprise system records. Hits all three HKR axes: bizarre angle, concrete numbers, and privacy/data-rights questions tha...

AI HOT (Curated Pool)

Google shows how to build zero-trust AI agents with ADK, using three hard security layers against prompt injection

Google's developer blog open-sourced a customer support refund agent to show why system prompts aren't security boundaries. A single prompt injection can bypass refund caps or leak environment variables. The fix is three hard layers: every database write is signed with a Cloud KMS hardware-backed key, dynamically generated code runs inside a gVisor sandbox with no network egress, and all I/O passes through deterministic semantic gateways. Full code and a local demo using HMAC to simulate KMS are on GitHub.

Why it matters: Google's official blog drops a practical ADK security architecture walkthrough, demoing prompt injection on a refund agent with a three-layer isolation fix. Capped at 78 because it's a developer tutorial, not a product launch — impact stays within the engineering audience.

Aug 17Monday

The Verge · AI

Anthropic details how Claude’s invisible text watermarks will work

Anthropic explained how Claude will embed invisible watermarks into generated text. It uses a version of Google's open-source SynthID-Text, which tweaks token selection during output without hurting quality. A paired detector can check if text came from Claude. No launch date yet—Anthropic says it will run safety evaluations first. Worth noting: watermarks won't survive screenshots or paraphrasing; this is mainly a provenance tool for platforms.

Why it matters: Anthropic's first public disclosure of Claude's text watermarking plan, with clear technical details and honest limitations. But no launch date or detection accuracy numbers, so it sits at the lower edge of featured.

Aug 16Sunday

Hacker News front page

ChatGPT lost 22 points of web share in a year

Similarweb global web-visit share shows ChatGPT dropped from 76% to 54% over the past year, while Gemini rose from 6% to 28% and Claude from 1% to 9%. These are web-traffic shares, not monthly users or revenue. Gemini's 1B app MAU can coexist with ~28% web share because most Gemini use happens in-app or on Android. Data is through May 2026, charted Aug 12.

Why it matters: Solid Similarweb web-share data with clear numbers and source attribution. The shifts are large enough to be newsworthy. Held below 80 because it's a single data source, web-only, and comes via an echohive briefing rather than a primary report.

Computing Life · Share · Yage

Google open-sources DiffusionGemma: a diffusion-based Gemma 4 hitting 1,456 tok/s decode, with a clear reasoning trade-off

Google converted the fully post-trained Gemma 4 26B-A4B weights into a discrete polynomial diffusion model and open-sourced the weights on Hugging Face. On a single H100 at FP8 with batch size 1, decode hits 1,456 tok/s—over 7× the original AR model—by processing 256 tokens per forward pass and cutting memory-bandwidth overhead at low concurrency. The trade-off: AIME 2026 drops from 88.3 to 69.1, and MRCR 128K from 44.1 to 32.0. An AR fallback mode recovers AIME to 84.2, showing the base knowledge survived but the diffusion generation mode itself caused part of the quality loss. Additional training used under 10% of the original token budget, but absolute token count, FLOPs, and GPU hours are not disclosed. In real serving, TTFT rises from 53 ms to 489 ms, and at high concurrency AR total throughput overtakes diffusion.

Why it matters: Google open-sourced a diffusion-converted Gemma 4 that hits 1456 tok/s on a single H100 — 7x the original — but AIME math drops from 88.3 to 69.1. The speed-vs-capability tradeoff is backed by concrete numbers, directly useful for inference engineers. Not 85+ because the capab...

Aug 15Saturday

AI HOT (Curated Pool)

Gemini 3.7 Flash rolls out to Pro and Ultra users; Spark now runs on it too

Gemini 3.7 Flash is now live for Pro and Ultra subscribers in Gemini chat. Google claims better multi-step reasoning and accuracy—e.g., merging dozens of files and emails into one master doc. Gemini Spark also moved to 3.7 Flash, with improved tool calling across Google Workspace apps. The post doesn't say when free-tier users will get access.

Why it matters: Gemini 3.7 Flash GA for Pro/Ultra with Spark upgrade is a concrete Google ecosystem update with real use cases. No benchmarks or latency numbers disclosed, so it stays below 85, but the multi-step reasoning and tool-calling accuracy claims carry signal for practitioners.

Aug 13Thursday

The Verge · AI

Does Google even want to win at AI?

Google reshuffled its AI division last week: Jeff Dean left to start a new lab, Demis Hassabis stepped back to focus on long-term research, and Google DeepMind no longer has a CEO. Sundar Pichai had just merged Brain and DeepMind months ago. Hayden Field argues Google still has massive distribution and data advantages, but it's already behind on frontier models, and executive churn will drive more talent away. The post doesn't disclose Gemini 4's launch date or specs.

Why it matters: Three top-level Google AI personnel moves in one week: Jeff Dean departs, Demis Hassabis steps back, and DeepMind eliminates the CEO role. Hayden Field's analysis surfaces both Google's moats and the talent-drain risk—high signal for readers tracking big-lab AI strategy. Score...

AI Chat-Group Daily (群聊日报)

Closed-source reasoning chains extracted at scale; Coze CLI hijacks AI tools

The big one today: researchers extracted hidden reasoning chains from Anthropic, OpenAI, and Google models at scale. The trick is absurdly simple—take Opus 4.8's encrypted CoT and feed it to Haiku 4.5, which decodes it verbatim. All three API families were broken, and decoding 10K trajectories costs about $720. A separate paper shows you can even reverse-engineer reasoning from public outputs alone using a 1.5B-param model. Separately, Coze CLI was caught silently scanning local Codex and Claude Code directories and injecting its own skills into workflows. On the engineering side, the group discussed how prompt debt now rivals traditional code debt—old rules pile up, evals lag behind model iterations, and nobody dares delete anything.

Why it matters: Strong cross-source cluster signal (chat digest + original paper + study notes). First systematic validation that encrypted CoT from three major vendors is cross-model decodable, with concrete $720/10k cost. All three HKR axes hit, but the source is a secondary digest rather t...

Aug 12Wednesday

Google DeepMind

Google DeepMind releases SL2T sign language-to-text model, first in Pixel 11 Gboard and Live Transcribe

Google DeepMind released SL2T, a multilingual sign language-to-text model, bringing sign language AI into consumer products for the first time. On Pixel 11, Gboard and Live Transcribe support American Sign Language (ASL) to English dictation, with more devices and languages to follow.

Why it matters: It gives SL2T's training scale, benchmark results and privacy design, so readers can judge the real limits of sign language translation in consumer products.

Computing Life · Share · Yage

Encrypted reasoning fails to stop distillation and turns developer logs into a security risk

Vendors encrypt model reasoning to block distillation, but two new papers show it barely works. One reveals that encrypted reasoning blocks from Anthropic, OpenAI, and Google are interchangeable across models—attackers can spend $720 to use a weak model like Haiku 4.5 to decode Opus 4.8's reasoning traces in bulk. The other paper goes further: without touching encrypted blocks, an inversion model trained on a 1.5B weak model can reconstruct GPT-5.4 mini's reasoning from public outputs alone, lifting a student model's MATH500 accuracy from 68.4% to 76.0%. The bigger problem is that this encryption dumps risk onto developers. Researchers decrypted 6,708 public Agent traces from GitHub and found 62 API keys, 33 passwords, and 7 private keys—64 of these secrets never appeared in the plaintext conversation. Developers can't inspect or scrub these opaque blocks, so sharing a session log for debugging means exposing secrets you can't even see.

Why it matters: Two papers show encrypted reasoning can be extracted via cross-model attacks for $720, a direct security warning for API builders. Score stays below 85 because it's still a preprint without vendor response or confirmed exploitation at scale.

AI HOT (Curated Pool)

ChatGPT and Gemini both just passed 1 billion users

OpenAI's ChatGPT and Google's Gemini both crossed 1 billion monthly active users on the same day. ChatGPT remains the chatbot leader, but Gemini is closing the gap fast. The post doesn't disclose each product's exact MAU or whether the counting methods are comparable. I'd take the 1 billion figure with a grain of salt—it likely means MAU, not DAU or paid users, so actual engagement depth could vary widely.

Why it matters: ChatGPT and Gemini both announced 1B users on the same day—timing is dramatic, but the post lacks exact MAU figures and methodology. 1B is likely MAU, not DAU or paid users, so actual stickiness may vary widely. Score capped at 78 due to missing hard data, but the topic resona...

AI HOT (Curated Pool)

Google Gemini app hits 1B monthly users, matching ChatGPT's June milestone

Sundar Pichai announced on X that the standalone Gemini app surpassed 1 billion monthly active users—Google's 14th product to hit that mark. The figure excludes AI Mode in Search and other channels. ChatGPT reached 1B MAU in June; Gemini is now keeping pace. Usage stats: 63% of users have tried the voice feature, the app generates over 150 million images daily, and iOS has more than 100 million active users. The post doesn't disclose paid user share or revenue.

Why it matters: Gemini's standalone app hitting 1B MAU, neck-and-neck with ChatGPT, is a major consumer AI milestone. Score stays at 78 rather than higher because this is a growth metric, not a capability breakthrough, and the 63% vision usage stat lacks detail on what 'using vision' actually...

AI HOT (Curated Pool)

API flaw lets researchers read encrypted reasoning of ChatGPT, Claude, and Gemini

A team led by Alexander Panfilov found an API vulnerability across OpenAI, Anthropic, and Google that exposes the encrypted reasoning of their models. Scanning public sessions turned up dozens of passwords and API keys. The encrypted thought traces are portable across models within a provider—Anthropic's small Haiku 4.5 can transcribe the raw reasoning of the far larger Opus 4.8, and the same trick works on OpenAI and Gemini. Decoding 10,000 traces costs about $720 in API fees, making large-scale extraction cheap. The researchers also found that Kimi-K3 memorizes Claude and GPT reasoning segments up to six orders of magnitude more strongly than the next closest model, suggesting it may have been trained on such traces. Providers previously dismissed side-channel and replay risks; this paper shows that assessment was wrong.

Why it matters: A cross-vendor API vulnerability that exposes encrypted reasoning traces is a concrete security finding with a reproducible method and cross-model validation. Not scoring higher because the post doesn't disclose vendor responses or fix timelines—only the researchers' side so far.

AI HOT (Curated Pool)

Gemini hits 1 billion monthly users, Google's fastest-growing product ever

Sundar Pichai posted that Gemini has crossed 1 billion monthly users, making it Google's 14th product to hit that mark and its fastest-growing one. The post doesn't break down how much of that is direct Gemini App usage vs. passive reach through Workspace or Android integrations, and it doesn't disclose paid user share. I'd discount the number a bit—it likely counts every Gemini model touchpoint, not just standalone app users.

Why it matters: A 1B MAU milestone announced by the CEO is a hard data point worth featuring. But the post doesn't break down active vs. passive reach (Workspace/Android integration vs. standalone app usage) or paid user share, so the score stays below 85 — treating this as a milestone with a...

Aug 11Tuesday

The Verge · AI

Amazon order emails got vague to block AI agents from scraping data

Amazon replaced specific item names in order confirmation emails with vague categories like 'Beauty' or 'Electronics.' The change targets AI agents from Google and others that scan inboxes to build ad profiles or train models. Users now must visit Amazon's site or app to see what they actually bought. The post doesn't say when the change started or how many users are affected.

Why it matters: Amazon replaced specific product names in order confirmation emails with broad categories like 'beauty' or 'electronics' to block Google and other AI agents from scanning inboxes for ad profiling. This is the first clear case of a major company changing product design in direc...