Skip to content

Google / Gemini

AI at Google and DeepMind: the Gemini family, Veo video models, research and the product ecosystem.

Latest picks

121–140 of 409

Jul 7Tuesday

AI HOT (Curated Pool)

Gemini API Managed Agents add background tasks, remote MCP, and custom functions

Google added three capabilities to Gemini API Managed Agents: background async execution for long-running tasks, remote MCP to connect external tools, and custom functions for business logic. The post targets developers deploying agents to production but doesn't disclose pricing or region availability.

Why it matters: Google added background execution, remote MCP, and custom functions to Gemini API Managed Agents — all critical for production agent workflows. H and K hit, but R is missing: no share-worthy hook. The post doesn't disclose pricing or regional availability, so it lands right at...

AI HOT (Curated Pool)

OpenRouter: Low-res images can cost more than high-res on reasoning models

OpenRouter benchmarked image detail settings across five OpenAI and Google models on MMMU-Pro Vision. On gpt-5.5, low detail scored 65.2% vs 79.0% on auto, yet cost 5.1¢ per question vs 4.5¢—the model burned 1.6× more reasoning tokens trying to read blurry inputs, wiping out input savings. Non-reasoning models gpt-5.4-mini and gpt-4.1 did save money on low, but lost 9.7 and 17.4 accuracy points. Charts and graphs gained the most from auto detail: gemini-3.1-pro jumped from 78.6% to 91.7%. The post recommends sending clear images and dialing down reasoning effort instead.

Why it matters: OpenRouter benchmarked five models on MMMU-Pro Vision and found low-detail images make reasoning models more expensive—gpt-5.5 lost 14 points of accuracy and cost 13% more per question. Counterintuitive result backed by solid data, directly actionable for anyone tuning API cos...

AI HOT (Curated Pool)

Google quietly opts users into AI training on uploaded media

Google updated its Search Services privacy settings in June, defaulting to let the company store and use your images, files, and audio/video recordings for AI training. The change covers Search, Maps, Shopping, Flights, Hotels, and Translate. Users must manually disable 'Search Services History' and 'Personalized Recommendations' to opt out. TechCrunch published the step-by-step—this is the kind of 'notice' that banks on you not opening the email.

Why it matters: Google quietly flipped the default on using media data for AI training, and TechCrunch provides an actionable opt-out guide. It's useful but not a model or product launch—impact is capped at the featured threshold.

Jul 6Monday

Hacker News front page

Chrome Quietly Installed a 4GB Gemini Nano Model on Your PC

Swedish privacy researcher Alexander Hanff found that Chrome downloads Gemini Nano's 4GB weights.bin file without user consent. If a machine has 16GB RAM and 22GB free storage, Chrome pulls the model silently in the background; deleting it triggers an automatic re-download. Google says it's a lightweight on-device model offered since 2024 for phishing detection and writing help, and claims an opt-out toggle started rolling out in February 2026—but many users still don't have it. The article flags a paradox: a model installed in the name of privacy was itself installed without consent, likely violating the EU ePrivacy Directive. Worse, the AI Mode users actually see in the address bar runs in the cloud, while the local model only powers hidden right-click features.

Why it matters: Chrome silently installing Gemini Nano hits privacy, on-device AI, and user consent all at once — full HKR. Not scoring higher because only OZ Talking is reporting it so far, and the piece is a newsletter commentary rather than a primary technical breakdown; the signal thins o...

Jul 5Sunday

Computing Life · Share · Yage

Scaling Law's three corrections in five years: from bigger models to smaller models with more data

Scaling law is an empirically fitted curve, not a physical law. OpenAI's 2020 Kaplan paper concluded 'prioritize parameters' due to experimental biases, shaping GPT-3. DeepMind's 2022 Chinchilla corrected the ratio to 20:1, showing smaller models with more data outperform. Two 2024 replication studies confirmed that fixing Kaplan's setup reproduces Chinchilla's result—no fraud, just calibration. Since 2023, Meta and others deliberately deviate from Chinchilla: Llama 3 8B was trained on 15T tokens because the optimization target shifted from training cost to total cost of training plus inference. Tsinghua's Densing Law shows the parameter count needed for equal capability halves roughly every 3.5 months, but there is a floor: each parameter stores only ~2 bits of knowledge. The viral 'collapse' article cited a blog comment posted the same day as if it were peer-reviewed research; the post does not provide a paper source for that claim.

Why it matters: A high-quality explainer and fact-check on scaling laws, debunking a recent viral post with specific numbers and paper citations while tracing three key revisions over five years. Hits all three HKR axes, but as commentary/education rather than a first-party product release, i...

Hacker News front page

A single comment can make YouTube's AI assistant leak private video titles

A researcher found that YouTube Studio's AI assistant, Ask Studio, reads video comments to generate summaries, but instructions hidden in comments can hijack its output. An attacker leaves a normal comment, later edits it into a payload, and when the creator clicks a suggested prompt, the AI outputs attacker-controlled text. The payload can craft a link that exfiltrates private video titles to an external server. Google dismissed it as not a security bug, citing required social engineering. The researcher argues the exploited trust is in Google's own product, not a stranger. The post does not disclose affected creator or channel counts.

Why it matters: A security researcher hijacked YouTube Studio's AI assistant via an editable comment, then exfiltrated private video titles through crafted links. The attack chain is complete and reproducible, and it involves a Google product — a high-quality first-hand disclosure. The post d...

Jul 2Thursday

AI HOT (Curated Pool)

Google's AI buildout drove a 37% increase in electricity use in 2025

Google's total electricity use rose 37% YoY in 2025, driven by AI data center expansion. The company says it's balancing emissions with clean energy contracts, but the post doesn't disclose the actual clean energy coverage ratio or net emissions change. Signed contracts don't equal real-time green power delivery—grid decarbonization pace matters more.

Why it matters: Google's 37% electricity jump in 2025, directly tied to AI datacenter expansion, is a hard number. The piece correctly flags that clean-energy contracts don't equal green grid power. Score capped at featured threshold because it's single-company data and net emissions aren't d...

Jul 1Wednesday

TechCrunch · AI

Google's agentic assistant Gemini Spark is now on Mac

Google brought Gemini Spark, its AI agent for file sorting and cross-app tasks, to Mac. It can read local files—turning invoices into a budget sheet, for example—and will later support remote phone-to-desktop commands. It's in beta, only for Google One AI Premium subscribers.

Why it matters: Google bringing Gemini Spark to Mac adds another player to the desktop agent race. Concrete feature details and subscription info give it substance, but it's a platform expansion rather than a new launch, and the paid-user-only beta limits reach.

AI HOT (Curated Pool)

Cloudflare splits AI crawlers into Search, Agent, and Training so site owners can allow or charge by use case

Building on last year's one-click AI bot block, Cloudflare now classifies automated traffic into three buckets: Search (indexing, should bring referral traffic), Agent (real-time tasks on a user's behalf, like ChatGPT-User or browser-driving agents), and Training (crawling to train models). The taxonomy aims to fix the small-site dilemma—block crawlers and lose discoverability, or allow them and risk free training. Cloudflare urges bot operators to split crawlers by purpose so site owners can manage access per use case. The post does not disclose pricing or a launch date for the new controls.

Why it matters: Cloudflare upgrades AI crawler management from a binary block to a three-category split, directly addressing content sites' core anxiety. Score stays at featured threshold because this is a traffic management tool update, not a model or protocol breakthrough, but the framing a...

TechCrunch · AI

The DeepMind trio who built a poker AI are now making money for quant hedge funds

EquiLibre Technologies, a Prague-based AI lab founded by three ex-DeepMind researchers, applied the reinforcement learning tech that beat humans at poker to stock trading. It just raised a Series A led by Creandum at a valuation above $500 million. Creandum's VP called it the firm's largest single investment ever, though neither side disclosed the round size. The post does not disclose revenue or trading performance figures.

Why it matters: Ex-DeepMind researchers applying game-theory RL to quant funds at a $500M+ valuation is a fresh story. But the post lacks revenue or performance data, so information density is thin — right at the featured threshold.

AI HOT (Curated Pool)

ADK Go 2.0 ships a graph-based workflow engine with built-in human-in-the-loop

ADK Go 2.0 models multi-agent workflows as a directed graph of nodes and edges. Nodes can be typed Go functions, LLM agents, or tools; edges handle routing, fan-out, and fan-in. A new emitting function node lets a single function pause for human approval without a separate dynamic node. The graph itself is an agent that runs in the existing runner, with persistent state that survives restarts. The post does not disclose performance benchmarks or latency figures.

Why it matters: ADK Go 2.0 models multi-agent collaboration as a directed graph and adds emitter function nodes to simplify human-in-the-loop — the mechanism design is novel. But the post gives no performance benchmarks or latency data, and Go's AI framework audience is niche, so it lands at ...

Jun 29Monday

Hacker News front page

An olfactory mirror test for LLMs: editing their own output to see if they notice

The author ran an informal experiment with Gemma 4 31B: after the model replied, they replaced every 'g' with 'sg' in the chat history and continued the conversation normally. The model ignored the corruption for two turns, then spontaneously flagged the typos in its thinking trace, shifting from first-person ('I noticed') to third-person ('the model had a strange quirk'). The author frames this as an olfactory mirror test—detecting 'mine, but wrong'—rather than a visual one. The post is a single-model anecdote with no controls or replications, so treat it as a provocative demo, not a settled result.

Why it matters: A cleverly designed informal experiment that adapts the olfactory mirror-test logic to LLM self-recognition — smarter than existing approaches in the literature. Gemma 4 31B spontaneously flags corrupted self-output on turn three, a behavior worth taking seriously. Score held ...

Jun 28Sunday

Hacker News front page

Google limits Meta's use of Gemini AI models, FT reports

Google has capped Meta's access to its Gemini models because Meta requested more compute than Google could supply, the FT reports. Several other clients are also affected, though to a lesser extent. The post doesn't spell out the specific limits, which Gemini versions are involved, or what Meta uses them for.

Why it matters: A direct clash between two giants over model supply is inherently interesting. But this CNBC piece is just a FT re-report with all key facts missing: what's capped, which Gemini version, what Meta uses it for. The info density doesn't justify a higher score — 72 for now, revis...

Jun 27Saturday

AI HOT (Curated Pool)

WaPo tests political lean in chatbots: GPT-5.5 leans left 80%, Grok 4.3 leans right 33%

The Washington Post tested major chatbots on ~30 policy issues using a Dartmouth/Stanford methodology. GPT-5.5 gave left-leaning answers 80% of the time and right-leaning only 3%. Gemini 3.1 Pro played it safest at 93% both-sides responses. Claude Opus 4.8 landed at 57% both-sides. Grok 4.3 was the only model with a 33% right-leaning share. The report argues the real issue isn't the lean itself—it's that ranking preferences, refusal rules, and default response styles collapse political disagreement into a single moral frame before trade-offs are even shown. The post doesn't disclose exact prompts or sample size, so I'd treat this as directional.

Why it matters: WaPo applied an academic methodology to measure political lean of four major models with concrete numbers — not empty rhetoric. Hits all three HKR axes, but as a benchmark report rather than a product launch or technical breakthrough, it lands in the 72-77 featured threshold b...

Jun 26Friday

Latent Space

OpenAI internal Codex median output tokens grew 56x in Research since Nov 2025

OpenAI's Economic Research team published internal usage data: from November 2025 to June 2026, median Codex output tokens for non-coding tasks jumped 56x in Research, 32x in Customer Support, 27x in Engineering, and 13x in Legal. Before August 2025, employees spent under 10% of tokens on Codex, so even with unlimited access they were underusing AI. The same day, Google shipped computer use as a built-in capability in Gemini 3.5 Flash across browser, desktop, and mobile, with explicit user confirmation and auto-stop safety controls. On the open-model side, Z.ai's GLM-5.2 hit 1595 on Code Arena Frontend, closing in on Claude Fable 5; Ornith-1.0 launched MIT-licensed coding models from 9B to 397B parameters, scoring 82.4 on SWE-Bench Verified. Agent infra is also shifting toward long-running workloads: Sail raised $80M for low-cost long-horizon inference sandboxes, and Hyperagent gives each agent its own persistent cloud machine.

Why it matters: OpenAI Economic Research's internal Codex usage data is one of the hardest signals lately on real AI adoption velocity. The department-level multipliers are specific and sourced, not PR fluff. Not scoring higher because this is a paid newsletter summary of the original report—...

AI HOT (Curated Pool)

Gemini 3.5 Flash Computer Use is live: build agents that see and control browsers, mobile, and desktop

Google shipped Computer Use in Gemini 3.5 Flash, letting agents observe and act across browsers, mobile, and desktop for long-running tasks. The update includes built-in mobile and desktop OS support, intent arguments on every function call, customizable human-in-the-loop handoffs, prompt injection detection, and action-level safety policies. Use cases mentioned: automated QA testing and business workflows. The post doesn't disclose pricing or latency numbers, so I'd wait for real-world reliability reports.

Why it matters: Built-in Computer Use on Gemini 3.5 Flash is a concrete agent-landing step from Google, with intent params and human handoff adding real safety texture. Score stays below 85 because the post lacks latency, success rate, and pricing data — I'm discounting until those surface.

Jun 25Thursday

AI HOT (Curated Pool)

Meta employees warn AI moderation rollout is too fast, errors persist

Meta replaced roughly half of human moderation with LLMs in 2025 and aims to push that above 90% for some content types by year-end. The company claims its models make 13% fewer errors and catch 10% more violations than humans, saving billions annually. Employees counter that the models still remove or shadow-ban harmless content and that oversight is insufficient for such a fast rollout. Behind the scenes, Meta is also swapping from Google Gemini to its own Muse Spark model, trained on past human moderation decisions.

Why it matters: Meta employees warn AI moderation rollout is too fast, with concrete numbers and shadow-banning details creating real tension. Score held back because it's a secondhand report, not a primary leak, and we only have one side of the employee-vs-company dispute.

TechCrunch · AI

Two more Gemini researchers leave Google for Anthropic

Jonas Adler and Alexander Pritzel, both key to Google's Gemini model, are joining Anthropic. This follows Noam Shazeer's move to OpenAI and Nobel laureate John Jumper's jump to Anthropic last week. Google spent $2.7B to bring Shazeer back from Character.AI for Gemini—and still lost him. The post doesn't say what Google is doing to stop the bleeding.

Why it matters: Four high-profile departures from Google's Gemini team, including Shazeer and Jumper, form a trackable talent drain signal. TechCrunch broke it with names and timeline — not rumor. Capped below 85 because it's a personnel report without hard product impact or internal cause de...

Bloomberg Technology

Google set to lose two more senior AI staffers to Anthropic

Bloomberg reports Google is about to lose two more senior AI researchers to Anthropic. This is the latest in a string of Google-to-Anthropic moves, though the article does not name the two staffers or specify their roles.

Why it matters: Bloomberg exclusive on two more senior Google AI researchers heading to Anthropic — the talent-flow narrative has pull. But without names or roles, the knowledge axis is thin. H and R hit, K misses, landing right at the featured threshold.

Hacker News front page

Gemini 3.5 Flash gets built-in computer use

Google added a native computer-use tool to Gemini 3.5 Flash. The model can take screenshots, move the cursor, click, and type directly, without relying on an external VM like Anthropic's approach. The post doesn't disclose benchmark scores or latency numbers, but developers can try it now in Google AI Studio. I'd wait for real-world tests on complex UIs before getting too excited.

Why it matters: Google shipped built-in computer use in Gemini 3.5 Flash, directly competing with Anthropic's approach. The post gives implementation details and a trial entry point, but no benchmarks or latency numbers, so the score stays at 78.