Skip to content

#编码

10 today

May 20Wednesday

Latent Space

Google I/O 2026: Gemini 3.5 Flash, Omni, Spark, and Antigravity 2.0

Google announced Gemini 3.5 Flash at I/O 2026 with a 1M-token context window, 65k max output, four thinking levels, and pricing of $1.50 per 1M input tokens and $9.00 per 1M output tokens.

Why it matters: HKR-H/K/R all pass: this is a Google I/O model-and-product bundle with concrete context, output, thinking-tier, and pricing facts. It has same-day relevance for Claude, OpenAI, and coding-agent competition, so it clears P1.

AI HOT (Curated Pool)

Microsoft reportedly warns internally that GitHub faces existential risk as AI coding tools reduce hosting need

Microsoft internally warned that GitHub faces an existential risk from AI coding assistants such as Cursor and Claude Code, and told some teams to stop using Claude Code by the end of June 2026 and move to GitHub Copilot CLI.

Why it matters: HKR-H/K/R all pass: the angle is sharp, the summary gives a Claude Code-to-Copilot CLI deadline, and the workflow stakes are real. Single-source “reported” framing and no Microsoft response keep it below the 85 must-write band.

Bloomberg Technology

SpaceX Is Planning to Buy Startup Cursor 30 Days After IPO

SpaceX plans to acquire AI coding startup Cursor 30 days after Elon Musk’s company begins public trading; the post does not disclose the deal price, IPO timeline, or regulatory conditions.

Why it matters: Bloomberg sourcing plus the odd SpaceX-after-IPO Cursor deal clears HKR-H/K/R. Price, IPO timing, and regulatory conditions are not disclosed, so it stays below the 85 P1 line.

AI HOT (Curated Pool)

Claude Code’s HTML Output: The Unreasonable Effectiveness of HTML

The Claude Code team is shifting its primary output format from Markdown to HTML, and the post names four mechanisms: tables, CSS styling, SVG charts, and JavaScript interactions.

Why it matters: Official Claude Code post with a concrete shift from Markdown to HTML and 4 output mechanisms; strong practitioner utility, but not a major product launch, so it sits in the featured threshold band.

AI HOT (Curated Pool)

Empirical Research Assistant ERA: From Nature Publication to Computational Discovery

Google Research published its Gemini-based Empirical Research Assistant in Nature and opened early access through the Google Labs trusted tester program.

Why it matters: HKR-H/K/R all pass: Google moves Gemini-based ERA from a Nature paper to a Labs trusted-tester trial. Score stays at 78 because the provided text lacks metrics, benchmark setup, or reproducible workflow details.

Hacker News front page

Gemini CLI will stop working from June 18, 2026

Google Developers Blog says Gemini CLI will stop working on June 18, 2026, and the post title points to a transition to Antigravity CLI; the RSS snippet only includes the article URL, Hacker News comments URL, 36 points, and 10 comments, and it does not disclose migration steps, compatibility details, or replacement behavior.

Why it matters: Google’s developer blog gives a Gemini CLI shutdown date and migration target, clearing HKR-H/K/R. Detail is thin beyond the deadline, so it stays in the low featured band.

AI HOT (Curated Pool)

Gemini 3.5 Series Launches With Stronger Agent and Coding Performance

Google AI launched the Gemini 3.5 series with Gemini 3.5 Flash first, stating it targets agent and coding performance; the post does not disclose parameters, pricing, benchmark scores, or context window size.

Why it matters: HKR-H/K/R all pass: Google’s Gemini 3.5 series is a first-tier model update. Sparse details on price, context, and benchmarks keep it at the low end of the must-write band.

AI HOT (Curated Pool)

Gemini 3.5 Flash launches with stronger performance and speed

Google opened Gemini 3.5 Flash after Google I/O across its products and API; the post says it outperforms Gemini 3.1 Pro on most benchmarks and generates tokens 4x faster than other frontier models.

Why it matters: HKR-H/K/R all pass: Sundar Pichai announced Gemini 3.5 Flash with product/API access and a 4x token-speed claim. This is same-day model-release signal, though price, context window, and full evals are not disclosed.

TechCrunch · AI

With Gemini 3.5 Flash, Google bets its next AI wave on agents, not chatbots

Google launched Gemini 3.5 Flash at its annual developer conference, describing it as its strongest coding and agentic AI model yet; the RSS snippet says it can autonomously execute complex tasks and build software from scratch, but the post does not disclose parameters, pricing, or context window.

Why it matters: HKR-H/K/R all pass: a Google model release with an agent-first and code-building claim is same-day material. Missing parameters, pricing, and context window keep it at the low end of the must-write band.

The Verge · AI

Google wants to compete with Anthropic’s Mythos

Google invited select experts at I/O to test the CodeMender API, an AI agent for code security that flags and fixes vulnerabilities; the RSS snippet does not disclose launch timing, pricing, benchmark results, or concrete details about Anthropic’s Claude Mythos Preview.

Why it matters: HKR-H/K/R all pass, but the post only confirms closed expert testing and the flag/fix mechanism; availability, pricing, and eval results are not disclosed, so this stays at the featured threshold.

AI HOT (Curated Pool)

Google releases Gemini 3.5 Flash for complex agent workflows

Google introduced Gemini 3.5 Flash at Google I/O for long-running agent workflows; it outscored 3.1 Pro on Terminal-Bench and MCP Atlas, runs up to 4x faster than other frontier models, and reaches up to 12x speed gains in Google Antigravity.

Why it matters: HKR-H/K/R all pass: Google launched Gemini 3.5 Flash for long-horizon agents with benchmark and speed claims. This is a same-day major model update, below industry-shaking tier.

The Verge · AI

Google can now vibe-code you an Android app

Google upgraded AI Studio today to build native Android apps from prompts, with an embedded Android emulator for preview and device connection for installation; the initial release focuses on personal utility apps, and tester invites are planned but not yet available.

Why it matters: HKR-H/K/R all pass: Google AI Studio now generates installable native Android prototypes with emulator preview and device install. Scope stays limited to personal utility apps, so this is a mid-weight product update, not P1.

TechCrunch · AI

Agentic app coding gets an upgrade with Google’s release of Android CLI

Google released Android CLI for AI coding agents, letting platforms such as Claude Code and OpenAI Codex build Android apps from the command line; the RSS snippet does not disclose version numbers, release timelines, pricing, or performance data.

Why it matters: HKR-H/K/R all pass, but the body lacks version, timeline, and performance data. Google plus Android plus agentic coding clears the featured line, not the must-write band.

TechCrunch · AI

Google’s AI Studio now lets anyone build Android apps in minutes

Google unveiled web-based AI tools in AI Studio that generate native Android apps in minutes. The RSS snippet does not disclose the model, pricing, availability, or supported development constraints.

Why it matters: Google AI Studio’s coding update clears HKR-H/K/R with a strong speed hook and a concrete capability. Model, pricing, and rollout are not disclosed, so it stays at the featured threshold for a mid-weight product update.

TechCrunch · AI

Google launches Antigravity 2.0 with updated desktop app and CLI tool at I/O 2026

Google launched Antigravity 2.0 with an updated desktop app and CLI tool, and introduced a $100 AI Ultra plan that gives users 5x the usage limit of AI Pro; the post does not disclose the desktop app or CLI feature details.

Why it matters: HKR-H/K/R pass, but the post does not disclose concrete desktop or CLI capabilities, so it stays below 78. Google I/O plus the $100 plan and 5x quota clear the featured bar.

AI HOT (Curated Pool)

Google AI Ultra plan gets a price cut and a new tier

Google cut the top AI Ultra plan from $250 to $200 per month and added a $100 monthly tier with 5x the Gemini app usage limit of Pro, 20TB of storage, early access to new features, and YouTube Premium under stated terms.

Why it matters: HKR-H/K/R all pass, but this is subscription pricing and quota packaging, not a model or capability launch. Official source and concrete prices put it at the featured threshold.

r/LocalLLaMA

Public repository Codegraph claims 94% fewer Claude, Cursor, Codex, and OpenCode tool calls locally

Codegraph uses a pre-indexed knowledge graph for symbol relationships, call graphs, and code structure. In the VS Code test, it reduced tool calls from 52 to 3 and runtime from 1m37s to 17s.

Why it matters: All HKR axes pass, but evidence is a Reddit/public-repo self-test without independent replication. The 94% reduction and 52→3 call count clear featured, not p1.

AI HOT (Curated Pool)

Gemini 3.5 launches with frontier intelligence and real-world action

Google DeepMind introduced Gemini 3.5 with 3.5 Flash as the first release; the post says it is the company’s strongest model for agents and coding, but does not disclose parameters, pricing, or context window.

Why it matters: HKR-H/K/R all pass because Google DeepMind shipped a major Gemini 3.5 update with 3.5 Flash first and agent/coding claims. Missing price, parameters, and context window keep it mid-85–94, not higher.

r/LocalLLaMA

Cursor and Claude Code Are Not Getting Dumber; Agent Loops Are Suffocating Context

A Reddit user says an API-log audit showed Cursor and Claude Code recursively grep about 40 files in 10k-plus-line repositories, sometimes load 2k-line files for 5-line edits, and spend roughly 30k tokens on tool definitions and logs before generating code.

Why it matters: HKR-H/K/R all pass: the hook is contrarian, the API-log numbers are concrete, and coding-agent context waste is a live practitioner pain. Reddit single-post sourcing and no shared logs keep it at the featured threshold.

May 19Tuesday

AI HOT (Curated Pool)

I really want to praise HTML!

The author used Claude Code to generate a single-file HTML project plan page in 2 minutes, with a dark theme, timeline, and collapsible tables; the comparable Notion template previously took 30-40 minutes.

Why it matters: HKR-H/K/R all pass: the post has a concrete Claude Code workflow hook, a 2-minute vs 30-40-minute comparison, and clear practitioner resonance. Scope is small, so it sits at the featured threshold.

AI HOT (Curated Pool)

Kimi's Latest Funding Adds State Capital and Central SOEs, Valuation Quadruples in Six Months

Moonshot AI’s Kimi is raising $2 billion, with Guozhitou and China Mobile added to the shareholder list; in January and February, Kimi completed three funding rounds totaling more than $3.9 billion.

Why it matters: HKR-H/K/R all pass: Kimi is a top Chinese model player, with a reported $2B raise, 4x valuation jump, and Guozhitou/China Mobile entering. Because the round is still in progress, it stays below a completed major launch or IPO.

Latent Space

[AINews] How to Land a Job at a Frontier Lab (on Pretraining)

Latent Space says Vlad Feinberg’s pretraining job-prep notes reduce frontier-lab readiness to kernel-level performance work: derive Chinchilla laws, compare dense and MoE architectures, code the solution in JAX, then write a Pallas kernel that beats jax.lax.ragged_dot for F > D by fusing up/down projections.

Why it matters: HKR-H/K/R all pass: the career hook is strong and the prep list is concrete. It is not a model release or major product update, and the kernel-heavy angle keeps it at the lower featured band.

Xinzhiyuan · WeChat

AI startup annualized revenue hits $80B, with OpenAI and Anthropic taking 89%

The Information says 34 leading AI startups reached about $80 billion in annualized revenue, with OpenAI and Anthropic taking 89%, while Anthropic exceeded $30 billion in April 2026 and surpassed OpenAI’s reported $25 billion.

Why it matters: HKR-H/K/R all pass: the story has a sharp Anthropic-vs-OpenAI hook, concrete revenue-concentration numbers, and startup-economics resonance. It is secondary financial reporting, not a model release, so it stays in 78–84.

Synced · WeChat

Recent LLM Architecture Changes: From Gemma 4 to DeepSeek V4

Jiqizhixin translated Sebastian Raschka’s blog on recent LLM architecture changes, covering long-context cost reductions in Gemma 4, Laguna XS.2, and ZAYA1-8B; the article states that Gemma 4 E2B saves about 2.7GB of KV cache at 128K context with bfloat16 precision.

Why it matters: HKR-H/K/R pass: notable model names, a concrete 128K bf16 KV-cache saving, and inference-cost relevance. As a translated survey rather than a release, it stays in the 72–77 featured band.

AI HOT (Curated Pool)

Cursor releases Composer 2.5, calling it its strongest model yet

Cursor released Composer 2.5, claiming a 10x efficiency gain at comparable capability, with larger training scale, more complex reinforcement-learning environments, and a text-feedback mechanism.

Why it matters: Cursor Composer 2.5 is a substantive model update for a front-line AI coding tool, with HKR-H/K/R from the 10x efficiency and RL-training details. The single social-source summary lacks benchmarks, pricing, and reproducible tests, keeping it in the 78–84 band.

TechCrunch · AI

Anthropic has acquired the dev tools startup used by OpenAI, Google, and Cloudflare

Anthropic acquired Stainless, a New York startup founded in 2022 that automates creation and maintenance of SDKs for developers using APIs; the post does not disclose the deal price or Anthropic’s integration plan.

Why it matters: HKR-H/K/R pass: the rival-used startup hook is strong, the SDK automation mechanism is concrete, and the Anthropic developer-stack angle resonates. Missing deal value and integration details keep it below the 78+ band.

AI HOT (Curated Pool)

Claude Code Fast Mode Defaults to Opus 4.7

Claude Code fast mode now defaults to Opus 4.7, and the post discloses the /fast invocation method but does not disclose pricing, context window, rate limits, or rollout conditions.

Why it matters: HKR-H/K/R pass because a Claude Code default changes to Opus 4.7 with a concrete /fast path. Sparse details on price, context window, and limits keep it in the 72–77 band.

r/LocalLLaMA

llama.cpp MTP support landed: Qwen3.6 27B reaches 2.44× on Strix Halo

llama.cpp merged MTP speculative decoding in PR #22673; Qwen3.6 27B Q8_0 rose from 7.4 to 18.1 tok/s on Strix Halo, while a dual RTX 3090 Q8_0 setup rose from 25.7 to 55.9 tok/s.

Why it matters: HKR-H/K/R all pass: llama.cpp adds MTP speculative decoding with Qwen3.6 27B speedups on Strix Halo and RTX 3090. The scope is local inference, not a broad model release, so 78 fits featured.

AI HOT (Curated Pool)

Cursor releases Composer 2.5 coding model

Cursor released Composer 2.5, claiming up to 10x higher efficiency on long coding tasks; the model is further trained on Moonshot’s Kimi K2.5 and uses text feedback for 100k-token-scale trajectories.

Why it matters: HKR-H/K/R all pass: Cursor is a core AI coding tool, and Composer 2.5 adds concrete claims around 10x long-task gains and Kimi K2.5 tuning. Limited sourcing and no independent eval keep it in the 78–84 band.

AI HOT (Curated Pool)

Take your local GitHub sessions anywhere

GitHub launched remote control sessions for Copilot, letting users start tasks in VS Code or the command line and continue them through github.com or GitHub Mobile.

Why it matters: GitHub Copilot session handoff from VS Code/CLI to web and mobile clears HKR-H/K/R, but the post only gives entry points and use case; permissions, pricing, and supported task scope are not disclosed.

May 18Monday

Hacker News front page

Show HN: InsForge – Open-source Heroku for coding agents

InsForge released an Apache 2.0 backend platform that lets coding agents deploy, operate, and debug backend systems through one CLI install command and Skills.

Why it matters: HKR-H/K/R all pass: the Heroku-for-agents framing, Apache 2.0 plus one-CLI install, and agent ops pain are concrete. Source is mainly Show HN/GitHub with no usage, benchmark, or production proof, so it sits at the featured threshold.

r/LocalLLaMA

Qwen 3.6 27B on 24GB VRAM: backend comparisons, quant choice, and settings

The author tested Qwen 3.6 27B on an RTX 3090 24GB and kept ik_llama.cpp with Qwen3.6-27B-MTP-IQ4_KS.gguf; at 156k context with q8_0 KV and MTP, a ~5.9k-token prompt plus 1024-token output reached about 1261 tok/s prefill and 72.9 tok/s decode, while vLLM lacked a clean single-card long-context run.

Why it matters: HKR-H/K/R all pass: this is a first-person local inference benchmark with concrete VRAM, quant, context, and speed numbers. Its reach is narrower than a model release, so it sits at the featured threshold.

AI HOT (Curated Pool)

OpenAI and Dell partner to bring Codex to hybrid and on-prem enterprise environments

OpenAI and Dell are partnering to bring the Codex coding agent to enterprise hybrid-cloud and on-premises deployments; the RSS snippet does not disclose launch timing, pricing, supported regions, or the specific security controls for sensitive data.

Why it matters: HKR-H/R are strong: Codex via Dell targets hybrid/on-prem enterprise code. HKR-K is limited to deployment path; launch date, price, regions, and security controls are absent, keeping it at the featured threshold.

r/LocalLLaMA

I built a coding agent that gets 87% on benchmarks with a 4B parameter model

SmallCode passes 87 of 100 benchmark tasks with Gemma 4 activating 4B parameters per token. The author attributes the result to compound tools, compile and lint feedback, task decomposition after two repeated failures, and optional escalation to Claude or OpenAI for one task.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the benchmark identity plus replication details are incomplete. It fits a concrete first-person experiment above the featured bar, not the 78+ band.

AI HOT (Curated Pool)

Project Glasswing: What Mythos Shows Us

The team applied Mythos and other security-focused LLMs to real-time code testing for critical infrastructure; the post reports vulnerability detection strengths, false positives, and unstable context handling, but does not disclose sample size or benchmark metrics.

Why it matters: Cloudflare offers applied observations on security LLMs testing critical-infrastructure code, so HKR-K/R pass. Missing sample size and metrics keep it low in the 72-77 band; no hard-exclusion rule applies.

AI HOT (Curated Pool)

Tencent AI Design Agent Ardot Enters Public Beta: Generates Editable Designs and Converts Them to Code

Tencent Cloud opened public beta for Ardot, an AI design agent that generates editable app pages, websites, and posters from one-sentence prompts, then converts designs to code.

Why it matters: HKR-H/K/R pass on a concrete Tencent product beta for editable design-to-code workflows. Missing pricing, model details, benchmarks, and field results keep it at the lower featured threshold.

AI HOT (Curated Pool)

Composer 2.5 release and technical analysis

Cursor released Composer 2.5, built on a Moonshot open-source checkpoint, trained with synthetic data from real codebases at 25 times the previous scale, and updated with text-feedback reinforcement learning and a sharded Muon optimizer.

Why it matters: HKR-H/K/R all pass: Cursor is a core coding-agent surface, and the post gives concrete training details around Moonshot, 25x data, RL, and Muon. It lacks benchmarks, pricing, or user-facing capability limits, so it stays in the 78–84 band.

May 17Sunday

Hacker News front page

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

MinishLab open-sourced Semble, a code-search tool for agents that combines Model2Vec embeddings, BM25, RRF fusion, and reranking; on a 63-repo benchmark, it used 98% fewer tokens than grep+read, reached 0.854 NDCG@10, and ran CPU queries in about 1.5 ms.

Why it matters: HKR-H/K/R all pass: the 98% token claim is clickworthy, the 63-repo benchmark adds substance, and coding-agent context cost is a real practitioner nerve. Impact is still toolchain-level, so it stays below must-write.

r/LocalLLaMA

DeepSeek V4's 1M Context Window: The Breaking Point

A Reddit user tested DeepSeek V4 on 45k, 180k, and 520k-token codebases and found 150k-250k tokens best for coding work. Past 300k tokens, line-number precision degraded; at 520k, outputs shifted toward architecture summaries and skipped implementation details.

Why it matters: A single Reddit post limits authority, but HKR-H/K/R all pass: it is a numbered first-person test with a concrete long-context failure pattern. The right band is featured, not 78+, because replication and model details are thin.

Synced · WeChat

AI agents may spend 1,000x more tokens without better results: the hidden bill

Researchers used OpenHands to analyze traces from 8 frontier models on 500 swe-bench-verified tasks, finding that agentic coding reached a 154:1 input-output token ratio and that human difficulty labels correlated weakly with token use at Kendall tau 0.32.

Why it matters: All HKR axes pass: strong cost-performance hook, concrete benchmark setup and correlation numbers, and direct resonance with coding-agent economics. It is not a model or platform launch, so it fits the 78–84 quality-recommendation band.