Skip to content

All news

25 today

May 19Tuesday

AI HOT (Curated Pool)

Qwen 3.7 Preview

The title identifies Qwen 3.7 Preview, but the post body is empty; it does not disclose parameter size, context window, release timing, access method, pricing, model card details, or benchmark results.

Why it matters: HKR-H/R pass because an official Qwen 3.7 preview matters for flagship-model competition, but HKR-K fails: no size, context, access path, or benchmarks are disclosed, so it stays at the featured threshold.

r/LocalLLaMA

llama.cpp MTP support landed: Qwen3.6 27B reaches 2.44× on Strix Halo

llama.cpp merged MTP speculative decoding in PR #22673; Qwen3.6 27B Q8_0 rose from 7.4 to 18.1 tok/s on Strix Halo, while a dual RTX 3090 Q8_0 setup rose from 25.7 to 55.9 tok/s.

Why it matters: HKR-H/K/R all pass: llama.cpp adds MTP speculative decoding with Qwen3.6 27B speedups on Strix Halo and RTX 3090. The scope is local inference, not a broad model release, so 78 fits featured.

Hacker News front page

We Let AIs Run Radio Stations

Andon Labs gave four AI agents tools to host live radio shows and run a media business without humans. The post says revenue is terrible, but it does not disclose amounts or operating metrics.

Why it matters: HKR-H is strong from the AI-run radio premise; HKR-K has a concrete 4-agent experiment but weak revenue disclosure; HKR-R lands on agent economics and media automation. Niche lab post, so it sits at the featured threshold, not same-day news.

AI HOT (Curated Pool)

Claude Console Adds Prompt Cache Diagnostics

Anthropic added prompt cache diagnostics to the Claude Console; when a request misses the cache, developers can see which prompt segment changed and how many tokens it consumed.

Why it matters: HKR-H/K/R pass: this is a small but practical Anthropic Claude developer-console update with concrete cache-miss diagnostics and token-cost visibility, enough for the featured threshold but below major product-release weight.

AI HOT (Curated Pool)

Cursor releases Composer 2.5 coding model

Cursor released Composer 2.5, claiming up to 10x higher efficiency on long coding tasks; the model is further trained on Moonshot’s Kimi K2.5 and uses text feedback for 100k-token-scale trajectories.

Why it matters: HKR-H/K/R all pass: Cursor is a core AI coding tool, and Composer 2.5 adds concrete claims around 10x long-task gains and Kimi K2.5 tuning. Limited sourcing and no independent eval keep it in the 78–84 band.

AI HOT (Curated Pool)

Take your local GitHub sessions anywhere

GitHub launched remote control sessions for Copilot, letting users start tasks in VS Code or the command line and continue them through github.com or GitHub Mobile.

Why it matters: GitHub Copilot session handoff from VS Code/CLI to web and mobile clears HKR-H/K/R, but the post only gives entry points and use case; permissions, pricing, and supported task scope are not disclosed.

AI HOT (Curated Pool)

NVIDIA fine-tunes Cosmos Predict 2.5 with LoRA/DoRA for robot video generation

NVIDIA published a Hugging Face post on fine-tuning Cosmos Predict 2.5 with LoRA and DoRA to generate robot first-person videos from text prompts; the post does not disclose dataset size, training cost, or evaluation results.

Why it matters: HKR-H/K/R pass: the robot POV video angle is clickable, and LoRA/DoRA on Cosmos Predict 2.5 is a concrete mechanism. Missing dataset scale and metrics keep it in the low featured band.

May 18Monday

Hacker News front page

Show HN: InsForge – Open-source Heroku for coding agents

InsForge released an Apache 2.0 backend platform that lets coding agents deploy, operate, and debug backend systems through one CLI install command and Skills.

Why it matters: HKR-H/K/R all pass: the Heroku-for-agents framing, Apache 2.0 plus one-CLI install, and agent ops pain are concrete. Source is mainly Show HN/GitHub with no usage, benchmark, or production proof, so it sits at the featured threshold.

AI HOT (Curated Pool)

Baidu Core AI Business Revenue Exceeded RMB 13.6B in Q1

Baidu reported that its core AI-driven business generated more than RMB 13.6 billion in Q1 2026, up 49% year over year, and accounted for more than half of Baidu’s general business revenue for the first time.

Why it matters: HKR-H/K/R all pass, but the source is a Baidu post with revenue framing only; segment breakdown and margin quality are not disclosed. This fits the low featured band as an AI commercialization signal.

QbitAI · WeChat

openJiuwen open-sources JiuwenSwarm, a multi-agent swarm coordination framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Skills Hub, and self-evolution, and the framework supports HOTS and HITS modes for human participation in multi-agent workflows.

Why it matters: HKR-H/K/R pass: the swarm angle is clickable, the post gives four modules plus HOTS/HITS, and agent builders care about orchestration choices. Lacking benchmarks or adoption data keeps it at the featured threshold.

Bloomberg Technology

Baidu AI Sales Eclipse Waning Legacy Ads for the First Time

Baidu reported a 1% revenue decline as growth in nascent AI businesses offset shrinking traditional internet revenue; the post does not disclose AI sales, advertising revenue, or details of the agentic AI pivot.

Why it matters: Baidu revenue fell 1% while AI sales topped legacy ads for the first time, so HKR-H/K/R pass. Missing AI/ad dollar splits and agentic-AI mechanics keep it in the low featured band, not p1.

AI HOT (Curated Pool)

Grok Now Supports Video Understanding and Analysis

Grok now supports full-video uploads for real-time analysis, summarization, translation, scene explanation, and context extraction; the post does not disclose duration limits, supported formats, or rollout scope.

Why it matters: HKR-H/K/R all pass, but duration limits, formats, and rollout scope are not disclosed, so this stays at the featured threshold for a mid-weight product update.

Synced · WeChat

openJiuwen releases JiuwenSwarm, an open-source multi-agent swarm framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Swarm Skills Hub, and self-evolving Swarm Skills, and reports a 94.2% PinchBench score versus 91.6% for OpenClaw.

Why it matters: HKR-H/K/R all pass: an open-source agent-swarm framework with named components and a PinchBench 94.2% claim. It stays at 78 because openJiuwen is not a top lab and the summary lacks license, reproduction setup, and baselines.

AI HOT (Curated Pool)

Tencent AI Design Agent Ardot Enters Public Beta: Generates Editable Designs and Converts Them to Code

Tencent Cloud opened public beta for Ardot, an AI design agent that generates editable app pages, websites, and posters from one-sentence prompts, then converts designs to code.

Why it matters: HKR-H/K/R pass on a concrete Tencent product beta for editable design-to-code workflows. Missing pricing, model details, benchmarks, and field results keep it at the lower featured threshold.

AI HOT (Curated Pool)

Alibaba Cloud launches HappyHorse video generation model

Alibaba Cloud launched HappyHorse on Model Studio, with prompt-to-1080p multi-shot video generation in one workflow; the post lists a limited-time 20% discount but does not disclose pricing, model parameters, or availability terms.

Why it matters: HKR-H/K/R pass on the named model, 1080p multi-shot capability, and cost/competition angle. Thin disclosure on price, parameters, and benchmarks keeps it near the featured threshold.

AI HOT (Curated Pool)

Open-source tool exposes security risks and detection gaps in AI API relays

api-relay-audit audits AI API relay risks with verifiable three-state decisions and transparent logs, covering AC-1 tool-call rewriting, AC-2 error-response leakage, and context truncation, while the author has published the methodology, comparison results, quick-reference table, and the open-source tool.

Why it matters: HKR-H/K/R all pass because the tool targets real AI API relay risks with concrete checks. Source is a single X post, and adoption or incident data is not disclosed, so it stays in the low featured band.

AI HOT (Curated Pool)

Grok launches Skills feature

xAI launched Grok Skills on May 18, 2026, letting users set preferences, formatting rules, or workflows once and keep them active across all conversations on web, iOS, and Android.

Why it matters: HKR-H/K/R all pass: Grok Skills adds persistent preferences and workflows across web, iOS, and Android. This is a mid-weight xAI product update; rollout scope, limits, and pricing are not disclosed.

AI HOT (Curated Pool)

Composer 2.5 release and technical analysis

Cursor released Composer 2.5, built on a Moonshot open-source checkpoint, trained with synthetic data from real codebases at 25 times the previous scale, and updated with text-feedback reinforcement learning and a sharded Muon optimizer.

Why it matters: HKR-H/K/R all pass: Cursor is a core coding-agent surface, and the post gives concrete training details around Moonshot, 25x data, RL, and Muon. It lacks benchmarks, pricing, or user-facing capability limits, so it stays in the 78–84 band.

Google DeepMind

Google DeepMind adds Street View grounding to Project Genie

Google DeepMind has added Street View real-scene grounding to its experimental prototype Project Genie. Users can pick a US location, then pair it with a style and characters to generate a world.

Why it matters: With Street View imagery wired in, agents and robots can train and navigate in virtual environments that track real places.

Google DeepMind

Introducing Google Antigravity 2.0

Google 发布智能体开发平台 Google Antigravity 2.0。该平台在 Google DeepMind 官网被列为面向开发者的 agentic development platform,与 Gemini 应用、Google AI Studio 并列。原文未披露版本功能、参数或可用性细节。

May 17Sunday

Hacker News front page

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

MinishLab open-sourced Semble, a code-search tool for agents that combines Model2Vec embeddings, BM25, RRF fusion, and reranking; on a 63-repo benchmark, it used 98% fewer tokens than grep+read, reached 0.854 NDCG@10, and ran CPU queries in about 1.5 ms.

Why it matters: HKR-H/K/R all pass: the 98% token claim is clickworthy, the 63-repo benchmark adds substance, and coding-agent context cost is a real practitioner nerve. Impact is still toolchain-level, so it stays below must-write.

Bloomberg Technology

Apple’s New ChatGPT-Like Siri App Will Have Auto-Deleting Chats

The title says Apple’s ChatGPT-like Siri app will support auto-deleting chats; the RSS snippet only adds that iOS 27 will include a Genmoji upgrade, and the post does not disclose retention periods, release timing, or feature details.

Why it matters: HKR-H and HKR-R pass because Bloomberg frames a specific Apple Siri privacy angle; HKR-K fails since retention and feature mechanics are missing, so this stays at the low featured threshold.

Google DeepMind

Google DeepMind launches Gemini for Science toolset

Google DeepMind released Gemini for Science, which includes three experimental tools on Google Labs: Hypothesis Generation, built on Co-Scientist.

Why it matters: Google is packaging research prototypes like Co-Scientist and AlphaEvolve into apply-to-use science tools, showing what agentic research looks like in practice.

Google DeepMind

Google expands content provenance and verification tools across Search, Gemini, Chrome and Pixel

Google is widening its content transparency and verification tools across Search, Gemini, Chrome, Pixel and Cloud, and deepening industry partnerships. SynthID has watermarked over 100 billion images and videos plus 60,000 years of audio. SynthID verification in the Gemini app has been used 50 million times, and the capability reaches Search today, with Chrome in the coming weeks.

Why it matters: The post lays out where SynthID and C2PA land across Search, Gemini, Chrome and Pixel, which shows the current limits of content provenance tools.

QbitAI · WeChat

A Robot Dog Challenges Nvidia's Compute Lead

Weilan Technology unveiled BabyAlpha A3, a consumer quadruped robot using a six-chip heterogeneous cluster that runs a 7B-parameter model on-device at 280 TPS; the article says it has 66MP vision, 2.232 million point-cloud samples per second, and a planned Q3 launch.

Why it matters: HKR-H/K/R pass: the robot-dog-versus-Nvidia angle is clickable, and 280 TPS on a local 7B model is concrete. Single-source summary lacks price, power draw, and benchmark setup, so it stays near the featured floor.

AI HOT (Curated Pool)

Grok Imagine image generation is officially released

Grok Imagine is now available on X for all users, with text-to-image generation for realistic images and multiple aspect ratios; the post does not disclose model parameters, pricing, or regional limits.

Why it matters: HKR-H/K/R pass, but the post only discloses availability and basic image features; model details, pricing, and regions are absent, so this lands at the featured threshold.

AI HOT (Curated Pool)

MagicPath Integrates with Codex to Combine Design and Development

MagicPath AI CEO @skirano demonstrated MagicPath running inside Codex as a native canvas, with users configuring it through one command, dragging UI elements, and letting Codex generate and edit code in real time.

Why it matters: HKR-H/K/R pass: MagicPath puts a draggable design canvas inside Codex with one-command setup and live code edits. Single-demo sourcing and missing framework support, permissions, and reproducible cases keep it at the lower featured band.

May 16Saturday

AI HOT (Curated Pool)

Codex adds multi-device remote control and shared context

Codex controls multiple devices through ChatGPT, switches by project to access each device’s context and files, and supports remote SSH setup for other VMs.

Why it matters: HKR-H/K/R all pass, but the item is a thin X-post summary with no official release note, pricing, permission model, or reproducible demo. Treat it as a mid-weight coding-agent product update at the featured threshold.

AI HOT (Curated Pool)

OpenAI Restructures as Brockman Takes Over Product Strategy

OpenAI merged ChatGPT, Codex, and API into one product organization, with Greg Brockman taking over product strategy; the post says Anthropic’s valuation reached $900 billion, but it does not disclose the restructuring timeline.

Why it matters: HKR-H/K/R all pass: this is an OpenAI top-level product reorg covering ChatGPT, Codex, and API. Single-source summary keeps it below the highest band, but it is same-day must-write news.

QbitAI · WeChat

Codex Integrates HeyGen for Prompt-Based Video Generation and Editing

Codex integrates the HeyGen plugin to run image generation, talking-avatar video, subtitles, and edits from natural-language prompts; the article tests roughly one-minute avatar generation, trimming content after 10 seconds, and deleting a blink at the eighth second.

Why it matters: HKR-H/K/R all pass, backed by a numbered hands-on test. The scope is still one Codex-to-HeyGen plugin workflow, not a model or platform release, so it lands in the 72-77 featured band.

QbitAI · WeChat

A new AI for 5 million doctors in China: exclusive journal partnership focuses on evidence sources

Alibaba Health launched the medical AI product Qinglizi for China’s 5 million doctors, with access to ten years of content from 70 BMJ Group journals and an evidence workflow constrained by PICO, GRADE, and review from more than 300 clinical experts.

Why it matters: HKR-H/K/R all pass: Alibaba Health and BMJ add concrete evidence sources and review mechanisms to a medical AI product. It remains a vertical product/partnership update, not a foundation-model or platform release.

Computing Life · Share · Yage

OpenAI Reaches Into Your Bank Account

OpenAI uses Plaid to let ChatGPT connect to bank accounts; the post does not disclose launch timing, authorization flow, or the exact data scope ChatGPT can access.

Why it matters: HKR-H/R are strong and HKR-K passes via the Plaid integration mechanism. Missing launch timing, authorization flow, and data scope keep it at the featured threshold rather than a higher OpenAI product-update score.

AI HOT (Curated Pool)

Ignoring Token Costs, Using 100 AI Instances to Automate an Open Source Project

The OpenClaw team runs about 100 Codex instances to handle code review, security analysis, issue deduplication, test reproduction, task creation from meetings, spam filtering, and performance regression monitoring.

Why it matters: HKR-H/K/R all pass: 100 Codex instances running open-source maintenance is a strong operational anecdote with concrete task types. Single X post, no cost, outcome metrics, or reproducible setup, so it stays in the lower featured band.

AI HOT (Curated Pool)

Runway Agent Generates Complete Ads in One Session

Runway Agent turns product photos and ideas into fully produced ads in one session; the post does not disclose the model, pricing, generation length, or regional availability.

Why it matters: Runway’s ad-generation Agent clears HKR-H/K/R as a mid-weight product update. Missing model, pricing, duration, and region details keep it at the featured threshold, not a must-write release.

The Verge · AI

OpenAI now wants ChatGPT to access your bank accounts

OpenAI previewed a ChatGPT feature that lets users connect financial accounts through Plaid, which links to 12,000 institutions. OpenAI says more than 200 million people ask ChatGPT finance questions each month; the post does not disclose a general release date.

Why it matters: HKR-H/K/R all pass: OpenAI is moving ChatGPT toward real financial-account access, with Plaid’s 12,000 institutions and 200M monthly finance askers as concrete facts. It stays below 85 because launch timing is not disclosed.

TechCrunch · AI

OpenAI launches ChatGPT for personal finance, will let users connect bank accounts

OpenAI launched ChatGPT for personal finance, and connected users can view portfolio performance, spending, subscriptions, and upcoming payments; the RSS snippet does not disclose supported banks, launch regions, pricing, or account-security terms.

Why it matters: HKR-H is strong because ChatGPT connects to bank accounts; HKR-K has concrete finance features; HKR-R hits privacy and fintech competition. Banks, regions, and pricing are undisclosed, so this stays in the low P1 band.

May 15Friday

r/LocalLLaMA

Fully Offline Suitcase Robot Built Around Jetson Orin NX SUPER 16GB

CreativelyBankrupt built Sparky as a fully offline suitcase robot on Jetson Orin NX SUPER 16GB, running Gemma 4 E4B Q4_K_M via llama.cpp with q8_0 KV cache, about 200 ms cached TTFT, 14-15 tok/s sustained output, 12K context, 30+ sensors, and no WiFi, Bluetooth, or cellular interface.

Why it matters: HKR-H/K/R all pass, with a named hands-on build and concrete latency/sensor numbers. It stays in low featured because this is a Reddit project post, not a product launch or research release.

AI HOT (Curated Pool)

X open-sources the “For You” feed recommendation algorithm

X open-sourced the For You recommendation pipeline on GitHub, using a Grok-based Phoenix Transformer to score candidate posts and predict engagement probabilities such as likes, replies, and reposts.

Why it matters: HKR-H/K/R all pass, but the item only gives the open-source claim and Phoenix Transformer ranking mechanism; repo details, license, and reproducible tests are not disclosed, so it stays low-featured.

AI HOT (Curated Pool)

Feishu Open-Source CLI Tool Gets 10,000 Stars in 45 Days with Visible AI Operations

Feishu’s open-source lark-cli gained over 10,000 GitHub stars in 45 days, letting AI create groups and documents through the command line with each operation previewable and reviewable.

Why it matters: HKR-H/K/R all pass, but the source is a single social post and lacks usage, contributor, or adoption data. This fits the lower featured band for an open-source agent tool update.

MIT Technology Review · AI

The Download: China’s AI Drama Factory and the WHO’s Missing Health Targets

China’s short-drama industry released an average of 470 AI-generated short dramas per day in January, while production timelines fell from months to weeks and costs dropped by up to 90%.

Why it matters: MIT Technology Review provides concrete output, cycle-time, and cost figures for China’s AI short-drama pipeline, clearing HKR-H/K/R. The story is application-layer, not a core model or product release, so it sits at the featured threshold.