Skip to content

Product updates

New features, redesigns and pricing in AI products — whose product got better, pricier or finally usable.

842 picksRelated topicsModel releasesIndustryAI coding

Latest picks

221–240 of 842

May 26Tuesday

QbitAI · WeChat

Chinese AI-Written Pretraining Framework ForgeTrain Trains MiniCPM5-1B

ModelBest released ForgeTrain and MiniCPM5-1B, saying ForgeTrain was written by AI and trains 10% faster than NVIDIA Megatron under the same hardware conditions. MiniCPM5-1B is a 1B-parameter edge model with about 2GB FP16 weights and about 0.5GB INT4/Q4 weights.

Why it matters: HKR-H/K/R all pass: an AI-written trainer, a 10% same-hardware Megatron speed claim, and a 0.5GB 1B edge model are concrete hooks. Score stays at 80 because the first-ever claim and benchmark lack third-party reproduction.

Synced · WeChat

Grok keeps updating after xAI disbandment as Musk announces a new model

Elon Musk said the 1.5T-parameter Grok V9-Medium has finished training, will enter reinforcement learning in a few days, and is planned for release in two to three weeks. Grok Build supports up to 8 parallel sub-agents, a 256K-token context window, Plan Mode, Arena Mode, MCP, and ACP.

Why it matters: HKR-H/K/R all pass, but this is a Grok V9-Medium preview before RL and release, with no benchmarked capability yet. That fits a strong model-race/product update at 82, featured but not p1.

Synced · WeChat

AI-written training framework trains 1B edge model MiniCPM5-1B

ModelBest open-sourced MiniCPM5-1B and ForgeTrain; the 1B edge model scores 17.9 on AA-Index, while the AI-written ForgeTrain framework matches Megatron’s training results and runs 10% faster on Nvidia H100 under the article’s reported setup.

Why it matters: HKR-H/K/R all pass: the AI-written training framework hook is strong, with concrete AA-Index and H100 speed claims. It is not a flagship model release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Chinese agent SkyClaw targets Opus 4.6-level performance with free trial

Kunlun Tech released SkyClaw-v1.0 and SkyClaw-v1.0-lite with a 2-4 week free trial, claiming SkyClaw-v1.0 input costs are 1/24 of DeepSeek V4 Pro and about 1/43 of Sonnet 4.6.

Why it matters: HKR-H/K/R all pass: SkyClaw-v1.0 has a sharp cost hook, concrete trial and pricing ratios, and budget resonance. Source facts remain vendor claims, so it stays at the low featured band.

AI HOT (Curated Pool)

OpenAI GPT-5.6 Reportedly Set for Next Month With 1.5M-Token Context

Developers found an unannounced OpenAI GPT-5.6 entry in Codex backend logs under the codename iris-alpha, with a 1.5 million-token context window, about 43% higher than GPT-5.5’s 1.05 million-token limit.

Why it matters: HKR-H/K/R all pass: the Codex-log leak, 1.5M-token window, and 43% increase are concrete and practitioner-relevant. It stays below 85 because this is not an official GPT-5.6 launch.

r/LocalLLaMA

Update on a 12×32GB SXM V100 Cluster for Local Legal Drafting

A lawyer runs a local legal-drafting pipeline across 16 GPUs, with Qwen3.5-122B-A10B reaching about 50 tok/s on four V100s, while a verifier blocks ungrounded citations, dates, and Bates numbers before any final document is used.

Why it matters: HKR-H/K/R all pass: this is a first-person local-LLM experiment with concrete numbers, not a vendor post. Reddit source limits authority, so it stays at the low featured band rather than p1.

AI HOT (Curated Pool)

Apple reportedly uses a custom 1.2T-parameter Google model for next-generation Siri

Apple is reportedly using a custom 1.2T-parameter Google model to run parts of the next-generation Siri, while simpler queries are expected to run on-device; the post says response speed for everyday questions is the key constraint.

Why it matters: HKR-H/K/R all pass, but this is a single X-sourced reported claim; the post gives architecture details but not sourcing documents, rollout timing, or scope. Keep it at the featured threshold, below the 78+ band.

AI HOT (Curated Pool)

Grok Build Beta Opens to SuperGrok Users

xAI opened Grok Build Beta to all SuperGrok and X Premium+ users, with Plan Mode, Imagine-based image and video creation, and a CLI for automation or orchestrator workflows at x.ai/cli.

Why it matters: HKR-H/K/R all pass: xAI opened a paid beta with named workflow features. The score stays at the featured floor because the post lacks capability limits, pricing detail, and test results.

May 25Monday

QbitAI · WeChat

Reasonix for DeepSeek V4 reaches 99.82% cache hit rate and cuts costs to 20%

Reasonix uses an append-only loop for DeepSeek V4 and reports a 99.82% cache hit rate in long coding sessions, cutting an example 400M-token bill from $61 to $12.

Why it matters: HKR-H/K/R all pass, but this is a third-party cost tool around DeepSeek V4, not a model launch or platform update. Concrete mechanism and billing numbers put it in the 72–77 featured band.

Bloomberg Technology

Huawei Touts Chipmaking Breakthrough to Shorten Gap With TSMC

Huawei said it has found a new path to narrow the gap with TSMC and make advanced semiconductors without advanced equipment; the RSS snippet does not disclose the process node, yield, cost, or production timeline.

Why it matters: Bloomberg authority plus the Huawei-vs-TSMC chip-gap angle clears HKR-H and HKR-R. HKR-K fails because node, yield, and production timing are not disclosed, so it sits at the featured threshold, not 78+.

May 24Sunday

r/LocalLLaMA

Using llama.cpp native tools for web RAG inside llama-server WebUI

A Reddit user describes using llama.cpp native tools for web RAG inside llama-server WebUI with a 7-step setup: enable get_datetime and exec_shell_command, then run wget through firejail, a separate Linux user, and an Alpine OCI VM sandbox.

Why it matters: HKR-H/K/R all pass: the post gives a concrete local web-RAG recipe with sandboxing. It is a community tutorial, not a model or product launch, so the narrow reach and source authority keep it at the low featured band.

Xinzhiyuan · WeChat

Anthropic’s Three Cards Surface: Mythos 1 Appears, Opus 4.8 Spotted

Xinzhiyuan says Anthropic’s claude-opus-4.8 appeared in Google Vertex AI, while a 59.8MB Claude Code source-map leak with 512,000 TypeScript lines exposed Sonnet 4.8 references and Mythos 1 clues tied to Claude Code and Claude Security.

Why it matters: HKR-H/K/R all pass, but this is a leak plus Vertex listing, not an Anthropic launch. No capability numbers, pricing, context window, or reproducible evals, so it stays in the 78–84 band.

r/LocalLLaMA

llama.cpp server has built-in native tools: exec_shell, edit_file, and more

llama.cpp server exposes an experimental --tools flag with 8 native tools, including file reads, grep search, shell execution, file edits, diffs, and datetime; the post says file operations are relative to the server launch directory and no command whitelist or strict sandbox is provided yet.

Why it matters: HKR-H/K/R all pass: llama.cpp adding native shell and file tools is a concrete agent-runtime shift with safety stakes. Reddit sourcing and experimental status keep it in the lower featured band.

AI HOT (Curated Pool)

StepAudio 2.5 Realtime Voice Released with Paralinguistic Awareness and Persona Interaction

StepFun released StepAudio 2.5 Realtime with Chinese and English real-time voice support, API-based custom personas, more than 10,000 native persona options, millions of composable traits, and 5 built-in preset personas.

Why it matters: HKR-H/K/R all pass, but the source is an official X post and lacks latency, pricing, benchmarks, and rollout scope. This fits the low featured band for a mid-weight product update.

May 23Saturday

The Verge · AI

Google’s New Anything-to-Anything AI Model Is Wild

The Verge tried Google’s new Gemini anything-to-anything model for a stuffed-deer deepfake video, but the RSS snippet discloses only one example and does not disclose model parameters, pricing, release timing, or safety controls.

Why it matters: HKR-H/R pass: a Google/Gemini multimodal hands-on has a strong deepfake hook and safety resonance. HKR-K fails because the feed discloses one example only, with no params, pricing, or launch timing.

Bloomberg Technology

DeepSeek To Make Permanent 75% Discount on Flagship AI Model

DeepSeek will make a 75% discount on its flagship AI model permanent, but the post does not disclose the specific model name, original price, discounted price, or effective date.

Why it matters: HKR-H/K/R pass on a concrete 75% permanent discount from DeepSeek, a cost and price-war story. Sparse extracted body lacks model name, list price, discounted price, and timing, so it stays in low featured.

r/LocalLLaMA

meituan-longcat/LongCat-Video-Avatar-1.5 on Hugging Face

Meituan LongCat released LongCat-Video-Avatar-1.5 on Hugging Face, supporting AT2V, ATI2V, and video continuation while replacing Wav2Vec2 with Whisper-Large and using DMD2 distillation to reduce inference to 8 NFE; the model weights are released under the MIT License.

Why it matters: HKR-H/K/R all pass: open MIT video-avatar weights plus 8 NFE inference give local multimodal builders real signal. This is a mid-weight open-source model update, not an 85+ same-day industry event.

AI HOT (Curated Pool)

Gemini update: over 900 million users and new agent features

Google announced that the Gemini app has surpassed 900 million monthly active users and introduced two agent features: Daily Brief for personalized daily summaries and Gemini Spark, a 24/7 personal agent that manages tasks under user authorization.

Why it matters: HKR-H/K/R all pass: Google gives a 900M MAU number and two agent features for Gemini. This is an entry-point product update with competitive weight, not a routine small feature.

AI HOT (Curated Pool)

v2.1.149 release summary

Claude Code v2.1.149 adds categorized /usage reporting, an enterprise allowAllClaudeAiMcps setting for cloud MCP connectors, and fixes three security issues involving PowerShell permission bypass, Git worktree sandbox allowlist overflow, and otelHeadersHelper failures when script paths contain spaces.

Why it matters: Official Claude Code point release with concrete changes but limited blast radius: /usage categories, an enterprise MCP allow switch, and PowerShell bypass fixes hit developer security and governance needs.

AI HOT (Curated Pool)

Claude Auto Mode Adds Pro Plan and Model Support

Claude Auto Mode is now available on the Pro plan and supports Sonnet 4.6 and Opus 4.7; users can start it with Shift+Tab, while the post does not disclose pricing changes or rollout scope.

Why it matters: HKR-H/K/R all pass: official Claude dev channel gives Pro access, two supported models, and a shortcut. This is a mid-weight Claude product update, not a major model or capability release.