Skip to content

All news

25 today

May 7Thursday

QbitAI · WeChat

Vidu Claw Generates Ad Videos From One Prompt and a Hundred-Yuan Budget

Shengshu Technology opened Vidu Claw, which generates ad scripts, voiceover, music, editing, and final videos from one prompt; its Video Plan includes up to 40 minutes of daily generation across video, image, and audio.

Why it matters: HKR-H has a concrete ad-test hook, HKR-K adds the 40-minute daily quota and one-prompt workflow, and HKR-R hits production-cost pressure. No benchmark or pricing detail, so this stays at the featured threshold.

Ben's Bites

Elon Doubled Limits

Ben’s Bites says Anthropic doubled Claude usage on paid plans via SpaceX’s Colossus 1. The issue also lists GPT-5.5 Instant, ChatGPT spreadsheet integration, and three Claude Managed Agents features. The title names Elon, but the post does not disclose exact limits.

Why it matters: HKR-H/K/R pass: the SpaceX Colossus 1 angle, 2x Claude usage, and quota pressure are all concrete. Missing exact caps, pricing, and rollout scope keep it in the low featured band.

OpenAI News

Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber

OpenAI expanded Trusted Access for Cyber to GPT-5.5 and GPT-5.5-Cyber. The RSS snippet says access is for verified defenders; the post does not disclose criteria, pricing, or benchmark data.

Why it matters: HKR-H/K/R all pass: OpenAI expands trusted cyber access to GPT-5.5 and GPT-5.5-Cyber. Kept below 85 because admission rules, pricing, evals, and reproducible tests are not disclosed.

AI HOT (Curated Pool)

Consistent web search and scraping for all models

OpenRouter released tools for tool-calling models to run web search and page scraping. The post says multiple search and scraping engines are supported, but does not disclose names, pricing, or limits. The key item is cross-model tool interface consistency.

Why it matters: HKR-H/K/R all pass, but engines, pricing, and limits are not disclosed. This is a mid-weight Product update: useful for model-agnostic agent stacks, not a major model or capability release.

r/LocalLLaMA

Running Qwen3.5/Qwen3.6 with NextN MTP in llama.cpp on one RTX 3090 Ti

A Reddit user posted a llama.cpp guide for Qwen3.5/3.6 with NextN MTP on one RTX 3090 Ti. It requires two unmerged PRs, #22400 and #22673; Qwen3.6-35B-A3B-MTP reaches 157 tok/s at 350W, 1700MHz, with q8 KV. The key reproducible detail is nextn=q8_0 quant override; missing it yields “////” output.

Why it matters: HKR-H/K/R all pass: single-GPU 157 tok/s is a strong hook, and the PR/power settings make it testable. Scope stays narrow because it is a Reddit guide using unmerged PRs.

Xinzhiyuan · WeChat

Claude Managed Agents Add Dreaming, With Reported Task Completion Up to 6x

Anthropic added Dreaming, Outcomes, and multi-agent orchestration to Claude managed agents; Harvey reports about 6x higher task completion. Dreaming reads up to 100 sessions; one demo distilled 5.3M tokens into 98 rules, while Outcomes raised success by up to 10 points. Opus 4.7 and Sonnet 4.6 require access, with $0.08 per session-hour runtime fees.

Why it matters: HKR-H/K/R all pass: Anthropic adds Dreaming, Outcomes, and multi-agent orchestration with 100-session memory, $0.08/session-hour runtime, and Harvey’s ~6x completion claim. This is a same-day Claude agent update.

AI HOT (Curated Pool)

Amp releases Neo CLI as coding agents shift toward long-horizon workflows

Amp released Neo, a CLI tool covering remote orchestration, automatic context compression, and a Plugin API. Neo lets local threads be controlled remotely, allows all operations by default, and moves safety control to plugins; the post does not disclose version, pricing, or performance gains.

Why it matters: HKR-H/K/R all pass: Neo adds remote orchestration, context compression, Plugin API, and default-allow permissions. Amp’s reach and missing price/version/perf data keep it in the 72–77 band.

AI HOT (Curated Pool)

Open Slide lets AI write PPT code

Open Slide builds PPTs with React, using a workflow designed for AI agents. It integrates SVGL with 1,500+ brand logos, supports manual edits, and lets AI read user comments for revisions.

Why it matters: HKR-H/K/R pass: the programmable-slide angle is clickable, with concrete React and 1500+ logo details, and deck work is a real practitioner pain. No usage metrics or hands-on test keeps it at the featured threshold.

OpenAI News

Introducing Trusted Contact in ChatGPT

OpenAI introduced Trusted Contact in ChatGPT, notifying a trusted person when serious self-harm concerns are detected. The feature is optional; the post does not disclose detection mechanics, contact setup, or rollout scope.

Why it matters: HKR-H/K/R all pass: the ChatGPT safety hook is concrete and emotionally charged. Importance stays in the low featured band because detection, setup, and rollout details are not disclosed.

Financial Times · Technology

Arm projects $2bn in sales of its new AI chip from next year

Arm projects $2bn in sales for its first in-house AI chip from next year. The RSS snippet says the SoftBank-backed UK group has strong demand; the post does not disclose customers, pricing, process node, or delivery cadence.

Why it matters: FT authority plus a $2bn sales projection gives HKR-K, and Arm’s own AI chip adds HKR-H/R. Missing customers, process, price, and delivery cadence keep it at the featured threshold.

The Verge · AI

Google shuts down Project Mariner

Google shut down Project Mariner on May 4, 2026. The experimental web-task agent once supported up to 10 concurrent tasks. Its technology moved into Google products, including Gemini Agent.

Why it matters: HKR-H/K/R all pass, but the disclosed facts are limited to shutdown timing, a 10-task limit, and migration into Gemini Agent. Strong source authority supports low featured, not a major launch.

r/LocalLLaMA

GB10 inference engine Atlas is open source, with Qwen3.6-35B-FP8 over 100 tok/s

Avarok open-sourced Atlas, an inference engine running Qwen3.5-35B at ~111 tok/s sustained on one DGX Spark. It uses Rust+CUDA, a ~2.5GB image, and sub-2-minute cold start; the author claims 3.0–3.3x vLLM in tests. The key details are Blackwell SM120/121 kernels, NVFP4/FP8, and MTP decoding.

Why it matters: HKR-H/K/R pass: open-source inference engine, 35B FP8 at 111 tok/s, and a direct vLLM comparison. Single Reddit sourcing and unreproduced benchmarks keep it at the lower featured band.

Hacker News front page

Higher usage limits for Claude and a compute deal with SpaceX

Anthropic’s post has 91 HN points; the title says Claude gets higher usage limits and a SpaceX compute deal. The RSS body only lists links, 37 comments, and HN metadata. The post does not disclose limit multiples, compute scale, pricing, or timing.

Why it matters: Official title gives HKR-H/R: higher Claude limits and a SpaceX compute deal. HKR-K fails because the feed omits limit multiples, compute scale, pricing, and rollout timing.

May 6Wednesday

r/LocalLLaMA

CopilotKit (MIT): Open-source building blocks for agent apps and generative UI

CopilotKit offers MIT-licensed React components and claims 30k GitHub stars. It covers chat, streaming, tool calls, HITL, and generative UI, with AG-UI support for LangGraph, CrewAI, LlamaIndex, and other backends. The key point is decoupling the UI layer from agent frameworks.

Why it matters: HKR-H/K/R all pass: MIT open source, 30k stars, and AG-UI links to major agent backends. Kept in 72–77 because the post lacks a new version, benchmark, or named production adopter.

The Verge · AI

Google’s AI Search Summaries Will Now Quote Reddit

Google updated AI Search to include firsthand views from Reddit, social media, and forums in summaries. The post says a “perspectives” preview links queries to related online discussions; it does not disclose rollout scope or timing. For search teams, the key issue is how AI summaries cite and rank UGC sources.

Why it matters: HKR-H is strong because Google AI summaries quoting Reddit alters the search surface. HKR-K has the perspectives mechanism, and HKR-R hits SEO/UGC traffic concerns; missing rollout scope keeps it in the 72–77 product-update band.

NVIDIA Blog

NVIDIA Spectrum-X AI-Native Ethernet Fabric Adds MRC for Gigascale AI

NVIDIA added MRC support to Spectrum-X Ethernet, letting one RDMA connection spread traffic across multiple paths. MRC ran in Blackwell deployments, with microsecond failure bypass and hardware rerouting. The key detail is the OCP open specification and multiplane support for clusters up to hundreds of thousands of GPUs.

Why it matters: HKR-K/R are solid: MRC stripes one RDMA flow across paths, detects failures in microseconds, and is tied to Blackwell deployments. HKR-H is narrow and the source is vendor-owned, so this stays below major release level.

QbitAI · WeChat

Boston Dynamics executives exit as Atlas output is reported at four units per month

Boston Dynamics showed a new Atlas gymnastics demo, while the post says output is only four units per month. Atlas has 56 DoF, weighs 90 kg, runs four hours, and 2026 capacity is allocated to Hyundai RMAC and Google DeepMind. The key issue is scale: Hyundai targets 30,000 units yearly, but today’s rate needs over 200 years for 10,000.

Why it matters: HKR-H, HKR-K, and HKR-R all pass: the hook is sharp, the piece has concrete production and spec numbers, and robotics scaling is a practitioner nerve. It stays below 85 because this is secondary reporting, not a major release.

r/LocalLLaMA

2.5x Faster Inference with Qwen 3.6 27B Using MTP on 48GB

A llama.cpp PR adds MTP support for Qwen 3.6 27B, with a reported 2.5x inference speedup. The author measured 28 tok/s on a Mac M2 Max 96GB and shared GGUF builds, compile steps, and a 262144-context server command. The key detail is turbo4 4.25-bit KV cache: a 48GB Mac runs Q5_K_M at 262K context.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the post names mechanisms and numbers, and local coding-agent cost resonates. Single Reddit source and setup complexity keep it in the low featured band.

Synced · WeChat

DeepSeek Version of Claude Code Tops Trending Chart With 8,700 Stars

DeepSeek TUI topped GitHub trending with over 8,700 stars. Hunter Bown built it in Rust for local terminal use with DeepSeek V4, supporting chat, file edits, shell commands, and task management. The key detail is RLM mode: up to 16 V4 Flash subtasks, plus a 1M-token context window and approval gates.

Why it matters: HKR-H/K/R all pass: the 8,700-star hook is strong, RLM adds concrete mechanisms, and coding-agent competition resonates. It is a third-party open-source tool, not an official DeepSeek model release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

GPT-5.5 Instant becomes ChatGPT’s free default model

OpenAI made GPT-5.5 Instant the default ChatGPT model, rolling it out free to all users. AIME 2025 rose from 65.4% to 81.2%, responses are 30.2% shorter, and hallucinations fell 52.5% versus GPT-5.3 Instant on high-risk prompts. Plus and Pro web users get chat, file, and Gmail personalization first; the API model ID is chat-latest.

Why it matters: HKR-H/K/R all pass: a free default ChatGPT model switch, concrete benchmark and behavior deltas, and direct impact on daily OpenAI workflows. This fits the 85–94 must-write band.

TechCrunch · AI

Apple plans to make iOS 27 a Choose Your Own Adventure of AI models

Apple reportedly plans to let iOS 27 users choose third-party AI models for multiple tasks. The RSS snippet does not disclose model names, task scope, launch timing, or integration mechanics.

Why it matters: HKR-H/K/R pass: system-level model choice on iOS has a strong platform hook and distribution stakes. Kept in 72–77 because the RSS summary lacks model names, task scope, launch timing, and API mechanics.

The Verge · AI

Apple could let you pick a favorite AI model in iOS 27

Apple plans to let third-party chatbots run system-wide Apple Intelligence in iOS 27, iPadOS 27, and macOS 27. Mark Gurman says Extensions can handle Siri, Writing Tools, and Image Playground this fall. The post does not disclose supported models, pricing, or developer APIs.

Why it matters: HKR-H/K/R all pass: the Apple system-level model picker is a strong hook, with named Extension targets. Scored 80 because model list, pricing, and developer APIs are not disclosed, and this remains a roadmap report.

Financial Times · Technology

Meta plans advanced agentic AI assistant for consumers

Meta plans a consumer agentic AI assistant; the RSS body has one sentence. It says Meta is funding an OpenClaw counterpart for everyday task execution. The post does not disclose model size, launch timing, pricing, regions, or permission controls.

Why it matters: FT reports Meta plans a consumer agentic assistant, with HKR-H/K/R present. Details on launch, pricing, model, and permission design are missing, so this sits at the lower featured band.

NVIDIA Blog

NVIDIA and ServiceNow Partner on Autonomous AI Agents for Enterprises

NVIDIA and ServiceNow expanded their partnership with Project Arc, an enterprise desktop agent. It connects via Action Fabric and uses OpenShell for sandboxed, policy-governed execution. Blackwell delivers over 50x Hopper’s token output per watt and nearly 35x lower cost per million tokens.

Why it matters: HKR-K/R pass: the post gives mechanisms and Blackwell economics. HKR-H misses because the angle is a standard vendor partnership, so this sits in the 72–77 featured-threshold band.

The Verge · AI

OpenAI claims ChatGPT’s new default model hallucinates way less

OpenAI says ChatGPT’s default GPT-5.5 Instant reduced hallucinations in internal evaluations. Versus GPT-5.3 Instant, hallucinated claims fell 52.5% on high-stakes prompts. Inaccurate claims fell 37.3% on flagged hard chats; the post does not disclose full eval size.

Why it matters: OpenAI changed ChatGPT’s default model and gave two hallucination-reduction figures, satisfying HKR-H/K/R. Internal evals lack set size and reproduction details, but a default ChatGPT model change is same-day material.

TechCrunch · AI

OpenAI releases GPT-5.5 Instant, a new default model for ChatGPT

OpenAI released GPT-5.5 Instant as ChatGPT’s new default model. The company says it reduces hallucinations in law, medicine, and finance while keeping prior low latency; the post does not disclose benchmarks, rollout scope, or pricing.

Why it matters: HKR-H/K/R all pass: a new ChatGPT default model, testable reliability claims, and direct workflow impact. Missing eval numbers, rollout scope, and pricing keep it in the mid 85–94 band.

r/LocalLLaMA

Gemma 4 MTP Released

Google released Gemma 4 MTP drafters with 4 Hugging Face checkpoints listed. MTP uses a smaller draft model to predict multiple tokens, then the target model verifies them in parallel, giving up to 2x decoding speedups with identical output quality.

Why it matters: HKR-H/K/R all pass: the practical hook is 2x lower-latency decoding, with 4 checkpoints and a clear speculative-decoding mechanism. It is a useful Gemma update, not a flagship model release, so 75 fits the featured lower band.

May 5Tuesday

Hacker News front page

Show HN: Airbyte Agents – context for agents across multiple data sources

Airbyte launched Airbyte Agents, using Context Store to index operational data for agents. Its public benchmark reports up to 80% fewer tokens for Gong and 90% for Zendesk versus vendor MCPs. The key point is pre-indexed context, not another MCP wrapper.

Why it matters: HKR-H/K/R all pass: a concrete pre-indexing angle, reproducible claims, and agent data-access pain. Airbyte is not a frontier lab, so this stays at the lower featured band.

The Verge · AI

OpenAI is reportedly launching a phone for ChatGPT

Ming-Chi Kuo says OpenAI is fast-tracking a ChatGPT phone for mass production in early 2027. It reportedly uses a customized MediaTek Dimensity 9600 with enhanced-HDR ISP; the post does not disclose price, design, or OS details.

Why it matters: HKR-H/K/R all pass, but this is a Kuo report rather than an OpenAI launch. Missing price, form factor, and OS details keep it below must-write territory.

TechCrunch · AI

Meta will use AI to analyze height and bone structure to identify underage users

Meta will use AI to analyze height and bone structure to identify underage users; the system runs in select countries. The post does not disclose countries, error rates, or appeals.

Why it matters: HKR-H comes from the biometric age-detection hook; HKR-K has a concrete mechanism; HKR-R hits privacy and child-safety concerns. Missing countries, false-positive rate, and appeals keep it in the low featured band.

OpenAI News

OpenAI Introduces MRC for Large-Scale AI Training Networks

OpenAI introduced MRC for large-scale AI training cluster networks. MRC stands for Multipath Reliable Connection and is released via OCP to improve resilience and performance; the post does not disclose throughput, latency, or cluster size.

Why it matters: HKR-H/K/R pass: OpenAI shared MRC via OCP, with a concrete multipath reliability mechanism. No throughput, latency, or cluster scale is disclosed, so this stays in the 72–77 featured band.

OpenAI News

GPT-5.5 Instant: smarter, clearer, and more personalized

OpenAI updated ChatGPT’s default model to GPT-5.5 Instant for default chat use. The RSS snippet says answers are more accurate, hallucinations are reduced, and personalization controls improved; the post does not disclose metrics, pricing, or context window.

Why it matters: HKR-H/K/R all pass: OpenAI changed ChatGPT’s default model to GPT-5.5 Instant. The post lacks evals, pricing, and context window details, so it stays at the low end of the 85–94 band.

r/LocalLLaMA

vibevoice.cpp: Microsoft VibeVoice ported to ggml/C++ with no Python at inference

LocalAI released vibevoice.cpp, a ggml/C++ port of Microsoft VibeVoice for CPU, CUDA, Metal, and Vulkan inference. TTS uses a 30s reference clip for 24kHz cloned speech; ASR uses a 7B model with diarized JSON and was tested on 17min audio. The key constraint is memory: 17min CPU Q8_0 peaks near 26GB, with no streaming output yet.

Why it matters: HKR-H/K/R all pass: a practical open-source VibeVoice C++ port with concrete runtime numbers. Reddit-source scope and niche audio deployment keep it in the 72–77 featured band, not same-day must-write.

QbitAI · WeChat

Doubao Tests Paid Subscriptions, With Top Tier at 500 Yuan per Month

Doubao listed three App Store subscription tiers at 68, 200, and 500 yuan per month, while keeping a free basic version. QbitAI says the paywall is not live, and ByteDance has only confirmed full details will come through official channels. Doubao app DAU passed 140 million in April, and model calls exceeded 120 trillion tokens per day by March 2026.

Why it matters: HKR-H/K/R all pass: the pricing leak is concrete and high-signal for China AI monetization. It stays below P1 because paid access is not live and model quotas or tier benefits are not disclosed.

r/LocalLLaMA

MTPLX: 2.24x Faster TPS Native MTP Inference Engine for Apple Silicon

MTPLX raises Qwen3.6-27B on a MacBook Pro M5 Max from 28 to 63 tok/s. The test used 4-bit MLX, temperature 0.6, top_p 0.95, top_k 20, with D3 as the best depth. The key detail is native MTP heads: no external drafter and no second-model memory.

Why it matters: HKR-H/K/R all pass: a 2.24x speed hook, concrete test conditions, and a local-inference cost nerve. Reddit single-post sourcing and narrow Apple Silicon scope keep it in low featured, not P1.

OpenAI News

New Ways to Buy ChatGPT Ads

OpenAI expanded ChatGPT ad buying with a beta self-serve Ads Manager, CPC bidding, and enhanced measurement tools. The post says ads protect privacy and keep chats separate; it does not disclose pricing, rollout scope, or timing.

Why it matters: HKR-H/K/R all pass: OpenAI is turning ChatGPT ads into buyable tooling. Price, placement scope, and rollout timing are not disclosed, so this stays a mid-weight business product update.

r/LocalLLaMA

FastDMS: 6.4X KV-cache compression running faster than vLLM BF16/FP8

FastDMS released an MIT implementation that cuts KV memory to 1/5–1/8 of vLLM BF16 at 8K context. A Llama-3.2-1B replication reports PPL 9.200 with 6.4x compression; Qwen3-8B c=1 drops KV from 1.406 GiB to 0.184 GiB. The key detail is physical reclamation of evicted slots, not just nominal KV-byte reduction.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, with compression, PPL, KV GiB deltas, and physical slot reclamation. Reddit/open-source sourcing keeps it in 78–84, below P1.

May 4Monday

r/LocalLLaMA

Deep research report with Hermes Agent and qwen3.6-35b-a3b Q6_K

A Reddit user used Hermes Agent and qwen3.6-35b-a3b Q6_K to produce a 21-page research report. The run took 6 loops and over 5 hours on an RTX 4060, at about 28 tokens/s. The repo includes prompts, scripts, intermediate artifacts, and the final report.

Why it matters: HKR-H/K/R all pass: this is a local-agent experiment with hardware, runtime, speed, and artifacts. Reddit source limits reach, so it stays in the 72–77 featured-threshold band.

QbitAI · WeChat

DeepSeek-TUI, a “DeepSeek Claude Code,” reaches 2.3k GitHub stars

DeepSeek-TUI reached 2.3k GitHub stars; the Rust project is MIT-licensed. It targets DeepSeek V4 with a 1M-token context, RLM up to 16 V4 Flash subtasks, MCP, Shell, Git, and three control modes. Watch cache misses: uncached tokens cost 10x cached tokens.

Why it matters: HKR-H/K/R all pass: the hook is a DeepSeek-flavored Claude Code, with 2.3k stars, 1M tokens, 16 subtasks, and a 10x cache-miss cost gap. Impact is developer-specific, so it sits in the 72–77 band.

r/LocalLLaMA

Gemma 4 E2B runs well on an 8GB Android phone, powering a private voice notes app

A Reddit user ran Gemma 4 E2B locally on an 8GB OnePlus CE 5 and built a private voice notes app. Whisper Small 244MB transcribes, Gemma 4 E2B 2.4GB splits and tags, and a 10-15s note takes 12-15s end to end. Search uses query expansion, FTS lanes, RRF, and optional Gemma top-K reranking with a 15s fallback.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person build, not an official Google release. Concrete hardware, latency, model size, and retrieval details place it near the top of the tutorial band.