Skip to content

#Agent

41 today

Jun 3Wednesday

TechCrunch · AI

OpenAI launches new Codex tools for white-collar work

OpenAI released six Codex app plug-ins for data analytics, creative production, sales, product design, equity investing, and investment banking; each tool bundles integrations, instructions, and context, while the post does not disclose pricing or rollout limits.

Why it matters: HKR-H/K/R all pass: OpenAI is expanding Codex into six white-collar plugin categories. Pricing, rollout scope, and measured performance are not disclosed, so this stays in the mid-weight product-update band.

Jun 2Tuesday

AI HOT (Curated Pool)

Holo3.1: Fast Local Computer-Use Agents

Holo3.1 releases Qwen-based computer-use agents in 0.8B, 4B, 9B, and 35B-A3B sizes, with FP8, Q4 GGUF, and NVFP4 quantized checkpoints for local inference and a 79.3% AndroidWorld score for the 35B-A3B model.

Why it matters: HKR-H/K/R all pass: Holo3.1 pairs a local computer-use agent with concrete model sizes and quantized checkpoints. It fits the 78–84 band, below major lab model-release weight.

Ben's Bites

Opus 4.8

Ben’s Bites says Claude Opus 4.8 is out, and Claude Code can write an orchestration script before launching subagents in parallel to work through complex tasks.

Why it matters: HKR-H/K/R all pass for a substantive Anthropic/Claude release and Claude Code agent update. The post is thin on benchmarks, pricing, and context window, so it stays low in the 85–94 band.

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash as an open-weight model for agentic coding

StepFun released the open-weight Step 3.7 Flash model for fast agentic coding, with tool calling and multimodal understanding, and the model is already available in Kilo alongside MiniMax M3.

Why it matters: HKR-H/K/R pass on the open-weight agentic-coding angle and Kilo availability. Missing benchmarks, size, license, and pricing keep it at the lower featured threshold.

The Verge · AI

Gemini Spark is the most impressive and terrifying AI experience I’ve had yet

The Verge tested Google’s always-on AI agent Gemini Spark for trip planning, but the RSS snippet only describes a different experience from generic itinerary demos and does not disclose launch timing, pricing, benchmarks, or reproducible test conditions.

Why it matters: HKR-H and HKR-R pass: The Verge’s hands-on has a strong click hook and hits agent safety/competition nerves. HKR-K fails because pricing, release timing, and reproducible test conditions are missing, keeping it at the featured threshold.

r/LocalLLaMA

Replaced Claude with local Qwen3.6-27B in my multi-agent orchestrator for 2 weeks

The author ran Qwen3.6-27B on one RTX 3090 across 47 multi-step coding workflows. Plan generation reached about 95% schema validity, but tool-call formatting errors were about 12%, and practical long-context use degraded past about 12k tokens.

Why it matters: HKR-H/K/R all pass: a named first-person local-vs-Claude experiment with concrete numbers. The single Reddit source and 47-workflow scope keep it below the 78–84 band.

Synced · WeChat

DataMaster: When AI Becomes Its Own Data Engineer

DataMaster searches, cleans, and combines data while keeping the model and training algorithm fixed; on MLE-Bench Lite, it raised the medal rate from 35.91% to 68.18%.

Why it matters: HKR-H/K/R all pass: DataMaster changes the data pipeline under fixed model and training code, lifting MLE-Bench Lite medal rate from 35.91% to 68.18%. This is still a single research release without production validation, so it lands at 78 featured.

Synced · WeChat

Turing Award Winner Sutton’s New Paper Argues AI Should Move Toward Enactive Cognition

Banafsheh Rafiee and Richard S. Sutton propose an enactive cognition framework for AI, naming four pillars: experience, perception-action inseparability, autonomy, and embodiment.

Why it matters: HKR-H/K/R all pass, but the article centers on a conceptual framework and does not disclose experiments, code, or reproducible tests. Sutton’s name and the four pillars put it in the 78–84 research-commentary band.

AI HOT (Curated Pool)

To Avoid Paying $120, I Turned a Computer Cleaner into an Open-Source Skill

The author open-sourced a cross-platform AI cleaning skill for Mac and Windows, generating interactive HTML reports from file scans; in a test, it freed nearly 120GB, compared with CleanMyMac identifying 15.8GB.

Why it matters: This is not a platform-level release, but HKR-H/K/R all land through the $120 replacement hook, concrete scan/report mechanism, and 120GB test result. It fits the practical open-source tool band near the featured threshold.

QbitAI · WeChat

Jensen Huang Brings NVIDIA CPUs Into the PC Market

NVIDIA RTX Spark will ship in Windows PCs this fall with 1 petaflop of AI compute and 128GB unified memory. The platform combines a Blackwell RTX GPU, a 20-core Arm-based Grace CPU, and NVLink-C2C, and NVIDIA says it can run 1-million-token-context, 120B-parameter language models locally.

Why it matters: HKR-H/K/R all pass: NVIDIA is moving RTX Spark into Windows PCs with concrete specs: 1 petaflop, 128GB unified memory, 1M context, and 120B local models. This is a strong hardware product update, not a foundation-model release, so it lands in 78–84.

Xinzhiyuan · WeChat

CAS Opens MobileGym, a Browser-Based Agent Training Environment for Mobile Apps

CASIA released MobileGym, a browser-based Android simulation environment covering 28 apps, with about 400MB per instance, 3-second cold start, JSON state snapshots, and programmatic task verification for mobile-agent training and evaluation.

Why it matters: MobileGym is practical open-source infrastructure for agent training and evaluation, with enough concrete numbers and mechanisms to pass HKR-H/K/R. It fits the 78–84 quality band, below major lab model-release weight.

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash for efficient inference

StepFun released Step 3.7 Flash with a 196B MoE architecture, using multi-matrix factorized attention to cut KV-cache cost to about 22% of DeepSeek models.

Why it matters: HKR-H/K/R all pass: Step 3.7 Flash has concrete specs, not just launch copy, with 196B MoE and ~22% KV-cache cost versus DeepSeek. It is below top-lab flagship weight, so 78 featured.

AI HOT (Curated Pool)

Anthropic Developer Shares a Claude Code Understanding-Verification Workflow

An Anthropic developer shared a Claude Code understanding-verification workflow with 8 steps, using incremental teaching, user restatement, checklists, and quizzes to confirm the human can defend the problem, solution, and impact before moving to the next stage.

Why it matters: HKR-H/K/R all pass: a concrete Claude Code workflow with an 8-step verification loop and a strong oversight hook. It is a practical tutorial, not a product release, so it sits at the lower featured band.

Computing Life · Share · Yage

AI agents don't need to be hacked; persuasion is enough

The article says an AI agent with password-reset permission can be abused when an attacker persuades it they are a legitimate user; the snippet only discloses a three-layer architecture that separates what from who, not concrete attack steps or implementation details.

Why it matters: HKR-H/K/R all pass: the hook is strong, the post offers a what/who three-layer design, and agent permissions are a live security worry. No real incident, success rate, or product comparison keeps it at the featured threshold.

Computing Life · Share · Yage

The Next Form of AI Agents: From Chat Windows to Background Daemons

Gemini Spark is described as the first consumer-facing always-on background agent from a major platform; the post covers four product generations and a periodic versus reactive automation framework.

Why it matters: HKR-H/K/R all pass, but this is a single commentary item; the body summary does not disclose launch date, rollout scope, or hands-on results for Gemini Spark. Score stays at the lower featured band.

r/LocalLLaMA

I spent months inside verl, forked it, then stopped: internals, fork costs, and an NCCL bug

ReinforcedKnowledge analyzes ByteDance’s verl RLHF loop, covering DataProto plus rollout, reward, advantage, and update paths. The author stopped a private fork because near-daily upstream changes made sync cost exceed refactoring work, and describes an NCCL hang fixed on one node by setting NCCL_SOCKET_IFNAME=lo.

Why it matters: Niche but useful RL post-training field report, not an industry release. HKR-H comes from the fork-then-quit twist; HKR-K has verl’s five paths and NCCL_SOCKET_IFNAME=lo; HKR-R hits the cost of maintaining open-source training forks.

AI HOT (Curated Pool)

Google AI Studio adds app-building support for Gmail and other apps

Google AI Studio has added app-building support for connected Gmail, Drive, and Sheets apps, and users can add testers inside AI Studio; the post does not disclose a launch date for full public sharing.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight product update: Workspace connections and tester support are confirmed, while sharing, permission details, and pricing are not disclosed.

TechCrunch · AI

Nvidia chases $200B CPU market with AI agent PCs from Microsoft, Dell, and HP

The title says Nvidia is targeting the $200B CPU market with AI agent PCs from Microsoft, Dell, and HP; the RSS snippet does not disclose specifications, pricing, launch timing, or the safety mechanism for bringing agents to consumer PCs.

Why it matters: HKR-H/K/R all pass, but specs, price, and launch timing are not disclosed. Treat it as a mid-weight product/ecosystem update, with Nvidia plus Microsoft/Dell/HP enough for low featured.

The Verge · AI

Gemini’s New AI Agent Is About as Good as Google’s Demo

The Verge tested Google Gemini Spark for one week and says the 24/7 agent can run multi-step tasks in the background, but the RSS snippet does not disclose pricing, privacy terms, or the full hands-on results.

Why it matters: HKR-H/K/R pass: a Verge hands-on stress-tests Google’s Gemini Spark demo claim and confirms background multi-step tasks. Missing price, privacy terms, and full results keep it in the 72–77 band.

AI HOT (Curated Pool)

Meta AI Exploit Used to Hijack Instagram Accounts

Meta’s AI chatbot was found vulnerable to an account-takeover exploit against Instagram accounts. Attackers could ask the AI to link a new email address, and the failure condition was the agent’s ability to execute account-management actions directly; the RSS snippet does not disclose affected account counts, patch status, or reproduction details.

Why it matters: HKR-H/K/R all pass: a Meta AI support agent allegedly enabled Instagram account takeover via add-email requests. Impact scale, fix timeline, and reproducible steps are not disclosed, so it stays in the 78–84 band.

Hacker News front page

Hackers Used Meta's AI Support Bot to Seize Instagram Accounts

The title says hackers used Meta's AI support bot to seize Instagram accounts; the RSS snippet lists 40 points and 14 comments, but the post does not disclose the attack mechanism.

Why it matters: HKR-H and HKR-R pass: a Meta AI support bot allegedly enabled Instagram account takeovers, a Krebs-sourced security angle. HKR-K fails because the feed lacks mechanism or scale, so it sits at the featured floor.

AI HOT (Curated Pool)

Perplexity Releases Search as Code Architecture

Perplexity released Search as Code, an architecture where agents write Python code to call its search stack directly instead of looping through function calls; it is now available in the Perplexity Agent API and is the default option for Computer.

Why it matters: HKR-H/K/R pass: Perplexity gives a concrete agent-search mechanism and Agent API integration. Single-source post lacks performance, pricing, and rollout scope, so this stays a low featured product update.

Jun 1Monday

Latent Space

Why Video Agent Models Are Next — Ethan He on xAI Grok Imagine

Ethan He says a small xAI team built Grok Imagine from zero to one in 3 months, and the episode discusses video agents, audio-video alignment, inference speedups, and the storage, egress, and GPU-hour costs behind large video datasets.

Why it matters: HKR-H/K/R all pass, but the body is interview-level signal: beyond the 3-month build and mechanism themes, it gives no benchmarks, cost figures, or reproducible test. Strong xAI video-agent context, not same-day must-write.

AI HOT (Curated Pool)

Open and Closed Models Are on Different Exponentials

Nathan Lambert argues that closed frontier labs will capture high-margin demand in coding-agent workflows, citing a personal willingness to pay $2,000 per month and projecting OpenAI and Anthropic valuations of $2-10 trillion over 5-10 years.

Why it matters: HKR-H/K/R all pass: the essay has a clear open-vs-closed hook, concrete price and valuation claims, and practitioner resonance. It remains single-source commentary, so it sits in the featured-threshold band.

AI HOT (Curated Pool)

Wang Xing: Meituan AI Agent Xiao Mei to Partner Deeply with Tencent Yuanbao

Meituan CEO Wang Xing said Xiao Mei will connect with Tencent Yuanbao, routing local service requests into food ordering, delivery, and related Meituan scenarios; Meituan reported Q1 2026 revenue of RMB 91.039 billion and a net loss of RMB 6.827 billion.

Why it matters: HKR-H/K/R all pass, but the deal is still “coming soon”; launch timing, UX entry point, and revenue split are not disclosed. This fits a mid-weight product partnership at the featured floor.

AI HOT (Curated Pool)

Tutorial: Turning Books into AI Skills with Claude Opus 4.8

The author used Claude Opus 4.8 to turn Nonviolent Communication into an AI Skill through a six-step workflow, taking about 45 minutes, using roughly 300,000 tokens, and costing under RMB 20.

Why it matters: HKR-H/K/R all pass: this is a numbered first-person Claude workflow with concrete cost and token details. It stays in the lower featured band because it is a personal tutorial, not an Anthropic release or model update.

AI HOT (Curated Pool)

Apache RocketMQ Releases an AI-Focused Messaging Engine

Apache RocketMQ released RocketMQ for AI, a messaging engine for long-running sessions, multi-agent workflows, and fair scheduling, with Lite-Topics, ordered messages, and traffic shaping; the post does not disclose a version number or performance figures.

Why it matters: HKR-H/K/R pass: the AI-specific RocketMQ angle has a real agent-infra hook and named mechanisms. Score stays in the 72–77 band because version, benchmarks, and production cases are not disclosed.

Xinzhiyuan · WeChat

400 tokens/s: StepFun Step 3.7 Flash cuts Agent task costs

StepFun released Step 3.7 Flash, a sparse MoE model with 196B parameters plus a 1.8B ViT, activating 11B parameters per inference and reaching up to 400 tokens per second.

Why it matters: HKR-H/K/R all pass with concrete speed and parameter numbers. The feed does not disclose pricing, benchmark setup, or open-source terms, so this stays in the 78–84 quality update band.

Synced · WeChat

OpenAI recruits for robotics team led by Sora creator Aditya Ramesh

OpenAI has listed more than a dozen San Francisco robotics roles for OpenAI Robotics, a team that evolved from Aditya Ramesh’s Worldsim work, with the actuator design engineer role offering $342,000 to $445,000 in base cash pay plus PPU incentives.

Why it matters: HKR-H/K/R all pass: OpenAI robotics hiring adds a strong hook, plus concrete roles, leader, and salary range. This is still a hiring signal, not a model or product release, so it stays in the featured-threshold band.

Synced · WeChat

World models get a “save state”: VAST releases Project Eden

VAST released Project Eden, a three-layer world-model architecture that separates persistent state evolution from visual rendering, and disclosed nearly $200 million across its A+ and A++ funding rounds.

Why it matters: HKR-H/K/R all pass: Project Eden has a product hook, architecture detail, and funding scale. VAST is not a top foundation-model lab, and benchmarks or access terms are not disclosed, so this lands in 78–84.

AI HOT (Curated Pool)

Tencent Hunyuan Releases Long-Term Memory Plugin Hy-Memory

Tencent Hunyuan released Hy-Memory for long-term collaborative agents such as OpenClaw, using a six-layer memory framework and System1/System2 dual system, with memory count reduced by over 70% and token consumption down 35% in ultra-long-context scenarios.

Why it matters: Tencent Hunyuan’s Hy-Memory clears HKR-H/K/R with a concrete memory architecture and cost-reduction figures. The score stays at the featured floor because the source is an official short post without reproducible tests, license details, or third-party benchmarks.

QbitAI · WeChat

VAST Raises Nearly $200M and Discloses Its Project Eden World Model Roadmap

VAST raised nearly $200 million in A+ and A++ rounds and disclosed Project Eden, a world model architecture that separates state evolution from visual rendering through a structured state layer, a conditional interface layer, and a generative rendering layer.

Why it matters: HKR-H/K/R all pass: the $200M A+/A++ financing is sizable, and Project Eden gives a concrete three-layer world-model mechanism. VAST is not a top-tier foundation-model lab and no metrics or release details are disclosed, so this stays in the 78–84 band.

AI HOT (Curated Pool)

NVIDIA Releases FOX Factory Operations Blueprint for Autonomous Factory Management Agents

NVIDIA released the FOX factory operations blueprint at GTC Taipei, and Foxconn used it to build the MoMClaw multi-agent system with an expected 80% reduction in root-cause analysis time.

Why it matters: HKR-H/K/R pass: NVIDIA is pushing an agent blueprint into factory ops, with Foxconn’s MoMClaw and an expected 80% RCA time cut. Kept at the featured floor because the source is a vendor blog and the result is projected.

AI HOT (Curated Pool)

NVIDIA Releases RTX Spark and Local AI Agent Security and Performance Updates

NVIDIA released RTX Spark, a Windows PC for local AI agents with 1 petaflops of AI compute and 128GB of unified memory. OpenShell uses new Windows security primitives with Microsoft, while llama.cpp optimizations raise Qwen 27B throughput by up to 2x.

Why it matters: HKR-H/K/R all pass: NVIDIA frames RTX Spark for local agents and gives hard specs: 1 petaflops, 128GB, and up to 2x llama.cpp throughput. Vendor-blog framing keeps it in the low 78–84 band.

AI HOT (Curated Pool)

MiniMax M3: Frontier coding, 1M-token context, and native multimodal model

MiniMax released M3 as an open-source unified model with coding, agent, and native multimodal capabilities, supporting a 1M-token context window and using MiniMax Sparse Attention to cut per-token compute at 1M context to 1/20 of its predecessor, with over 9x faster prefill and over 15x faster decoding.

Why it matters: HKR-H/K/R all pass: MiniMax M3 has a 1M-token context hook, MSA with a claimed 20x cost cut, and open-source China-model resonance. Single official-source release keeps it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Qwen3.7-Plus: Multimodal Agent Intelligence

Qwen Studio lists seven capability areas: chatbots, image and video understanding, image generation, document processing, web search integration, tool use, and artifact generation; the post does not disclose Qwen3.7-Plus parameters, pricing, or release timing.

Why it matters: HKR-H/K/R pass, but the facts are thin: 7 capability categories, no params, pricing, benchmarks, or launch terms. A Qwen flagship update clears featured, not p1.

AI HOT (Curated Pool)

MWC26 Shanghai to Host First Humanoid Robot Penalty Shootout With Unitree and 7 Other Teams

MWC26 Shanghai will host a humanoid robot penalty shootout in June 2026, with eight Chinese embodied intelligence teams competing under rules that require autonomous play without human control or preset scripts.

Why it matters: HKR-H/K/R all pass: the robot penalty shootout is clickable, with rules banning teleoperation and scripts. It stays in 72–77 because this is an event preview, not a model release or reproducible result.

May 31Sunday

AI HOT (Curated Pool)

Apple WWDC AI Upgrade: Gemini-Distilled Model Runs Locally, With Heavy External Dependencies

Apple will present Siri and on-device AI upgrades at next month’s WWDC, with iPhones running a smaller Gemini-distilled model locally while complex queries route to Google Cloud using Nvidia confidential computing.

Why it matters: HKR-H/K/R all pass: the Apple-Google-Nvidia stack is a strong WWDC AI hook with a concrete routing mechanism and clear industry tension. Capped at 82 because this is a single X-sourced claim with no model size, latency, pricing, or contract terms disclosed.

r/LocalLLaMA

PolyRange: Contamination-resistant offensive-AI benchmark for web targets

PolyRange v1.0 ships 84 WSTG-derived classes across 12 OWASP testing-guide categories. It generates fresh targets per deploy with a chosen LLM, adds two defense tiers, uses an agent-submits-flag oracle, and runs via a single-command CLI on Fly.io or Docker.

Why it matters: HKR-H/K/R all pass: PolyRange turns web-security targets into a dynamic agent benchmark with 84 WSTG classes and two defense levels. Single-source Reddit origin and security niche keep it at 78.

r/LocalLLaMA

Use any model and provider with the official OpenAI Codex Desktop App without modifying its code

Reddit user thibautrey describes a 3-step setup: edit Codex Desktop config.toml, store an API key, and use a multicodex proxy alias to map gpt-5.3-codex to MiniMax-Latest. The post lists a local base_url of 127.0.0.1:1455 and says the proxy disguises returned model names as gpt-5.3-codex.

Why it matters: This is a reproducible developer workflow trick, not an official release. HKR-H comes from the lock-in workaround, HKR-K has concrete config details, and HKR-R hits cost and model-choice pressure, placing it at the tutorial featured threshold.