Skip to content

#编码

10 today

Jun 4Thursday

QbitAI · WeChat

Beyond TurboQuant: Together AI Brings 2-bit KV Cache to Real Serving

Together AI, the University of Sydney, and UIUC introduced OSCAR, a 2-bit KV Cache quantization method that uses about 2.28 effective bits per KV element and scores 71.86 on Qwen3-4B-Thinking, 40.1 points above TurboQuant.

Why it matters: HKR-H/K/R all pass: OSCAR links 2-bit KV cache to serving and provides concrete scores. The topic is still low-level inference optimization, so it lands in featured rather than same-day must-write.

AI HOT (Curated Pool)

Hugging Face redesigns hf CLI output format for coding agents

Hugging Face redesigned hf CLI output for coding agents including Claude Code and Codex, using environment-variable detection and compact untruncated TSV output; in complex multi-step tasks, agents without the CLI used up to 6 times more tokens.

Why it matters: HKR-H/K/R pass: the story has a clear agent-CLI hook, a concrete TSV/token mechanism, and strong developer cost resonance. It stays in the featured band because this is a tooling update, not a model or platform release.

Latent Space

Scaling Past Informal AI - Carina Hong, Axiom Math

Axiom solved all 12 Putnam problems in 2025 and scored 8/12 within the time limit; Carina Hong says its Verina ProofGen result reached 187/189, while the last disclosed OpenAI o3 result on that benchmark was 4.9%.

Why it matters: HKR-H/K/R all pass: Putnam results, the o3 comparison, and 187/189 give it a real hook. It stays at 80 because this is a Latent Space interview/research story, not a broad model release.

r/LocalLLaMA

I built a compiler that rewrites Python into a model-facing representation

The author released Vulpine, a compiler that converts Python into a compact model-facing representation for coding LLMs. Tests on about 13,000 held-out files showed roughly 14% token reduction and 99.8% AST-equivalent round-trip success, with code published on GitHub.

Why it matters: HKR-H/K/R all pass, with a named experiment and concrete numbers. Source authority is low and the post does not disclose real-task gains, speed, or failure cases, so it stays at the featured threshold.

Jun 3Wednesday

r/LocalLLaMA

google/gemma-4-12B on Hugging Face

Google DeepMind released Gemma 4 open-weight models in five sizes, with the 12B variant supporting text, image, and audio input, instruction-tuned and pre-trained variants, native system prompts, function calling, and a context window of up to 256K tokens.

Why it matters: Gemma 4 clears HKR-H/K/R: open weights, multimodal input, and 256K context make it more than a routine update. Missing benchmarks, license detail, and fuller official context keep it in the 78–84 band.

Alibaba Technology · WeChat

Rethinking R&D Infrastructure When Agents Become First-Class Citizens

Xu Xiaobin argues that agent-based development compresses the intent-to-code loop from weeks or months to minutes, using a weekly-report system, a multi-role agent development setup, and image-repository provisioning as examples; the article identifies mismatches in Git, CI, code review, release flows, permissions, harness setup, and dry-run validation.

Why it matters: HKR-H/K/R all pass, but this is infrastructure commentary rather than a model or product launch. The named cases and week/month-to-minutes claim put it in the 72–77 featured band.

Latent Space

[AINews] Microsoft Build: MAI-Thinking-1 and MAI Family Models

Microsoft announced seven MAI models at Build, with MAI-Thinking-1 described as a 35B-active-parameter MoE with a 256K context window, and released a 109-page technical report covering training, data lineage, and performance claims.

Why it matters: All HKR axes pass: Microsoft’s MAI family has concrete specs, a long technical report, and clear competitive stakes around its model stack. This clears the 85+ same-day bar, but no weights, pricing, or external evals are disclosed, so it lands at 87.

QbitAI · WeChat

Coze 3.0 test: phone can remotely control agents on your computer

Coze 3.0 adds project-based agent collaboration across iOS, Android, Mac, Windows, and web, supports importing local agents such as Claude Code, Codex CLI, and OpenClaw, and can read a desktop PDF from a phone after user authorization.

Why it matters: HKR-H/K/R all pass: Coze 3.0 adds cross-device agent control and imports Claude Code/Codex CLI. It remains a single product update, with price, rollout scope, and security limits not disclosed, so it sits at the featured threshold.

Xinzhiyuan · WeChat

OpenAI’s Greg Brockman and a 9-Year Rift With Anthropic Co-Founder Dario Amodei

A WSJ-based profile says Dario Amodei once barred Greg Brockman from an internal OpenAI project that later led to ChatGPT, and the article says Brockman now oversees OpenAI product strategy with nearly 1,500 people under that function.

Why it matters: HKR-H/K/R all pass: the WSJ-sourced ban detail and the nearly 1,500-person scope give this more signal than gossip. It is not a model release or current executive departure, so it stays in the good-quality featured band.

AI HOT (Curated Pool)

Complete Practical Tips for Agent Engineering

@mvanhorn shared an agent engineering workflow centered on a Research→Plan→Work loop, plan.md constraints, and 22 practical tips; the snippet says it covers planning, parallel execution, input methods, and remote control, but the post does not disclose the full tool stack list.

Why it matters: HKR-H/K/R all pass, but this is a practitioner methods post, not a model or product release. The full tool stack is not disclosed, so it sits at the featured threshold.

Computing Life · Share · Yage

After vibe coding: the industrialization of AI programming

MAI filtered 265,000 trainable tasks from 4.87 million open-source PRs and built a three-layer judging system. The key change after vibe coding is the industrialization of training infrastructure.

Why it matters: HKR-H/K/R all pass via the post-vibe-coding angle, 4.87M PR corpus, 265K tasks, and code-agent infra stakes. No model scores, open-source scope, or product access are disclosed, so it stays below P1.

AI HOT (Curated Pool)

Intelligence Cost-Performance

Microsoft added average token usage to its model release card; the model scored 71.6 on SWE-Bench Verified while using about one-third of Claude Haiku 4.5’s tokens.

Why it matters: HKR-H/K/R all pass: the score-per-token angle is clickable, with concrete 71.6 and one-third-token claims. The article is thin on full test setup and pricing, so it lands at 78.

AI HOT (Curated Pool)

OpenAI launches Codex Sites to turn ideas into interactive websites

OpenAI opened Codex Sites in preview to Business and Enterprise subscribers, letting users turn ideas into hosted interactive sites such as dashboards, planners, and project boards, with URL sharing for specified team members.

Why it matters: HKR-H/K/R all pass, but this is an OpenAI Codex enterprise-preview feature rather than a model or core capability release. It sits in the mid-weight product-update band.

AI HOT (Curated Pool)

Claude Code Adds Dynamic Workflows

Claude Code added dynamic workflows that execute JavaScript files at runtime to create and coordinate multiple subagents; each subagent has its own context window, and the feature is described for research, security analysis, and code review tasks.

Why it matters: HKR-H/K/R all pass: Claude Code gets runtime JS workflows coordinating isolated-context subagents. Anthropic update earns a bump, but this is a feature release rather than a model or platform launch, so it sits in the 78–84 band.

AI HOT (Curated Pool)

Microsoft releases MAI-Thinking-1 model

Microsoft released MAI-Thinking-1, an MoE model with 35B active parameters and 1T total parameters, pretrained from scratch on 30T tokens without third-party model distillation.

Why it matters: HKR-H/K/R all pass: Microsoft released MAI-Thinking-1 with concrete MoE scale and training-token figures. Benchmarks, access, and pricing are not disclosed, so it stays in the 78–84 band rather than P1.

AI HOT (Curated Pool)

Claude Code launches dynamic workflows for task-specific frameworks

Claude Code added dynamic workflows that execute JavaScript files to coordinate subagents, with configurable model choice and workspace isolation level, but the post does not disclose token overhead figures or release availability details.

Why it matters: HKR-H/K/R all pass, but the post gives mechanism-level detail only; token overhead, rollout scope, and pricing are not disclosed. Claude Code relevance lifts this to the high end of a mid-weight product update.

Hacker News front page

Microsoft's MAI-Code-1-Flash Scores 51% SWE-Bench Pro with Just 5B Active Params

The title says Microsoft's MAI-Code-1-Flash scores 51% on SWE-Bench Pro with 5B active parameters; the post does not disclose the evaluation setup, training data, release date, or deployment conditions.

Why it matters: HKR-H/K/R pass on the 51% SWE-Bench Pro with 5B active params claim from Microsoft. Missing eval setup, training data, and release timing keep it in the 72–77 band.

AI HOT (Curated Pool)

Claude Platform Adds CLI Tool

Claude Platform added a CLI that runs every API endpoint from the terminal, calls the Messages API, launches Claude-hosted agents, and pipes results directly into the shell.

Why it matters: Claude Platform CLI clears HKR-H/K/R as a practical developer-tooling update, but the post only gives capability scope; install flow, permissions, safety limits, and pricing are not disclosed.

AI HOT (Curated Pool)

Microsoft releases its first advanced reasoning AI model, MAI-Thinking-1

Microsoft released MAI-Thinking-1 at Build 2026, describing it as a medium-sized reasoning model that matches leading models on key software engineering benchmarks.

Why it matters: HKR-H/K/R all pass: Microsoft released its first advanced reasoning model with a mid-sized design and SWE benchmark claim. Exact scores, access, and pricing are not disclosed, so it stays below 85.

Bloomberg Technology

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs

Uber Technologies set usage caps on staff AI tools including Claude Code after the company exceeded its AI budget earlier this year; the post does not disclose the cap size, affected teams, or budget amount.

Why it matters: HKR-H/K/R all pass: the Bloomberg item gives a named enterprise cost-control case for Claude Code-like tools. Budget size, cap rules, and affected headcount are not disclosed, keeping it at the featured threshold.

Latent Space

GitHub's Plan for Agents — Kyle Daigle, GitHub

GitHub COO Kyle Daigle said AI-driven code commits grew 14x in 2026, and the interview covers Copilot, Actions, MCP, WorkIQ, cloud agents, and the infrastructure availability pressure created when code review, CI/CD, and open-source contribution volume scale beyond human-speed workflows.

Why it matters: HKR-H/K/R all pass: a GitHub executive gives a 14x AI code-submission figure and ties Copilot, Actions, MCP, WorkIQ, and cloud agents into one roadmap. Not a major release, so it stays at 80.

AI HOT (Curated Pool)

Claude Code Team Practice: How Agentic Coding Changes Engineering Organizations and Processes

The Claude Code engineering team described process changes after making agentic coding the default at Code w/ Claude SF 2026: JIT planning, asking Claude first for context collection, Claude handling style and tests in code review, and humans focusing on legal and safety judgments.

Why it matters: First-party Claude Code workflow post with concrete engineering mechanisms and strong HKR-H/K/R fit. It is not a model or major product release, so it stays in the 78–84 band.

AI HOT (Curated Pool)

OpenAI Codex releases Python SDK for direct app integration

OpenAI Codex released a Python SDK with the install command pip install openai-codex, and the snippet says it can reuse the Codex login state; the post does not disclose API pricing, model versions, or rate-limit conditions.

Why it matters: HKR-H/K/R pass: a Codex SDK for embedded app use is practical and discussable. Sparse sourcing keeps it in the mid-weight product-update band: package and auth are given, but price, model, and rate limits are not.

AI HOT (Curated Pool)

OpenAI Codex Sites feature launches

OpenAI launched Codex Sites, which turns work, ideas, and plans into an interactive website or app that a team can access through one URL; the feature rolls out first to Business and Enterprise plans, and the post does not disclose pricing or broader availability timing.

Why it matters: HKR-H/K/R all pass, but the post gives launch framing without pricing, permission boundaries, or quality examples. Treat it as a mid-weight OpenAI product feature, above the featured threshold.

TechCrunch · AI

OpenAI launches new Codex tools for white-collar work

OpenAI released six Codex app plug-ins for data analytics, creative production, sales, product design, equity investing, and investment banking; each tool bundles integrations, instructions, and context, while the post does not disclose pricing or rollout limits.

Why it matters: HKR-H/K/R all pass: OpenAI is expanding Codex into six white-collar plugin categories. Pricing, rollout scope, and measured performance are not disclosed, so this stays in the mid-weight product-update band.

Jun 2Tuesday

Ben's Bites

Opus 4.8

Ben’s Bites says Claude Opus 4.8 is out, and Claude Code can write an orchestration script before launching subagents in parallel to work through complex tasks.

Why it matters: HKR-H/K/R all pass for a substantive Anthropic/Claude release and Claude Code agent update. The post is thin on benchmarks, pricing, and context window, so it stays low in the 85–94 band.

AI HOT (Curated Pool)

Anthropic Expands Project Glasswing Program

Anthropic expanded Project Glasswing to about 150 new organizations across more than 15 countries, covering electricity, water, healthcare, communications, and hardware infrastructure, after an initial group of about 50 partners.

Why it matters: Anthropic expanded Project Glasswing to about 150 new organizations across 15+ countries, giving HKR-H/K/R enough substance. No concrete safety mechanism or Claude capability change is disclosed, so it stays in the lower featured band.

r/LocalLLaMA

Replaced Claude with local Qwen3.6-27B in my multi-agent orchestrator for 2 weeks

The author ran Qwen3.6-27B on one RTX 3090 across 47 multi-step coding workflows. Plan generation reached about 95% schema validity, but tool-call formatting errors were about 12%, and practical long-context use degraded past about 12k tokens.

Why it matters: HKR-H/K/R all pass: a named first-person local-vs-Claude experiment with concrete numbers. The single Reddit source and 47-workflow scope keep it below the 78–84 band.

AI HOT (Curated Pool)

To Avoid Paying $120, I Turned a Computer Cleaner into an Open-Source Skill

The author open-sourced a cross-platform AI cleaning skill for Mac and Windows, generating interactive HTML reports from file scans; in a test, it freed nearly 120GB, compared with CleanMyMac identifying 15.8GB.

Why it matters: This is not a platform-level release, but HKR-H/K/R all land through the $120 replacement hook, concrete scan/report mechanism, and 120GB test result. It fits the practical open-source tool band near the featured threshold.

AI HOT (Curated Pool)

Anthropic Developer Shares a Claude Code Understanding-Verification Workflow

An Anthropic developer shared a Claude Code understanding-verification workflow with 8 steps, using incremental teaching, user restatement, checklists, and quizzes to confirm the human can defend the problem, solution, and impact before moving to the next stage.

Why it matters: HKR-H/K/R all pass: a concrete Claude Code workflow with an 8-step verification loop and a strong oversight hook. It is a practical tutorial, not a product release, so it sits at the lower featured band.

Hacker News front page

OpenAI frontier models and Codex are now available on AWS

OpenAI made its frontier models and Codex available on AWS; the RSS body only provides the article link, 56 Hacker News points, and 17 comments, and the post does not disclose regions, pricing, or the model list.

Why it matters: HKR-H/K/R all pass because OpenAI-on-AWS changes distribution optics and enterprise options. The post lacks regions, pricing, model list, and access path, so it stays below the 85 band.

AI HOT (Curated Pool)

Perplexity Releases Search as Code Architecture

Perplexity released Search as Code, an architecture where agents write Python code to call its search stack directly instead of looping through function calls; it is now available in the Perplexity Agent API and is the default option for Computer.

Why it matters: HKR-H/K/R pass: Perplexity gives a concrete agent-search mechanism and Agent API integration. Single-source post lacks performance, pricing, and rollout scope, so this stays a low featured product update.

Jun 1Monday

AI HOT (Curated Pool)

MiniMax Releases Open-Source M3 with Coding, Long-Context, and Multimodal Capabilities

MiniMax released the open-source M3 model with coding, a 1M-token context window, and native multimodal support; M3 scores 59.0% on SWE-Bench Pro, 83.5% on BrowseComp, and costs about one-twelfth per token versus GPT-5.5.

Why it matters: HKR-H/K/R all pass: M3 has open source, 1M context, multimodal support, and 59.0% on SWE-Bench Pro. A single X post without official docs or third-party tests keeps it in the 78–84 band.

AI HOT (Curated Pool)

Open and Closed Models Are on Different Exponentials

Nathan Lambert argues that closed frontier labs will capture high-margin demand in coding-agent workflows, citing a personal willingness to pay $2,000 per month and projecting OpenAI and Anthropic valuations of $2-10 trillion over 5-10 years.

Why it matters: HKR-H/K/R all pass: the essay has a clear open-vs-closed hook, concrete price and valuation claims, and practitioner resonance. It remains single-source commentary, so it sits in the featured-threshold band.

AI HOT (Curated Pool)

OpenBMB Releases Two UltraData Open Datasets, Tops HuggingFace Trending

OpenBMB, Tsinghua NLP, and Modelbest released two UltraData open datasets: Ultra-FineWeb-L3 contains 600B+ tokens, including 400B+ English and 200B+ Chinese tokens, while UltraData-SFT-2605 contains 15M+ SFT samples with thinking and non-thinking labels.

Why it matters: HKR-H/K/R pass: two open datasets, 600B+ tokens, and 15M+ SFT samples are concrete practitioner signal. Single-source release with no evals or license detail keeps it at the lower featured band.

AI HOT (Curated Pool)

MiniMax M3: Frontier coding, 1M-token context, and native multimodal model

MiniMax released M3 as an open-source unified model with coding, agent, and native multimodal capabilities, supporting a 1M-token context window and using MiniMax Sparse Attention to cut per-token compute at 1M context to 1/20 of its predecessor, with over 9x faster prefill and over 15x faster decoding.

Why it matters: HKR-H/K/R all pass: MiniMax M3 has a 1M-token context hook, MSA with a claimed 20x cost cut, and open-source China-model resonance. Single official-source release keeps it in the 78–84 band, not P1.

OpenAI News

OpenAI banned a likely PRC-origin cluster using ChatGPT to generate anti-US-data-center social media content

OpenAI's June threat report details a banned cluster of ChatGPT accounts likely originating in China. The operators used Simplified Chinese prompts to generate English posts and images on X, posing as ordinary Americans and claiming data centers and AI are driving up electricity costs for households. They also used ChatGPT for image editing, automation scripts, and harassing overseas dissidents. An internal work report they uploaded outlined tactics for building credible personas on Facebook and evading platform detection.

Why it matters: OpenAI's official threat report details a likely PRC-linked AI influence op with concrete tradecraft and a topic — AI driving up living costs — that's already a public flashpoint. Hits all three HKR axes, but as a security incident report rather than a product or research brea...

May 31Sunday

r/LocalLLaMA

Use any model and provider with the official OpenAI Codex Desktop App without modifying its code

Reddit user thibautrey describes a 3-step setup: edit Codex Desktop config.toml, store an API key, and use a multicodex proxy alias to map gpt-5.3-codex to MiniMax-Latest. The post lists a local base_url of 127.0.0.1:1455 and says the proxy disguises returned model names as gpt-5.3-codex.

Why it matters: This is a reproducible developer workflow trick, not an official release. HKR-H comes from the lock-in workaround, HKR-K has concrete config details, and HKR-R hits cost and model-choice pressure, placing it at the tutorial featured threshold.

AI HOT (Curated Pool)

Run Python ASGI Apps in the Browser with Pyodide and Service Workers

Simon Willison demonstrated running Python ASGI apps in the browser with Pyodide and Service Workers, with Claude Opus 4.8 assisting development, and showed two working demos: a basic ASGI FastCGI demo and Datasette 1.0a31.

Why it matters: HKR-H/K/R all pass: the post has a surprising browser-runtime hook, concrete mechanisms, and developer resonance. Impact stays in the 72–77 band because this is a developer experiment, not a model or platform launch.

AI HOT (Curated Pool)

“What a joke”: GitHub Copilot’s new token-based billing draws developer backlash

GitHub Copilot changed billing to token-based metering, and the RSS snippet says developers are unhappy; the post does not disclose pricing, per-token rates, or the rollout date.

Why it matters: HKR-H/K/R all pass: Copilot’s token billing creates conflict, a concrete mechanism, and a cost nerve for developers. Missing price, unit economics, and start date keep it in the lower featured band.