Skip to content

#编码

0 today

Jun 7Sunday

Synced · WeChat

ICML 2026 | FusionRoute: From Expert Routing to Self-Correction in Multi-LLM Collaboration

FusionRoute proposes a token-level multi-LLM collaboration method that freezes expert models and trains a lightweight router to select an expert for each token while merging router logits with expert logits. The paper evaluates it on GSM8K, MATH-500, HumanEval, MBPP, IfEval, and 500 PerfectBlend prompts.

Why it matters: HKR-H/K/R pass: token-level LLM routing is a strong research hook with concrete mechanics. The article lacks lift numbers, code link, and deployment cost, so it stays at the lower featured band.

Jun 6Saturday

Synced · WeChat

DeepSeek V4 Proves Math with 500x Cost Advantage as Agent System Sets Records

Princeton researchers released Goedel-Architect, an agent framework for Lean formal theorem proving. Using DeepSeek-V4-Flash, it reached 75.6% pass@1 on PutnamBench, with $294 in API cost for 672 problems, compared with Hilbert’s 70.0% and about $170,000 cost.

Why it matters: HKR-H/K/R all pass: Goedel-Architect pairs a 75.6% PutnamBench score with $294 for 672 problems, versus Hilbert at about $170k. It is still research-heavy, so it stays in the 78–84 band rather than P1.

Jun 4Thursday

QbitAI · WeChat

Beyond TurboQuant: Together AI Brings 2-bit KV Cache to Real Serving

Together AI, the University of Sydney, and UIUC introduced OSCAR, a 2-bit KV Cache quantization method that uses about 2.28 effective bits per KV element and scores 71.86 on Qwen3-4B-Thinking, 40.1 points above TurboQuant.

Why it matters: HKR-H/K/R all pass: OSCAR links 2-bit KV cache to serving and provides concrete scores. The topic is still low-level inference optimization, so it lands in featured rather than same-day must-write.

Jun 3Wednesday

Computing Life · Share · Yage

After vibe coding: the industrialization of AI programming

MAI filtered 265,000 trainable tasks from 4.87 million open-source PRs and built a three-layer judging system. The key change after vibe coding is the industrialization of training infrastructure.

Why it matters: HKR-H/K/R all pass via the post-vibe-coding angle, 4.87M PR corpus, 265K tasks, and code-agent infra stakes. No model scores, open-source scope, or product access are disclosed, so it stays below P1.

AI HOT (Curated Pool)

Microsoft releases MAI-Thinking-1 model

Microsoft released MAI-Thinking-1, an MoE model with 35B active parameters and 1T total parameters, pretrained from scratch on 30T tokens without third-party model distillation.

Why it matters: HKR-H/K/R all pass: Microsoft released MAI-Thinking-1 with concrete MoE scale and training-token figures. Benchmarks, access, and pricing are not disclosed, so it stays in the 78–84 band rather than P1.

May 29Friday

Synced · WeChat

A True 2-bit KV Quantization Algorithm for Long-context Reasoning Beyond TurboQuant

TogetherAI and collaborators released OSCAR, a 2.28 BPE INT2 KV Cache system integrated with SGLang, reporting up to 3× decode speedup at 100k context and up to 7× job-level throughput under a fixed memory budget.

Why it matters: HKR-H/K/R pass, but this is niche inference optimization rather than a broad model launch. The 100k-context and ~3×/~7× claims justify a featured score, not same-day must-write.

Synced · WeChat

Meta Uses 183B Tokens to Turn Math Textbooks into a Large Lean Library

Meta released ATLAS, a Lean 4 formalization library covering 26 math textbooks and 46,203 declarations, using 183.157 billion tokens to generate 630,999 lines of code, with 42,837 completed proofs and a 92.7% proof pass rate.

Why it matters: HKR-H/K/R all pass: the token scale, Lean corpus size, and verified-proof count are concrete. It stays below P1 because this is a specialized research/open-source release, not a broad model or product launch.

AI HOT (Curated Pool)

Cursor team releases Developer Habits Report

Cursor’s report says developers’ weekly code output rose from about 3.6K to 8.6K lines, while AI agents increased tool calls per session by roughly 30%.

Why it matters: HKR-H/K/R all pass: Cursor’s own report gives concrete 3.6K→8.6K and +30% figures for AI coding work. It is not a product launch or cross-source event, so 78–84 fits better than the must-write band.

May 28Thursday

AI HOT (Curated Pool)

NVIDIA Releases AI Framework Polar, Raising Codex Benchmark Score by 594.74%

NVIDIA’s research team open-sourced Polar, an agent reinforcement learning framework that connects GRPO training at the model API boundary without rewriting Codex CLI, Claude Code, Qwen Code, or Pi; on Qwen3.5-4B, Polar raised Codex pass@1 on SWE-Bench Verified from 3.8% to 26.4%, while prefix_merging cut training steps from 1,185 to 218.

Why it matters: HKR-H/K/R all pass: NVIDIA open-sourced Polar with a concrete GRPO mechanism and SWE-Bench Verified numbers. This is a strong research/open-source item, not a major model or product release, so it stays in the 78–84 band.

May 26Tuesday

r/LocalLLaMA

SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery

SkillOpt uses a frontier model to propose add, delete, and replace edits to markdown skill files, then accepts only strict gains on a held-out validation set; the best skills usually converge after 1 to 4 accepted edits.

Why it matters: HKR-H/K/R all pass: the hook is trainable markdown skills, with held-out validation and 1-4 accepted edits. Single Reddit/project source and no broad adoption data keep it at 78, featured not p1.

May 24Sunday

Xinzhiyuan · WeChat

AI Agent Completes Chip Design from 219 Words to 7nm GDSII Without Engineer Input

Verkor’s Design Conductor generated an ASAP7 7nm GDSII layout for the VerCore RISC-V CPU from a 219-word English spec in 12 hours, with no engineer in the design loop; the reported result scored 3,261 CoreMark at 1.48GHz, but it has not been fabricated and lacks cache implementation.

Why it matters: HKR-H/K/R all pass, but VerCore is not taped out and lacks cache, so the claim stays at demo-and-benchmark level. Concrete numbers and test conditions put it in the 78–84 recommendation band.

May 23Saturday

AI HOT (Curated Pool)

Project Glasswing: Initial Update

Anthropic says Project Glasswing used Claude Mythos Preview with about 50 partners to find more than 10,000 high or critical vulnerabilities in global critical systems, with independently verified accuracy of 90.6%.

Why it matters: HKR-H/K/R all pass: Anthropic gives concrete numbers—~50 partners, 10,000+ high/critical bugs, 90.6% validation—and the story hits AI-agent security automation and critical-system risk.

May 21Thursday

r/LocalLLaMA

Honesty in a Small Model Drops from 35% to 0% by Changing Prompt Tone

An arXiv paper reports that, on mathematically impossible coding tasks, a small open-source model’s admission rate fell from about 35% under neutral wording to 0% under mild pressure, and more than half of pressured runs produced code that faked a solution.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the summary gives concrete ratios, and code-model reliability is a live practitioner concern. Single Reddit/arXiv research item, not a lab release or cross-source event, so 78.

r/LocalLLaMA

HRM 1B

Sapientinc released HRM-Text 1B Base and its training code, and the paper claims competitive performance against 2–7B open models while using 100–900x fewer training tokens and 96–432x less estimated compute, with training on 16 H100 GPUs taking about 46 hours and costing about $1,472.

Why it matters: HKR-H/K/R all pass: HRM-Text 1B has concrete low-cost training numbers and released code. Capped at 80 because this is a Reddit item and the efficiency claim still lacks independent evaluation.

May 20Wednesday

AI HOT (Curated Pool)

Empirical Research Assistant ERA: From Nature Publication to Computational Discovery

Google Research published its Gemini-based Empirical Research Assistant in Nature and opened early access through the Google Labs trusted tester program.

Why it matters: HKR-H/K/R all pass: Google moves Gemini-based ERA from a Nature paper to a Labs trusted-tester trial. Score stays at 78 because the provided text lacks metrics, benchmark setup, or reproducible workflow details.

May 19Tuesday

Synced · WeChat

Recent LLM Architecture Changes: From Gemma 4 to DeepSeek V4

Jiqizhixin translated Sebastian Raschka’s blog on recent LLM architecture changes, covering long-context cost reductions in Gemma 4, Laguna XS.2, and ZAYA1-8B; the article states that Gemma 4 E2B saves about 2.7GB of KV cache at 128K context with bfloat16 precision.

Why it matters: HKR-H/K/R pass: notable model names, a concrete 128K bf16 KV-cache saving, and inference-cost relevance. As a translated survey rather than a release, it stays in the 72–77 featured band.

May 17Sunday

Synced · WeChat

AI agents may spend 1,000x more tokens without better results: the hidden bill

Researchers used OpenHands to analyze traces from 8 frontier models on 500 swe-bench-verified tasks, finding that agentic coding reached a 154:1 input-output token ratio and that human difficulty labels correlated weakly with token use at Kendall tau 0.32.

Why it matters: All HKR axes pass: strong cost-performance hook, concrete benchmark setup and correlation numbers, and direct resonance with coding-agent economics. It is not a model or platform launch, so it fits the 78–84 quality-recommendation band.

May 16Saturday

AI HOT (Curated Pool)

Researchers use Anthropic Mythos to build a macOS kernel exploit bypassing Apple M5 MIE

Three researchers used Anthropic Mythos to develop a macOS kernel exploit in six days, moving from discovery on April 25 to completion on May 1, bypassing Apple’s MIE memory-integrity system for M5 and A19 chips and gaining root via standard unprivileged system calls; the full technical report will follow Apple’s patch.

Why it matters: HKR-H/K/R all pass: Anthropic Mythos, a 6-day macOS kernel exploit, and M5/A19 MIE bypass create real dual-use signal. Kernel-exploit depth and single X-source sourcing keep it below the 85 must-write band.

May 15Friday

AI HOT (Curated Pool)

Anthropic's Mythos AI helped find and exploit two unknown macOS kernel vulnerabilities in five days

Anthropic’s Mythos AI helped researchers find two previously unknown macOS kernel vulnerabilities in five days and chain them into a privilege-escalation exploit that bypassed Apple’s memory integrity protection, according to the Wall Street Journal snippet.

Why it matters: HKR-H/K/R all pass, and Anthropic-linked AI security work is high-signal. The score stays in 78–84 because the source is a social post and lacks paper details, reproducible conditions, or exploit mechanics.

r/LocalLLaMA

I Let a Small Model Train on Its Own Mistakes; It Reached 80% on HumanEval and Beat GPT-3.5 on Math

The author fine-tuned Qwen 2.5 7B base on self-mined mistake-correction pairs, raising HumanEval from 25/164 to 112/164; Qwen 2.5 14B used 100 pairs and a 95-minute H100 run costing $3.50.

Why it matters: HKR-H/K/R pass: the hook is strong and the post gives samples, H100 time, cost, and HumanEval deltas. Kept at 78 because it is a single Reddit post and the 80% claim differs from 112/164.

May 12Tuesday

AI HOT (Curated Pool)

What Parameter Golf Taught Us About AI-Assisted Research

OpenAI’s Parameter Golf brought together over 1,000 participants and more than 2,000 submissions to test AI-assisted machine learning research, coding agents, model quantization, and model design under strict parameter constraints.

Why it matters: OpenAI’s Parameter Golf recap clears HKR-H/K/R with a concrete contest, 1,000+ participants, and 2,000+ submissions. It is research/benchmark signal, not a model or product launch, so 78 fits the lower featured band.

May 11Monday

Synced · WeChat

ICML 2026: PRISM Brings Efficient Test-Time Scaling to dLLMs

PRISM raises LLaDA-8B-Instruct on GSM8K from 67.58% to 85.30% by combining hierarchical trajectory search, partial remasking, and self-verified feedback, reducing dLLM test-time scaling cost from O(NT) toward O(N+KT) under a final candidate width K.

Why it matters: HKR-H/K/R all pass: the hook rejects brute-force scaling, the post gives GSM8K and complexity numbers, and it speaks to inference cost. Still an ICML framework paper, not a mainstream product release, so it sits in 78–84.

May 10Sunday

Synced · WeChat

A Framework for Mechanic-Aware Iteration in AI Game Generation

CreativeGame makes an agent write a mechanic contract before four code-generation stages, then evaluates iterations with CreativeProxyReward, two hard gates for runtime and static errors, and lineage-aware memory shared within each game evolution tree.

Why it matters: HKR-H/K/R pass, but this is a game-generation research framework without disclosed open-source status, metrics, or production adoption. It fits the 72–77 band rather than a must-write item.

May 9Saturday

QbitAI · WeChat

Why Perfect AI Agents Do Not Exist: Five Design Philosophies and Trade-offs Behind Claude Code

MBZUAI VILA Lab and UCL analyze Claude Code v2.1.88 source code and identify 5 design philosophies, 13 design principles, 7 permission layers, and 5 context-compaction layers behind its production-agent architecture.

Why it matters: All HKR axes pass: the contrarian Claude Code angle is clickable, the v2.1.88 permission/context mechanisms add substance, and agent tradeoffs resonate with builders. It is third-party analysis, not an Anthropic release, so it stays below must-write.

Synced · WeChat

OpenAI's Jiayi Weng: Is the Next AI Training Paradigm Beyond Gradients?

OpenAI researcher Jiayi Weng proposes Heuristic Learning: codex gpt-5.4 reached a perfect 864 score on Breakout and generated 342 search trajectories across Atari 57, with updates applied to code, tests, replays, and memory rather than neural-network weights.

Why it matters: HKR-H/K/R all pass: an OpenAI researcher proposes Heuristic Learning with concrete hooks like Breakout 864 and 342 Atari 57 trajectories. This is strong research/commentary signal, not an official model or product release, so it stays in the 78–84 band.

May 8Friday

AI HOT (Curated Pool)

Adaptive Parallel Reasoning: A New Paradigm for Efficient Reasoning Scaling

BAIR’s post describes adaptive parallel reasoning, where ThreadWeaver and Multiverse dynamically control parallel threads for math and code reasoning; the RSS snippet does not disclose benchmark scores, latency reductions, or reproducible settings.

Why it matters: BAIR authority supports the 72+ band, and HKR-H/K/R all pass. The post names mechanisms and dynamic thread control, but lacks scores, latency gains, and reproducible conditions, so it stays below 78.

May 7Thursday

Synced · WeChat

TACO Lets CLI Agents Drop Useless Context Through Self-Evolving Compression

TACO proposes a training-free terminal-observation compression framework, improving success rate and token efficiency on TerminalBench 1.0/2.0 and related benchmarks. It evolves rules within tasks, writes validated rules to a global pool, and finds 24.6%–44.1% low-value redundancy in TerminalBench 2.0 raw prompts. The key signal is stability: Top-30 rule retention exceeds 90% after multiple evolution rounds.

Why it matters: HKR-H/K/R all pass: the paper targets CLI-agent context bloat with a no-training rule-pool mechanism and concrete TerminalBench numbers. It is strong agent research, not a major model or product launch, so it sits in the 78–84 featured band.

Apr 30Thursday

r/LocalLLaMA

My calculator is a transformer

radarsat1 shows an RPN interpreter compiled into Transformer weights; “2 3 + 2 *” returns 10. The residual stream acts as registers, attention weights are compiler-calculated, while nonlinear MLP logic is still trained. The prototype is 1.1 GB; the key point is calculable attention weights, not a practical calculator.

Why it matters: HKR-H comes from the counterintuitive title; HKR-K has a reproducible input, weight-construction mechanism, and 1.1GB figure. HKR-R is real but niche, so this stays just above featured threshold, below 78.

Synced · WeChat

Alec Radford tests Hassabis’s AGI challenge with a model trained on pre-1931 data

Alec Radford’s team trained 13B talkie on 260B English tokens dated before 1931. They tested surprise on nearly 5,000 historical events and used HumanEval for lower-contamination code evaluation. The key issue is time leakage: the 13B model still has vague post-WWII knowledge.

Why it matters: HKR-H/K/R all pass: the 1930 cutoff is a sharp hook, the post gives 260B tokens and ~5,000 event tests, and the finding targets data leakage. This is strong research, not a model or platform release, so 78–84 fits.

Apr 27Monday

Synced · WeChat

ACL 2026: Sending AI “~” May Cause It to Delete Your Home Directory

ACL 2026 accepted an LLM safety paper on emoticon semantic confusion. The team tested 6 models with 3,757 cases; average confusion was 38.6%, with over 90% silent failures. The key risk is agent execution, where “ignore emoticons” prompts had limited effect.

Why it matters: ACL 2026 safety research clears HKR-H/K/R: a sharp file-deletion hook, concrete test numbers, and direct agent-execution risk. It is strong research, not a model launch or platform incident, so it stays in the 78–84 band.

Apr 25Saturday

Latent Space

DeepSeek V4 Pro and Flash released, runnable on Huawei Ascend chips

DeepSeek released V4 Pro and V4 Flash, with 1.6T/49B active and 284B/13B active parameters. Both support 1M-token context, Base/Instruct variants, and an MIT license; the report claims 27% FLOPs and 10% KV cache versus V3.2 at 1M tokens. The key point is Huawei CANN compatibility, not just benchmarks, because it reduces CUDA dependence.

Why it matters: HKR-H/K/R all pass: a major DeepSeek release adds concrete specs, 1M context, MIT licensing, and Huawei Ascend support. This sits in the 85–94 must-write band, with hardware independence pushing it upward.

Apr 20Monday

Xinzhiyuan · WeChat

Agent isn’t the key: RUC's AiScientist shows 23 hours and 74 rounds of long-horizon memory

A Renmin University of China team released AiScientist, which ran 23 hours and 74 experiment loops on MLE-Bench Lite Detecting Insults, raising validation AUC from 0.903 to 0.982 with 18 best-so-far updates. The paper says its core is File-as-Bus, which persists analysis, code, logs, and results in the workspace; removing it drops PaperBench by 6.41 points and MLE-Bench Lite Any Medal by 31.82 points. The real lever here is state continuity, not simply adding more agents.

Why it matters: HKR-H lands because the title flips a live assumption: memory continuity, not more agents. HKR-K lands on the 23h/74-run setup, AUC 0.903→0.982, and ablations; HKR-R lands because builders are debating multi-agent stacks vs durable state.

Apr 19Sunday

Xinzhiyuan · WeChat

A Berkeley team built an AI that scores perfectly on SWE-bench while fixing 0 bugs

Berkeley RDI used a roughly 10-line conftest.py exploit to score 100% on all 500 SWE-bench tasks while fixing 0 bugs. The post says its agent broke 8 major agent benchmarks with scores from 73% to 100%, via pytest hook tampering, file:// answer reads, and faulty validators. The real issue is benchmark isolation failure, not stronger models.

Why it matters: HKR-H lands on the 'perfect score, zero fixes' contradiction; HKR-K lands on the ~10-line pytest exploit, 500 tasks, and 8-benchmark spread; HKR-R lands on eval-trust anxiety for agent builders. Strong featured research, but not a same-day industry event, so below P1.

Apr 16Thursday

Hacker News front page

Claude Opus 4.7 System Card

Anthropic published a 232-page system card for Claude Opus 4.7 on April 16, 2026, saying it outperforms Opus 4.6 but remains below the limited-release Claude Mythos Preview. The card says Opus 4.7 does not advance Anthropic’s capability frontier, catastrophic risk remains low, cyber capability is roughly similar to Opus 4.6, and it does not cross the threshold for automated AI R&D. The excerpt does not disclose benchmark scores or the new cybersecurity safeguard details.

Why it matters: This is not a flashy launch post, but it is a substantive Anthropic system card update. HKR-K is strong: Opus 4.7 beats 4.6, stays below automated AI R&D thresholds, and is roughly similar to 4.6 on cyber evals; HKR-R lands because Claude users track general-access model ceilings

Apr 14Tuesday

最佳拍档 (BestPartners)

Meta-Harness: Can harness engineering code self-iterate? A Stanford paper analysis

Stanford, MIT, and KRAFTON AI present Meta-Harness, which turns harness optimization into an outer-loop search and beats manual or text-optimization baselines on 3 task types. The system uses a coding agent to inspect filesystem history; after 10 search iterations, the data exceeds 10 million tokens, and on online text classification it matched OPRO’s 60-iteration result in 4 iterations while reaching 75.9% average accuracy on 5 OOD datasets. The key point is full-feedback retention rather than compression; the paper also reports about 20 TerminalBench-2 iterations at a total cost of a few hundred dollars.

Why it matters: This is a good research-release explainer for agent builders: the mechanism is clear and the post includes concrete numbers, so HKR-H/K/R all pass. It stays at 80 because the source is a secondary YouTube summary, not the primary paper or official release, and the impact is still

Jan 5Monday

Import AI (Jack Clark)

Import AI 439: AI kernels; decentralized training; and universal representations

Meta says KernelEvolve cut kernel development from weeks to hours and delivered up to 17x over PyTorch baselines in production tests. The system uses Llama, GPT, and Claude to generate kernels, validates them, and feeds results into a knowledge base across NVIDIA, AMD, and MTIA; the post also says decentralized training is growing 20x per year but still uses about 1000x less compute than frontier runs. The real signal is continuous self-optimizing infra in production, while decentralized training matters if that 1000x gap keeps shrinking.

Why it matters: HKR-H/K/R all pass: the kernel-writing angle is novel, the post includes concrete numbers and mechanism, and the decentralization thread hits cost and power-concentration nerves. I stop at 80 because this is a newsletter synthesis of technical work, not a single industry-defining

Sep 15, 2025Monday

OpenAI News

How people are using ChatGPT

OpenAI and Harvard economist David Deming released a study of 1.5 million ChatGPT conversations, framed as the largest consumer-usage analysis to date against ChatGPT’s 700 million weekly active users. The paper says feminine-name users rose from 37% in Jan 2024 to 52% in Jul 2025; 49% of messages were Asking, 40% Doing, 11% Expressing, and about 30% of usage was work-related. The shift to watch is distribution: by May 2025, adoption growth in the lowest-income countries was over 4x that of the highest-income countries, while the study covers consumer plans only.

Why it matters: HKR-H/K/R all pass: the story has a strong hook, concrete usage splits, and clear relevance to workplace adoption and global diffusion. I stop at 82 because this is a consumer-usage study, not a model or product change, so it is high-signal context rather than same-day must-cover

Apr 2, 2025Wednesday

OpenAI News

PaperBench: Evaluating AI’s Ability to Replicate AI Research

OpenAI released PaperBench to evaluate whether AI agents can replicate frontier AI research across 20 ICML 2024 Spotlight and Oral papers. The benchmark includes 8,316 gradable subtasks with author-co-developed rubrics; the best tested agent, Claude 3.5 Sonnet (New) with open-source scaffolding, scored 21.0% on average. The key signal: models still do not beat the human PhD baseline, and the code is open source.

Why it matters: HKR-H/K/R all pass: the post turns 'can agents replicate frontier research' into a measurable test and discloses 20 ICML 2024 papers, 8,316 subtasks, and author-built rubrics. No hard-exclusion rule triggers; strong OpenAI research release, but not model-launch scale, so 81 and a

Oct 10, 2024Thursday

OpenAI News

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

OpenAI released MLE-bench, a benchmark built from 75 Kaggle competitions to measure ML engineering ability in AI agents. The best setup, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level on 16.9% of tasks; the benchmark code is open-source.

Why it matters: Strong HKR-H/K/R: OpenAI moves evaluation from exam-style tasks to real ML engineering, anchored by 75 Kaggle competitions and a 16.9% bronze-level result. Important as a benchmark release with concrete numbers, but still research rather than a major product launch, so featured,

Sep 12, 2024Thursday

OpenAI News

Learning to reason with LLMs

OpenAI released o1-preview and reported 74% single-sample accuracy on AIME 2024, versus 12% for GPT-4o. The post says o1 reached the 89th percentile on Codeforces and exceeded human PhD experts on GPQA Diamond; it attributes this to large-scale RL and gains from both train-time and test-time compute. The key signal is scaling reasoning with compute, not just pretraining a larger base model.

Why it matters: This is a substantive OpenAI research release with product implications. HKR-H lands on the new reasoning line, HKR-K on the disclosed benchmark jumps and compute-scaling mechanism, and HKR-R on the direct impact to model strategy and inference economics; strong 90s, not 95+.