Skip to content

#编码

10 today

Jun 10Wednesday

AI HOT (Curated Pool)

Claude Fable 5 and Claude Mythos 5

Anthropic launched Claude Fable 5 and Claude Mythos 5 at $10 per million input tokens and $50 per million output tokens. Fable 5 leads FrontierCode among frontier models, while Mythos 5 reports about 10x acceleration in drug design and about 80% scientist preference in blinded molecular biology hypothesis tests.

Why it matters: HKR-H/K/R all pass: this is an official Anthropic dual-model release with pricing, coding benchmark, and drug-design speed claims. As a major Claude model update plus Anthropic substantive-update bump, it sits in the 85–94 band.

AI HOT (Curated Pool)

Cohere’s First Coding Model North Mini Code Is Free and Open Source

Cohere released its first coding model, North Mini Code, on OpenCode for free, with a 256K context window and full open-source availability.

Why it matters: HKR-H/K/R all pass: Cohere’s first code model has a 256K context and free open-source access in OpenCode. Missing benchmarks, model size, and license detail keep it at the low end of the 78–84 band.

Hacker News front page

System Card: Claude Fable 5 and Claude Mythos 5

Anthropic published a 319-page system card for Claude Fable 5 and Claude Mythos 5, stating that Fable 5 is for general use with biology and cybersecurity safeguards, while Mythos 5 lifts relevant safeguards and is limited to trusted partners starting with Project Glasswing.

Why it matters: HKR-H/K/R all pass: Anthropic documents two Claude 5 configurations, calls Mythos 5 its most capable model, and gives safety-gating details. This is a same-day Claude substantive update, placed in the 85–94 band.

AI HOT (Curated Pool)

GitHub Copilot CLI Adds Custom AI Agents to Turn One-Off Terminal Prompts into Workflows

GitHub Copilot CLI added custom AI agents that understand a developer’s tech stack and team workflows; the post does not disclose configuration details, rollout scope, or pricing.

Why it matters: Official GitHub product update with HKR-H/R: custom Copilot CLI agents matter for developer workflows. HKR-K is weak because setup, rollout, and pricing are missing, so it sits at the featured threshold.

Jun 9Tuesday

AI HOT (Curated Pool)

Cohere Releases North Mini Code, an Open Coding Model for Developers

Cohere released North Mini Code, a 30B-parameter MoE coding model with 3B active parameters, under Apache 2.0; it supports 64K/128K context lengths and reaches 80.2% pass@10 on SWE-Bench Verified.

Why it matters: HKR-H comes from a compact MoE code model with a strong SWE-Bench claim; HKR-K has params, license, context, and benchmark. Cohere is notable but not a frontier-lab launch, so this fits the 78–84 open-source code-model band.

AI HOT (Curated Pool)

Qwen3.7-Max Delivers Mobile and Web Apps from Scratch Using One Document

Qwen3.7-Max delivered mobile and web applications from a roughly 150,000-character product research document without design files or backend code; each client took about 4 hours, used staged constraint injection and error feedback, and the web app passed typecheck, build, and 34 reachable routes.

Why it matters: HKR-H/K/R all pass: the coding-agent claim is clickable, quantified, and emotionally relevant to developers. The summary lacks eval setup, failure rate, and human-intervention detail, so it stays in the 78–84 band.

Hacker News front page

Microsoft's Open Source Tools Were Hacked to Steal AI Developers' Passwords

The title says Microsoft's open source tools were hacked to steal passwords from AI developers; the RSS snippet does not disclose the affected tools, attack mechanism, timeline, or victim count.

Why it matters: TechCrunch plus HN front-page placement supports source weight, and the title hits HKR-H and HKR-R. HKR-K fails because tools, mechanism, and victim scale are missing, so the score stays at the featured floor.

Latent Space

Cognition launches FrontierCode: a coding benchmark that asks 'would you actually merge this?'

Cognition built FrontierCode, a benchmark that scores code on mergeability and maintainability, not just passing unit tests. Tasks were designed with open-source maintainers, each taking 40+ hours, and evaluated on regression safety, cleanliness, scope, test correctness, and maintainability. The best model, Opus 4.8, hits only about 13% on the hardest tier—far below the 50%+ common on SWE-Bench-style evals. The post also notes METR found many SWE-bench-passing PRs wouldn't actually be merged, and FrontierCode directly measures that false-positive problem.

Why it matters: Cognition's FrontierCode shifts code eval from 'passes tests' to 'mergeable,' with 40+ hour task design and scoring on maintainability. Opus 4.8 leads the hardest tier. A real addition to the benchmark landscape, but too new for community replication — 78 feels right.

AI HOT (Curated Pool)

AI coding unicorn Cursor picks London for European HQ; SpaceX holds $60B acquisition option

Cursor set its European headquarters in London and plans to hire about 200 people; SpaceX holds an option to acquire Cursor for $60 billion or pay $10 billion for a new partnership.

Why it matters: HKR-H/K/R all pass: Cursor is a core AI coding player, and the $60B option plus 200-person London expansion lifts this above routine office news. Thin sourcing and no disclosed trigger terms keep it below the 78 band.

AI HOT (Curated Pool)

Xiaomi MiMo and TileRT Release UltraSpeed Mode, 1T Model Exceeds 1,000 Tokens/s

Xiaomi MiMo and TileRT released MiMo-V2.5-Pro-UltraSpeed, a 1T-parameter model mode exceeding 1,000 tokens/s, with API access open from June 9 to June 23, 2026, at 3× the MiMo-V2.5-Pro price and about 10× the speed.

Why it matters: HKR-H/K/R all pass, with a domestic flagship-model bump for Xiaomi. Missing hardware, batch, concurrency, and test conditions keep it in the 78-84 band rather than p1.

AI HOT (Curated Pool)

FrontierCode benchmark sets a new AI coding evaluation bar, with top maintainer approval at 13.4%

Cognition released FrontierCode, a coding benchmark built from 150 tasks by more than 20 open-source maintainers and judged against over 3,000 rules, with Claude Opus 4.8 reaching 13.4% approval in the hardest tier and GPT-5.5 reaching 6.3%.

Why it matters: HKR-H/K/R all pass: FrontierCode has a strong 13.4% hook, concrete maintainer-built methodology, and clear coding-agent resonance. Single-source benchmark news keeps it in the 78–84 band, not must-write territory.

AI HOT (Curated Pool)

Migrating GitHub CI to Hugging Face Jobs

Hugging Face describes using huggingface/jobs-actions to run GitHub Actions CI as HF Jobs, where the Trackio project cut CPU job time by about 30% and added a GPU test suite using CPU, t4-small, or h200 hardware.

Why it matters: HKR-H/K/R pass via a concrete CI-to-HF Jobs workflow, ~30% speedup, and GPU-test pain point. Scope is ML tooling, not a major platform release, so it sits at the featured threshold.

AI HOT (Curated Pool)

Claude Supports Apple Foundation Models Framework With New Swift Package

Anthropic released a Swift package that lets Apple developers call Claude inside the Foundation Models framework with three lines of code, returning typed Swift values and handing off multi-step reasoning, code generation, web search, and data analysis on iOS 27, macOS 27, and related platforms.

Why it matters: HKR-H/K/R all pass: Anthropic is shipping a concrete Claude Swift package for Apple Foundation Models, but this is a developer integration rather than a model release, so it sits high in the 78–84 featured band.

The Verge · AI

Apple is using AI to fix Safari’s extension problem

Apple demonstrated Safari using Apple Intelligence to generate an extension from a text prompt, with a Recipe Keeper example for saving recipes and notes; the RSS snippet does not disclose release timing, required OS versions, or developer restrictions.

Why it matters: HKR-H/K/R pass, but the post gives only a demo and the Recipe Keeper example; launch timing, OS version, and developer limits are not disclosed. This fits a mid-weight product update at 73, below the 78 band.

AI HOT (Curated Pool)

OpenAI plans AI-led research by 2028

Sam Altman said OpenAI plans to have AI perform a large share of its research by March 2028, and the post lists three goals: building automated AI researchers, using them for science and production, and giving each person a personal AGI.

Why it matters: HKR-H/K/R all pass: dated OpenAI AGI-research roadmap with March 2028 and three goals. It stays below P1 because the item is an X repost/summary, not a primary launch or detailed Sam Altman essay with mechanisms.

r/LocalLLaMA

Levi: Run AlphaEvolve on Your Local Qwen 30B

LEVI runs an AlphaEvolve-like search system with Qwen3-30B-A3B and reports tests on ADRS, IFBench, and HotpotQA, claiming up to 35x lower cost overall and up to 12x fewer evals under the same single-model, same-budget comparison.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with model, benchmarks, and cost ratios only; code maturity and reproducibility details are not disclosed. Scores as a strong open-source agent/inference item, not a major release.

Jun 8Monday

AI HOT (Curated Pool)

Hivemind launches continuous learning for AI coding agents

Hivemind released continuous learning for AI coding agents, collecting trajectories from Claude Code, Codex, Cursor, Hermes, and Pi, converting them into reusable skills stored in users’ cloud storage, with SkillOpt matching or leading all 52 test settings.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight Hivemind feature launch without major-lab weight or cross-source lift. The 52-setting result gives it enough substance for low featured.

r/LocalLLaMA

DFlash Speculative Decoding and KV Cache Compression on RTX 5090 Show 3.26x Speedup

The author tested Qwen3.6-27B on an RTX 5090 with DFlash plus KV cache compression, reaching up to 3.26x speedup; q4_0/turbo4 delivered 3.18x speedup with only +0.02% PPL on WikiText-2.

Why it matters: HKR-H/K/R all pass: RTX 5090 testing, DFlash speculative decoding, KV cache compression, 3.26x speedup, and PPL delta are concrete. Single Reddit source keeps it near the featured floor.

r/LocalLLaMA

Weird to get near-linear scaling by adding another GPU?

A Reddit user benchmarked qwen3.6-27b-autoround-int4 on 1x3090 versus 2x3090. Narrative decode rose from 53 TPS to 94 TPS, and code decode rose from 62 TPS to 120 TPS, under no NVLink, 8x/8x PCIe, P2P automatically enabled, tensor parallelism set to 2, and different KV-cache settings.

Why it matters: HKR-H/K/R all pass: the result is counterintuitive, includes concrete TPS and TP conditions, and speaks to local-inference cost. Single Reddit test lacks multi-model replication and full setup details, so it stays near the featured threshold.

AI HOT (Curated Pool)

ChatGPT Is Set to Become AgentGPT

OpenAI is preparing ChatGPT’s largest redesign since its 2022 launch, shifting it toward an agent platform that integrates Codex, image generation, Canva, and Booking, with web and mobile rollout planned in the coming weeks. ChatGPT has 900 million weekly active users, 50 million paid users, and $2 billion in monthly revenue, but the post says it remains unprofitable.

Why it matters: HKR-H/K/R all pass, but this is a single X post and the body lacks official timing, access scope, and pricing. It sits at the top of 78–84 rather than P1 because the revamp is not yet shipped.

Jun 7Sunday

r/LocalLLaMA

Qwen3.6 35B-A3B on a Laptop: My Zero-to-One Moment

A Reddit user ran Qwen3.6 35B-A3B on an ASUS Zenbook Pro 14 with RTX 4060 8GB VRAM and 64GB RAM, reaching about 27 TPS at 32k context and 18 TPS at 256k context. The setup uses llama.cpp, unsloth’s IQ3_XXS GGUF quantization, and a 262144-token context flag.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment, not an official release or paper. Concrete hardware, quantization, context, and TPS clear the featured bar, but keep it in the 72–77 band.

AI HOT (Curated Pool)

A Hokkaido Broccoli Farmer’s 8 Real AI Uses with ChatGPT and Codex

Hokkaido farmer Hiroki Tomiyasu uses ChatGPT and Codex for 8 farm tasks, including broccoli disease recognition, NDVI monitoring, ESP32 greenhouse control, LINE chatbots, sowing-count tracking, RTK-GPS steering study, and an Airtable farm database.

Why it matters: HKR-H/K/R all pass: the hook is unusual, the post names 8 farm workflows, and Codex moving into physical operations will travel among practitioners. Single-X sourcing and missing outcome metrics keep it near the featured floor.

Xinzhiyuan · WeChat

Anthropic co-founder says Claude now writes 80% of merged code

Jack Clark said Claude now produces 80% of Anthropic’s merged code and projected the share may reach 100% within two years; the article also says Anthropic engineers merged 8 times more code per person per day in Q2 2026 than in 2024.

Why it matters: HKR-H/K/R all pass: Jack Clark’s Anthropic coding numbers give a strong hook, concrete facts, and clear labor-productivity resonance. This is not a model launch or major product update, so it stays in the 78–84 band.

Synced · WeChat

ICML 2026 | FusionRoute: From Expert Routing to Self-Correction in Multi-LLM Collaboration

FusionRoute proposes a token-level multi-LLM collaboration method that freezes expert models and trains a lightweight router to select an expert for each token while merging router logits with expert logits. The paper evaluates it on GSM8K, MATH-500, HumanEval, MBPP, IfEval, and 500 PerfectBlend prompts.

Why it matters: HKR-H/K/R pass: token-level LLM routing is a strong research hook with concrete mechanics. The article lacks lift numbers, code link, and deployment cost, so it stays at the lower featured band.

r/LocalLLaMA

Cohere's Unreleased Coding Model Gets Early Access for LocalLLaMA

Cohere employee Nick Frosst opened early testing of BLS-Mini-Code-1.0 to LocalLLaMA, with weights on Hugging Face before public launch. The coding model has 30B total parameters and 3B active parameters, and Cohere says token output tests are in line with similar models in its size class.

Why it matters: HKR-H/K/R all pass: early-access Cohere coding weights with 30B/3B specifics matter to local-model users. Reddit sourcing and missing evals, license, and training details keep it in the low featured band.

Jun 6Saturday

AI HOT (Curated Pool)

GitHub open-sources Spec Kit to guide AI coding with product specifications

GitHub released the open-source Spec Kit, shifting AI coding from direct implementation to product specifications, gap clarification, technical planning, task breakdown, and agent execution, with support for 30+ agent integrations including Copilot, Claude Code, Codex, Gemini, Cursor, and Qwen, and 109K+ GitHub stars.

Why it matters: HKR-H/K/R all pass: GitHub’s Spec Kit gives a concrete spec-first agent workflow plus 30+ integrations and 109K+ stars. It is a strong tooling story, not a model- or platform-level launch.

r/LocalLLaMA

The Gap Between Claude and Local: Can a Self-Hosted Coding Agent Compete?

The author compared five coding-agent setups on a Laravel 12 + Livewire Playwright E2E task; Claude Opus 4.7 with 1M context produced 203 tests, while the strongest local OpenCode arm on a 24GB RTX 4090 produced 140 tests, compacted context four times, and needed seven manual nudges.

Why it matters: HKR-H/K/R all pass: a first-person Claude-vs-local coding-agent test with concrete counts. It stays below P1 because it is a single Reddit experiment, not a standardized benchmark or major release.

Xinzhiyuan · WeChat

$280 per task: 1,000 engineers teach Claude to write better code

Anthropic is using Snorkel’s Marlin project to recruit about 1,000 software engineers who review Claude Code outputs for $280 per task, with a workflow covering GitHub repository pull requests, A/B comparisons of two generated code versions, and scoring for correctness, security, reliability, and maintainability.

Why it matters: HKR-H/K/R all pass: price, scale, and review mechanics are concrete, and the Claude Code labor angle lands with AI coders. It fits featured, but not p1, since this is not a new model or capability launch.

Synced · WeChat

DeepSeek V4 Proves Math with 500x Cost Advantage as Agent System Sets Records

Princeton researchers released Goedel-Architect, an agent framework for Lean formal theorem proving. Using DeepSeek-V4-Flash, it reached 75.6% pass@1 on PutnamBench, with $294 in API cost for 672 problems, compared with Hilbert’s 70.0% and about $170,000 cost.

Why it matters: HKR-H/K/R all pass: Goedel-Architect pairs a 75.6% PutnamBench score with $294 for 672 problems, versus Hilbert at about $170k. It is still research-heavy, so it stays in the 78–84 band rather than P1.

Jun 5Friday

r/LocalLLaMA

Microsoft released MAI models instead of something like Qwen3.6-27B or Gemma-4-31B

Microsoft AI released seven MAI models, with MAI-Thinking-1 listed as 1T A35B with a 256K context window and MAI-Code-1-Flash listed as 137B A5B with a 256K context window.

Why it matters: Microsoft shipping 7 MAI models with reasoning/code variants and 256K context clears HKR-K/R, and the Qwen/Gemma catch-up angle clears HKR-H. Reddit sourcing and missing benchmarks, license, and pricing keep it below P1.

Xinzhiyuan · WeChat

Anthropic warns of AI self-acceleration as OpenAI is said to cross a reliability threshold

Xinzhiyuan cites a Yann Dubois interview saying OpenAI crossed a reliability threshold around last December, while Anthropic’s internal data says per-person quarterly code contribution reached 8× the Q1 2024 level by Q2 2026.

Why it matters: HKR-H/K/R all pass: the cliff-edge framing is clickable, and the summary includes a timing claim plus Anthropic’s 8x coding metric. Capped at 82 because this is second-hand interview analysis, not an official release or reproducible test.

Hacker News front page

Show HN: I benchmarked LLM agents on fixing real-world security vulnerabilities

Giovanni Gatti benchmarked 5 LLM agents on 20 real CVEs across 18 Python projects, and the best solve rate across 300 runs was 50%.

Why it matters: HKR-H/K/R all pass: real vulnerabilities, a reproducible test scale, and a 50% best fix rate. As a Show HN individual benchmark rather than a lab release, it stays in the lower featured band.

AI HOT (Curated Pool)

Tencent's Dowson Tong: Most Tencent Code This Year Is AI-Generated

Dowson Tong said Tencent generated most of its code with AI this year, while engineers spent more time on architecture design and regularly guided and corrected AI outputs. Tencent invested 18 billion yuan in AI new products last year, and President Martin Lau said this year’s spending will at least double.

Why it matters: HKR-H/K/R all pass: a Tencent executive claims AI now generates most code and cites RMB 18B spend plus a doubling plan. It stays below P1 because the share is unquantified and self-reported.

Computing Life · Share · Yage

Grok Build 0.1: xAI’s Bet on Parallel Breadth

xAI launched Grok Build 0.1 in May 2026 as a coding agent built around parallel subagents; the post does not disclose benchmark results, cost figures, or specific privacy-policy terms.

Why it matters: HKR-H/K/R pass because xAI entering coding agents with parallel subagents is clickable, concrete, and relevant to developers. Missing benchmarks, cost, and privacy terms keep it at the featured floor.

AI HOT (Curated Pool)

Co-Existence and the End of Co-Intelligence

Ethan Mollick announced Co-Existence for an October 20 release and argues that co-intelligence is giving way to autonomous agents, citing late-2025 coding agents that a study links to 17x more code and Anthropic’s claim that AI now writes 80% of its code.

Why it matters: HKR-H/K/R all pass: Ethan Mollick’s essay has authority, a sharp framing, and concrete coding-productivity claims. It stays below 85 because it is commentary plus a book announcement, not a model release or reproducible experiment.

Hacker News front page

Anthropic's open-source framework for AI-powered vulnerability discovery

Anthropic published an open-source framework for AI-powered vulnerability discovery, and the HN item shows 58 points and 19 comments; the post does not disclose the framework mechanism, benchmark results, or deployment scope.

Why it matters: Anthropic source plus an open GitHub artifact clears HKR-H/R and the featured bar. HKR-K fails because mechanism, benchmarks, and scope are not disclosed, keeping it in the 72–77 band.

Financial Times · Technology

US National Security Agency Using Anthropic’s Mythos for Cyber Attacks

The title says the US National Security Agency is using Anthropic’s Mythos for cyber attacks; the RSS snippet only says Anthropic is in a legal battle with the Pentagon over the Claude model and does not disclose deployment scope.

Why it matters: Single-source FT story with strong HKR-H/R; HKR-K reaches a named Mythos/Claude-Pentagon dispute, but deployment scope is absent, keeping it in the 78–84 band.

AI HOT (Curated Pool)

Codex launches iOS app build plugin

Codex integrated the Build iOS Apps plugin, which lets users test iOS apps in an in-app browser, open SwiftUI previews, and hot-reload edits without leaving Codex.

Why it matters: HKR-H/K/R all pass: the hook is Codex handling iOS app testing, with concrete SwiftUI preview and hot reload details. This is a mid-weight OpenAI dev-tool update, not a model release; pricing and rollout scope are not disclosed.

Jun 4Thursday

Hacker News front page

Show HN: Cost.dev (YC W21) Makes Agents Cost-Aware and Cheaper to Call

Infracost launched Cost.dev, a local CLI for cloud-cost estimates in coding-agent workflows, and says it cut Claude output-token use by up to 79% and API cost by up to 67% versus a bare-Claude baseline.

Why it matters: HKR-H/K/R all pass: the local CLI cost-estimation mechanism and 79%/67% reduction claims are concrete. It is still a small vendor launch, so it sits at the featured floor, not same-day news.

Xinzhiyuan · WeChat

Silicon Valley CEO backs MiniMax M3 as it tops open-source rankings amid Chinese community debate

MiniMax M3 ranks first among open-source models on Artificial Analysis, and the article says it supports a 1M-token context window, used 100T-scale pretraining, and will open-source its weights and full technical report within 10 days.

Why it matters: HKR-H/K/R all pass: the hook is an open-source No.1 claim amid debate, with 1M context, 100T pretraining, and weights promised in 10 days. Since weights and full report are not out, this stays in 78–84, not P1.