Skip to content

#编码

10 today

May 3Sunday

r/LocalLLaMA

Local LLM Benchmark for Backend Generation via Function Calling: GLM vs Qwen vs DeepSeek

AutoBe posted a controlled backend-generation benchmark and says qwen3.5-35b-a3b matches gpt-5.4 on DB/API design. One shopping-mall run uses 200–300M tokens, costing $1,000–$1,500 per model at GPT 5.5 pricing. The key caveat is n=4 projects and self-scoring harness bias.

Why it matters: HKR-H/K/R all pass, but Reddit sourcing, n=4 projects, and self-eval harness bias keep it at the low featured band. Concrete cost and test constraints carry the score.

Synced · WeChat

Why CTOs at Billion-Dollar Companies Are Joining Anthropic as Engineers

Jiqizhixin lists at least six CTOs who joined Anthropic as individual contributors. Cases include Workday, You.com, Box, Super.com, and Adept AI from Jan 2025 to Apr 2026. The key issue is career leverage, not just AGI mission talk.

Why it matters: HKR-H/K/R all pass: the career-status reversal is clickable, the post gives 6 cases, and it touches AI talent competition. No hard exclusion, but it is commentary, not a model or product release.

Xinzhiyuan · WeChat

Claude Code helps Anthropic double revenue pace in two months

Semi Analysis says Anthropic’s ARR reached $44B, adding $35B over 12 months. Claude Code hit $2.5B annualized revenue by Feb 2026, while inference gross margin rose from 38% to over 70%. The key test is keeping enterprise usage, coding-agent revenue, and inference margin together.

Why it matters: HKR-H/K/R all pass: SemiAnalysis gives hard ARR, Claude Code revenue, and inference-margin numbers. Not a model launch, but it materially shifts the view of Claude Code monetization.

r/LocalLLaMA

Built a C++17 transformer from scratch with 0.83M params and CPU training

Reddit user Suspicious_Gap1121 released Quadtrix.cpp, a C++17 GPT-style model with 0.83M parameters. It uses 4 layers, 4 heads, 200d width, and a 128-character context; one CPU core trained on 31.4M characters for 76.2 minutes to 1.6371 nats val loss. The key detail is handwritten backprop for LayerNorm, attention, Q/K/V, dropout, and AdamW without PyTorch, BLAS, or autograd.

Why it matters: HKR-H/K/R all pass: the no-framework C++17 build is clickable, the training setup is specific, and local-LLM builders care about dependency-free control. It stays in the 72–77 band because it is a small personal project.

May 2Saturday

QbitAI · WeChat

Apple Support App Accidentally Shipped Claude.md, Revealing Internal Claude Code Use

Apple Support v5.13 shipped a Claude.md file on May 1 and was pulled within 24 hours. The file describes Juno AI and Live Agents switching through a Protocol layer, with client, agent, and assistant messages handled in one flow. The key issue is release review; the post does not disclose how the file entered production.

Why it matters: HKR-H/K/R all pass, but this is still an app-packaging incident, not a model or platform release. Apple scale and Claude.md details clear the featured bar; the review-chain failure is not disclosed.

TechCrunch · AI

Replit's Amjad Masad on the Cursor deal, fighting Apple, and why he'd rather not sell

Replit grew from $2.8M in 2024 revenue to a billion-dollar annualized target. The excerpt says Cursor is reportedly discussing a $60B SpaceX acquisition; the post does not disclose Masad's full Apple or sale comments.

Why it matters: HKR-H/K/R pass: TechCrunch has the Cursor $60B hook, Replit revenue target, and coding-tool exit tension. The excerpt lacks Masad’s full Apple and sale comments, so this stays in the 72–77 band.

May 1Friday

r/LocalLLaMA

PFlash: 10x prefill speedup over llama.cpp at 128K on an RTX 3090

PFlash cuts Qwen3.6-27B Q4_K_M 128K TTFT to 24.8s on an RTX 3090, versus 248.4s cold for llama.cpp. It uses a Qwen3-0.6B drafter to score token importance, keeps 5% of spans, and runs C++/CUDA without Python, Triton, or PyTorch. The quality caveat is clear: only NIAH single-needle passes from 32K to 128K; RULER and multi-needle results are not disclosed.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit claim with quality evidence limited to single-needle NIAH 32K–128K. RULER and multi-needle results are not disclosed, so it stays at featured threshold.

Xinzhiyuan · WeChat

Claude Code's Real Story: 98.4% of What Works Is Engineering, Not AI

VILA-Lab analyzed 512,000 lines of Claude Code v2.1.88 and found 1.6% tied to AI decision logic. The other 98.4% is deterministic infrastructure: permissions, context, tool routing, and error recovery. The key shift is harness design, not longer prompts.

Why it matters: Strong HKR: the Claude Code teardown has a sharp counter-narrative and concrete 512k LOC plus 1.6%/98.4% split. It is not an official Anthropic release and lacks full reproduction details, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

OpenAI upgrades Codex to control Macs and run cross-app tasks

OpenAI upgraded Codex with Slack, Google Workspace, and Microsoft 365 integrations. Mike Russell tested Codex on a Mac across Adobe Audition, Photoshop, and Firefly, finishing in about 8 minutes with an 85–90 score. The key shift is OS-level computer control, not code completion.

Why it matters: All HKR axes pass: OpenAI Codex moves from coding into Mac-level control, with Slack, Google Workspace, and Microsoft 365 integrations. Single-source sourcing caps the score, but the 8-minute test and OS-agent angle justify P1.

Latent Space

[AINews] Agents for Everything Else: Codex for Knowledge Work, Claude for Creative Work

OpenAI expanded Codex to non-coding work, with CUA reported 42% faster. The update connects Microsoft, Google, and Salesforce, covering docs, slides, spreadsheets, research, and planning. The key signal is GUI-agent productization, not one benchmark score.

Why it matters: HKR-H/K/R all pass: Codex moves into non-code GUI work, with a 42% speed claim and named integrations. Price, rollout scope, and reproduction details are not disclosed, so it stays below P1.

Hacker News front page

Show HN: Pu.sh – a full coding-agent harness in 400 lines of shell

Pu.sh ships a coding-agent harness in about 400 lines of shell, using only sh, curl, and awk. It supports Anthropic and OpenAI, 7 tools, REPL, auto-compaction, checkpoint/resume, pipe mode, and 90 no-API tests. It excludes TUI, streaming, images, OAuth, and Windows.

Why it matters: HKR-H/K/R all pass, but this is a small Show HN open-source tool, not a model or platform release. HN frontpage plus a reproducible 400-line implementation clears the featured bar.

r/LocalLLaMA

Long-context coding on RTX 5080 16GB: Qwen3.6-35B-A3B holds 30 t/s at 128K

A Reddit user tested a local coding-agent setup on RTX 5080 16GB; the title says Qwen3.6-35B-A3B reaches 30 t/s at 128K. The post lists Ryzen 9700X, 96GB DDR5, Windows 11, and CUDA 12.9.1 as required. Qwen3.6-27B dense hit only 3.2 t/s at 128K, so the key path is KV quantization plus MoE offload.

Why it matters: HKR-H/K/R all pass: 30 t/s at 128K on a 16GB RTX 5080 is a strong hook, with hardware/CUDA details and a dense baseline. Single Reddit run lacks multi-source reproduction, so featured not P1.

Apr 30Thursday

r/LocalLLaMA

My calculator is a transformer

radarsat1 shows an RPN interpreter compiled into Transformer weights; “2 3 + 2 *” returns 10. The residual stream acts as registers, attention weights are compiler-calculated, while nonlinear MLP logic is still trained. The prototype is 1.1 GB; the key point is calculable attention weights, not a practical calculator.

Why it matters: HKR-H comes from the counterintuitive title; HKR-K has a reproducible input, weight-construction mechanism, and 1.1GB figure. HKR-R is real but niche, so this stays just above featured threshold, below 78.

The Verge · AI

OpenAI talks about not talking about goblins

OpenAI explained instructions telling its coding model to avoid goblins and similar creatures after Wired reported them. OpenAI says GPT-5.1’s “Nerdy” personality began using creature metaphors; the post does not disclose the full fix.

Why it matters: HKR-H/K/R all pass: the goblins prompt is unusual, OpenAI names the GPT-5.1 Nerdy persona behavior, and coders care about hidden prompt reliability. No full fix mechanism is disclosed, so it stays in the low featured band.

Ben's Bites

Building Gets Easier

Ben’s Bites lists agent tooling updates from Cloudflare, Stripe, Cursor SDK and others, with over 10 product leads. Cloudflare lets agents create accounts, buy domains, get API tokens and deploy; Stripe adds Agentic Commerce Suite, Link CLI and agent-ready Treasury accounts. The key shift is external permissions becoming agent-readable interfaces.

Why it matters: HKR-H/K/R pass, but this is a roundup rather than one major launch. Concrete Cloudflare and Stripe agent-permission details keep it in the featured-low band.

r/LocalLLaMA

Actual comparison between locally run Qwen-3.6-27B and proprietary models

The author compared 5 model setups on an autoresearch-loop task; only Qwen-3.6-27B via OpenRouter nearly solved it. The local q4_k_m run took about 8 hours and used 39k/45k tokens; full-quality Qwen used 4.4M tokens and cost $0.939. The useful signal is failure quality: both Qwen runs needed small fixes, while Gemma, Codex-Spark, and Claude Haiku 4.5 missed tests or key logic.

Why it matters: HKR-H/K/R all pass: the post has a concrete agent-test surprise, token and cost data, and local-vs-proprietary tension. Single Reddit run limits source authority, so it stays in the lower featured band.

r/LocalLLaMA

Notes on what actually breaks when you run a coding agent on small local models

A Reddit user tested small local and free-tier cloud models for weeks on multi-file coding tasks. Sub-7B structured output was unreliable; failures included markdown fences, wrong-file edits, and read/write misclassification, with post-processing and validation as fixes.

Why it matters: HKR-H/K/R pass: the post names real local coding-agent failure points, a sub-7B threshold, four failure classes, and mitigations. Reddit single-post scope keeps it below release-tier news, so 75.

Xinzhiyuan · WeChat

AI Raw Proofs Pile Up on GitHub as Terence Tao Says Solving Alone Is Not Enough

Terence Tao says math is shifting from proof scarcity to proof abundance, with 20-plus AI solutions pending assessment on an Erdős problems GitHub page. The post says GPT-5.4 Pro generated an Erdős #1196 approach in 80 minutes, and Tao verified the core within 24 hours. The key issue is verification and digestion workflow, not raw proof count.

Why it matters: All HKR axes pass: Tao plus GitHub proof backlog gives HKR-H, while 20+ pending AI solutions and an 80-minute GPT-5.4 Pro claim give HKR-K. This is not a model release, so it stays below 85.

Synced · WeChat

Alec Radford tests Hassabis’s AGI challenge with a model trained on pre-1931 data

Alec Radford’s team trained 13B talkie on 260B English tokens dated before 1931. They tested surprise on nearly 5,000 historical events and used HumanEval for lower-contamination code evaluation. The key issue is time leakage: the 13B model still has vague post-WWII knowledge.

Why it matters: HKR-H/K/R all pass: the 1930 cutoff is a sharp hook, the post gives 260B tokens and ~5,000 event tests, and the finding targets data leakage. This is strong research, not a model or platform release, so 78–84 fits.

Latent Space

[AINews] The Inference Inflection

Latent Space argues inference demand has hit an inflection point, citing its Apr 28-29, 2026 AINews roundup. Jensen Huang is quoted saying per-task compute rose about 10,000x in two years, with usage up about 100x. The key watchpoints are CPU sandboxes, agent harnesses, and split inference workloads.

Why it matters: HKR-H/K/R all pass, but this is a Latent Space AINews roundup and trend read, not a model launch or major product release. It fits the upper featured-threshold band for insightful commentary.

r/LocalLLaMA

inclusionAI/Ling-2.6-1T · Hugging Face

inclusionAI open-sourced Ling-2.6-1T on Hugging Face, with 1 trillion parameters. It uses MLA plus Linear Attention and Contextual Process Redundancy Suppression to reduce CoT overhead. The post cites AIME26 and SWE-bench Verified but does not disclose scores.

Why it matters: HKR-H/K/R all pass, but benchmark scores for AIME26 and SWE-bench Verified are not disclosed. A 1T open model with a named architecture mechanism fits featured, not P1.

Apr 29Wednesday

X · @dotey

OpenAI Expands AWS Partnership, Bringing GPT-5.5, Codex, and Managed Agents to Bedrock

OpenAI expanded its AWS partnership, bringing GPT-5.5, Codex, and Managed Agents to Amazon Bedrock in limited preview. Codex supports Bedrock across CLI, desktop, and VS Code, with over 4M weekly active users. The key detail is reuse of AWS compliance, billing, and cloud commitments.

Why it matters: HKR-H/K/R all pass: this is more than a routine cloud listing, with OpenAI bringing GPT-5.5, Codex, and Managed Agents to AWS Bedrock. Limited preview, 4M weekly Codex users, and IDE/CLI entry points justify same-day coverage.

X · @dotey

AI terminal tool Warp open-sources client code with OpenAI as founding sponsor

Warp open-sourced its client code under AGPL; only the client is open, while server code stays closed. The Rust terminal has 700,000+ developers, and its Oz cloud AI handles coding, planning, and tests. The key signal is its AI-first contribution workflow.

Why it matters: HKR-H/K/R all pass: the OpenAI sponsorship hook, AGPL/client-only detail, and 700K-developer signal are concrete. This is a strong dev-tool open-source update, not a major model or capability release.

Apr 28Tuesday

Ben's Bites

Builders

Ben’s Bites published one newsletter on AI builders. It says OpenAI released GPT-5.5 at 2x GPT-5.4 pricing, with a claimed 40% token-efficiency gain. Claude Managed Agents memory entered public beta, and Cursor’s SpaceX/xAI deal includes a $60B 2026 purchase option.

Why it matters: HKR-H/K/R all pass: GPT-5.5 cost/efficiency figures, Claude Managed Agents Memory beta, and a Cursor deal term. It stays in 85–94 because this is a newsletter roundup, not a primary release.

QbitAI · WeChat

Xiaomi open-sources MiMo-V2.5 series; Pro builds a macOS-like desktop in 4 hours

Xiaomi open-sourced MiMo-V2.5 weights, covering Pro Agent, multimodal base, TTS, and ASR models. MiMo-V2.5-Pro built a 54-app macOS-like desktop in 4 hours without human takeover; it scored 233/233 on SysY with 672 tool calls in 4.3 hours. Key details for practitioners are the 1M context, 100T-token program, and free Agent-framework access.

Why it matters: HKR-H/K/R all pass: Xiaomi open-sourced MiMo-V2.5 weights with concrete agent and coding-task numbers. Domestic flagship model release bump puts it in the must-write same-day band.

Hacker News front page

Xiaomi releases MiMo-v2.5 weights with strong coding and agent benchmarks

Xiaomi released MiMo-v2.5 family weights; the title cites strong coding and agent benchmarks. The RSS body only lists URLs, 13 HN points and 2 comments; the post does not disclose size, license, or scores.

Why it matters: HKR-H/K/R pass because a Xiaomi coding/agent weights release is concrete and practitioner-relevant. Sparse sourcing holds it near the featured floor: no parameters, license, or benchmark numbers are disclosed.

The Verge · AI

Attack of the Killer Script Kiddies

The Verge discusses Claude Mythos and AI bug finding, citing DARPA AIxCC scans over 54 million code lines. Teams found most seeded flaws plus over a dozen unseeded bugs; the RSS snippet does not disclose Mythos benchmarks, pricing, or access terms.

Why it matters: HKR-H/K/R all pass: the hook is strong, DARPA AIxCC supplies concrete numbers, and the security angle resonates. No Claude Mythos benchmark, pricing, or access terms are disclosed, so it stays in the featured-threshold band.

Hacker News front page

GitHub Copilot code review will start consuming GitHub Actions minutes

GitHub will make Copilot code reviews consume GitHub Actions minutes starting June 1, 2026. Private-repo reviews use plan entitlements, with overages billed at standard Actions rates; public repos stay free. The change covers Copilot Pro, Pro+, Business, and Enterprise, including direct org billing for unlicensed users.

Why it matters: Official GitHub billing change for Copilot code review hits CI quotas and org invoices; HKR-H/K/R all pass, but it is a pricing rule, not a capability release, so it sits low in 72–77.

Xinzhiyuan · WeChat

Claude bans hit 110-person firm; Cursor incident deletes database in 9 seconds

Anthropic allegedly suspended 110 Claude accounts at a US agtech firm, while API billing continued. The post says appeals went unanswered for 36 hours, and PocketOS says Claude Opus 4.6 via Cursor deleted production data and volume backups in 9 seconds. The key issue is access control: no RBAC, no environment isolation, and no delete confirmation.

Why it matters: HKR-H/K/R all pass: the incident has a strong hook and concrete details: 110 accounts, 36 hours, 9 seconds, and no RBAC. Kept at 82 because it is still a single-source allegation without an Anthropic postmortem.

X · @op7418

Xiaomi open-sources the MiMo-V2.5 model series

Xiaomi open-sourced the MiMo-V2.5 model series under the MIT license for commercial use, retraining, and fine-tuning. It also launched Orbit 100T Token, offering approved AI builders up to 1.6B credits worth 659 yuan. Agent framework teams can apply for free MiMo token access; the post does not disclose model size or benchmark results.

Why it matters: HKR-H/K/R all pass: Xiaomi MiMo-V2.5 open source, MIT terms, and Orbit 100T credits matter to builders. Missing params and benchmarks keep it in the 78–84 band, below P1.

r/LocalLLaMA

Local coding models have reached a threshold for real work

Antigma tested 27B–32B open-weight models; Qwen 3.6-27B scored 38.2% on Terminal-Bench 2.0. The run used 89 tasks and the default per-task timeout, while verified SOTA is about 80%. The key claim is deployment lag: offline coding is about 6–8 months behind hosted frontier models.

Why it matters: HKR-H/K/R all pass: the post gives a real-work threshold claim, a 38.2%/89-task Terminal-Bench result, and a 6–8 month offline gap. Reddit single-post sourcing keeps it in the low featured band.

Hacker News front page

Claude Pro: Opus Requires Extra Usage in Claude Code

Anthropic lists 6 Claude Code models, and Pro users need extra usage enabled and purchased to use Opus. The guide gives 3 configuration paths: /model, --model, and ANTHROPIC_MODEL in zsh or bash. The post does not disclose extra usage pricing or quotas.

Why it matters: HKR-H/K/R all pass, but the facts come from a help doc and cover Claude Code access/configuration, not a new model or major capability. Anthropic relevance lifts it to the lower featured band.

X · @dotey

The West forgot how to build things, and may forget how to write code

Denis Stetskov compares Western defense production gaps with AI coding, citing Stinger orders placed in 2022 for 2026 delivery. He says Europe’s 1M-shell target was 9 months late, and METR found senior developers 19% slower with AI. The key risk is the junior-engineer pipeline, not code generation speed.

Why it matters: HKR-H/K/R all pass: the analogy is clickable, the post gives concrete defense and METR numbers, and the junior-engineer pipeline resonates. X translation/commentary limits authority, so it sits just above the featured threshold.

X · @dotey

GitHub Copilot switches to usage-based billing on June 1

GitHub Copilot will switch to AI Credits billing on June 1 while keeping subscription prices unchanged. Credits count input, output, and cached tokens; Pro includes $10 monthly credits and Pro+ includes $39. Watch Copilot Agent long-task costs.

Why it matters: HKR-H/K/R all pass: Copilot billing moves from subscription expectations to token/cache consumption with date and credit amounts. Single-source X context lacks enterprise details and overage rates, so it stays in the 78–84 band.

X · @dotey

Cursor 3 feedback: users want a reliable AI development workspace

Eric Zakariasson’s Cursor 3 feedback thread summarizes 431 replies, with users asking for a stable AI development workspace. Requests center on Agent Window retaining LSP, debugging, Git, terminal and diff workflows, plus multi-agent worktrees and model-cost transparency. The key issue is workflow reliability, not a flashier IDE.

Why it matters: All HKR axes pass: 431 user replies, concrete workflow requests, and strong resonance for Cursor users. Kept in the low featured band because this is feedback synthesis, not an official Cursor release or roadmap.

Hacker News front page

GitHub Copilot is moving to usage-based billing

GitHub said on 2026-04-27 that GitHub Copilot will move to usage-based billing. The captured post only shows the title, time, and navigation. It does not disclose the launch date, usage metric, prices, or overage rules.

Why it matters: GitHub Copilot billing affects a large developer base. HKR-H and HKR-R are strong, while HKR-K is limited to the usage-based mechanism with no date, metering unit, or price details disclosed.

Apr 27Monday

Dwarkesh Patel podcast

What I've been Thinking About This Weekend: Open Questions, Intelligence vs Power, Verification in Science

Dwarkesh lists open AI questions, including that five hyperscalers own over 70% of global AI compute. He asks about coding agents, KV cache costs, merging training with inference, and online learning; the post gives questions, not experimental answers.

Why it matters: HKR-H/K/R all pass: Dwarkesh adds a concrete compute-concentration claim and practitioner-relevant questions. No experiment, release, or policy change, so it stays in the 72–77 commentary band.

Hacker News front page

Running Local LLMs Offline on a Ten-Hour Flight

Dmitri Lerko ran Gemma 4 31B and Qwen 4.6 36B locally during a 10-hour flight with no Wi‑Fi. The MacBook Pro M5 Max had 128GB unified memory and a 40-core GPU; sustained load used about 1% battery per minute, and performance degraded past 100k tokens. The sharp finding is instrumentation: an iPhone cable delivered 60W, while a MacBook cable delivered 94W under the same load.

Why it matters: HKR-H/K/R all pass: this is a named first-person local-inference test with concrete hardware, model, battery, and power numbers. Scope stays practical rather than industry-shaking, so it lands in the 72–77 band.

Hacker News front page

Show HN: OSS Agent Dirac topped TerminalBench on Gemini-3-flash-preview

Dirac-run released Dirac and says it topped TerminalBench using Gemini-3-flash-preview. The repo claims 50-80% lower API costs via Hash Anchored edits, parallel operations, and AST manipulation; the post does not disclose full scores.

Why it matters: HKR-H/K/R all pass: an OSS coding agent claims a TerminalBench lead with cost and mechanism details. Held to 78 because the post relies on repo claims and lacks full leaderboard scores or reproduction logs.

Hacker News front page

AI can cost more than human workers now

Axios says some firms now spend more on AI than salaries; Nvidia's Bryan Catanzaro says compute costs exceed employee costs. Gartner forecasts 2026 IT spending at $6.31T, up 13.5%, driven by AI infrastructure, software, and cloud. Watch token costs: Uber's CTO has already exhausted the 2026 AI budget.

Why it matters: HKR-H/K/R all pass: the piece turns AI cost anxiety into budget facts, including Nvidia compute costs and Uber’s token-budget issue. It stays in the 72–77 band because this is trend reporting, not a launch or hard news event.