Skip to content

#编码

10 today

May 17Sunday

Synced · WeChat

Peter Steinberger Says His Monthly Token Bill Hit $1.3M, Covered by OpenAI

Peter Steinberger used 603 billion tokens across 7.6 million requests in 30 days, with the bill exceeding $1.3 million; he said disabling fast mode cut the price by 70%, and OpenAI does not charge him for the tokens.

Why it matters: HKR-H/K/R all pass: the story has a sharp cost hook, concrete usage numbers, and strong practitioner resonance. It is a first-person bill disclosure, not an OpenAI pricing or product launch, so it sits just above the featured threshold.

AI HOT (Curated Pool)

Anthropic CEO discusses AI’s dual impact: high growth and high unemployment

Dario Amodei said AI may drive 5%-10% GDP growth while increasing unemployment and inequality, and near-free software costs would challenge the assumptions behind traditional software business models.

Why it matters: HKR-H/K/R all pass: Dario Amodei’s 5%-10% GDP and near-free software claims are concrete and highly discussable. The source is an X summary, not a full primary transcript, so it stays at 78.

Computing Life · Share · Yage

Vibe Coding’s Security Crisis

AI coding platforms exposed sensitive data from thousands of enterprise applications through public-by-default deployment settings; the snippet names hospital schedules, bank financial data, and clinical trial data, and identifies one-click deployment defaults rather than AI-generated code as the core mechanism.

Why it matters: HKR-H/K/R pass: the public-by-default deployment angle is clickable, concrete, and practitioner-relevant. Lack of named platform detail or top-tier sourcing keeps it in the lower good-quality band.

AI HOT (Curated Pool)

MagicPath Integrates with Codex to Combine Design and Development

MagicPath AI CEO @skirano demonstrated MagicPath running inside Codex as a native canvas, with users configuring it through one command, dragging UI elements, and letting Codex generate and edit code in real time.

Why it matters: HKR-H/K/R pass: MagicPath puts a draggable design canvas inside Codex with one-command setup and live code edits. Single-demo sourcing and missing framework support, permissions, and reproducible cases keep it at the lower featured band.

AI HOT (Curated Pool)

Eric Jang shares lessons from building AlphaGo from scratch

Eric Jang spent several months implementing AlphaGo from scratch and says that in 2026, training a strong Go AI requires only a few thousand dollars in rented compute rather than DeepMind-scale resources.

Why it matters: All three HKR axes pass: the hook is a from-scratch AlphaGo rebuild, and K has concrete claims on months of work and few-thousand-dollar compute. It stays in 78-84 because this is a social post, not a model release or full paper.

May 16Saturday

TechCrunch · AI

OpenAI co-founder Greg Brockman takes charge of product strategy

Greg Brockman has officially taken charge of OpenAI’s product strategy, and Wired reports that he described a plan in a staff memo to combine ChatGPT and Codex into one unified experience.

Why it matters: HKR-H/K/R all pass: OpenAI co-founder product control plus a reported ChatGPT-Codex unification matters. No launch date, feature boundary, or rollout plan is disclosed, so this stays below a major product release.

AI HOT (Curated Pool)

Anthropic Founder’s Playbook warns AI can raise startup failure rates

Anthropic published Founder’s Playbook, arguing that AI tools such as Claude Code reduce prototyping cost but increase startup failure risk across the Idea, MVP, Launch, and Scale stages through false validation, confirmation bias, agentic technical debt, and founder decision bottlenecks.

Why it matters: HKR-H/K/R pass: the Anthropic founder playbook has a sharp counterintuitive angle, a four-stage mechanism, and clear founder resonance. It stays near the featured floor because no dataset or reproducible test is disclosed.

AI HOT (Curated Pool)

Researchers use Anthropic Mythos to build a macOS kernel exploit bypassing Apple M5 MIE

Three researchers used Anthropic Mythos to develop a macOS kernel exploit in six days, moving from discovery on April 25 to completion on May 1, bypassing Apple’s MIE memory-integrity system for M5 and A19 chips and gaining root via standard unprivileged system calls; the full technical report will follow Apple’s patch.

Why it matters: HKR-H/K/R all pass: Anthropic Mythos, a 6-day macOS kernel exploit, and M5/A19 MIE bypass create real dual-use signal. Kernel-exploit depth and single X-source sourcing keep it below the 85 must-write band.

AI HOT (Curated Pool)

Codex adds multi-device remote control and shared context

Codex controls multiple devices through ChatGPT, switches by project to access each device’s context and files, and supports remote SSH setup for other VMs.

Why it matters: HKR-H/K/R all pass, but the item is a thin X-post summary with no official release note, pricing, permission model, or reproducible demo. Treat it as a mid-weight coding-agent product update at the featured threshold.

r/LocalLLaMA

Qwen3.6-35B-A3B and 9B land on the public Terminal-Bench 2.0 leaderboard

little-coder × Qwen3.6-35B-A3B scored 24.6% ±3.2 on Terminal-Bench 2.0, above Gemini 2.5 Pro on Gemini CLI at 19.6% and Qwen3-Coder-480B on Terminus 2 at 23.9%.

Why it matters: HKR-H/K/R all pass, but this is a Reddit post with leaderboard numbers only; test setup and reproducibility details are not disclosed. Strong code-agent benchmark signal, not a 78+ release story.

AI HOT (Curated Pool)

OpenAI Restructures as Brockman Takes Over Product Strategy

OpenAI merged ChatGPT, Codex, and API into one product organization, with Greg Brockman taking over product strategy; the post says Anthropic’s valuation reached $900 billion, but it does not disclose the restructuring timeline.

Why it matters: HKR-H/K/R all pass: this is an OpenAI top-level product reorg covering ChatGPT, Codex, and API. Single-source summary keeps it below the highest band, but it is same-day must-write news.

QbitAI · WeChat

Codex Integrates HeyGen for Prompt-Based Video Generation and Editing

Codex integrates the HeyGen plugin to run image generation, talking-avatar video, subtitles, and edits from natural-language prompts; the article tests roughly one-minute avatar generation, trimming content after 10 seconds, and deleting a blink at the eighth second.

Why it matters: HKR-H/K/R all pass, backed by a numbered hands-on test. The scope is still one Codex-to-HeyGen plugin workflow, not a model or platform release, so it lands in the 72-77 featured band.

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

AI HOT (Curated Pool)

Ignoring Token Costs, Using 100 AI Instances to Automate an Open Source Project

The OpenClaw team runs about 100 Codex instances to handle code review, security analysis, issue deduplication, test reproduction, task creation from meetings, spam filtering, and performance regression monitoring.

Why it matters: HKR-H/K/R all pass: 100 Codex instances running open-source maintenance is a strong operational anecdote with concrete task types. Single X post, no cost, outcome metrics, or reproducible setup, so it stays in the lower featured band.

The Verge · AI

OpenAI keeps shuffling executives to win the AI agent battle

OpenAI announced a reorganization Friday that makes Greg Brockman the official lead for product, and its memo says the company will combine ChatGPT and Codex into one unified agentic experience.

Why it matters: HKR-H/K/R all pass: the power shuffle is clickable, the product merge is new, and OpenAI's agent roadmap matters. It stays at 83 because no capability shipped and timing, pricing, and technical details are absent.

May 15Friday

Alibaba Technology · WeChat

Qoder 1.0 launches as an agentic development workspace beyond AI IDE

Alibaba released Qoder 1.0 with downloads for Windows, macOS, and Linux, adding a standalone Quest workspace, cross-project parallel agent tasks, a team knowledge engine, and Experts mode with five roles for planning, research, coding, review, and testing.

Why it matters: Alibaba’s Qoder 1.0 is a mid-weight AI coding product release with concrete agent-workflow features and developer resonance. No pricing, benchmark, or task-success data is disclosed, so it stays near the featured threshold.

r/LocalLLaMA

Used over a million tokens in three sessions to test Qwen 3.6 35B MTP

A Reddit user tested Qwen3.6-35B-A3B MTP across three million-token-scale sessions, using 300k context and KV Q8_0, and reported about 1.5x the tok/sec of earlier tests.

Why it matters: HKR-H/K/R all pass: the million-token test is clickable, 300k context and KV Q8_0 add testable detail, and local speed maps to cost. Source is one Reddit post, so it stays below the high-importance band.

AI HOT (Curated Pool)

Codex lands on mobile with preview in the ChatGPT app

OpenAI brought Codex to a preview inside the ChatGPT mobile app; the post does not disclose supported platforms, feature scope, pricing, or rollout schedule.

Why it matters: Official OpenAI product update with HKR-H/K/R, but detail is thin. The post does not disclose platform support, feature scope, pricing, or rollout timing, so it sits at the featured threshold.

QbitAI · WeChat

Understand LeCun’s JEPA World Model in 160 Lines of Code

A developer released the keon/jepa teaching repository with five JEPA variants implemented as standalone PyTorch files, ranging from 160 to 278 lines, depending only on PyTorch and torchvision; the post reports iJEPA runs on CIFAR-10 for 100 epochs and reaches 52.7% linear-probe accuracy, while V-JEPA, C-JEPA, and LeWorldModel use toy or synthetic datasets.

Why it matters: HKR-H/K/R pass via the 160-line JEPA hook, reproducible repo, and non-LLM world-model angle. It is a tutorial artifact, not a model or paper release, so it sits at the featured threshold.

AI HOT (Curated Pool)

Anthropic's Mythos AI helped find and exploit two unknown macOS kernel vulnerabilities in five days

Anthropic’s Mythos AI helped researchers find two previously unknown macOS kernel vulnerabilities in five days and chain them into a privilege-escalation exploit that bypassed Apple’s memory integrity protection, according to the Wall Street Journal snippet.

Why it matters: HKR-H/K/R all pass, and Anthropic-linked AI security work is high-signal. The score stays in 78–84 because the source is a social post and lacks paper details, reproducible conditions, or exploit mechanics.

AI HOT (Curated Pool)

Claude Agent Tool v2.1.142 Release

Claude Agent Tool v2.1.142 adds eight command-line flags for configuring background sessions, upgrades Fast mode’s default model to Opus 4.7, and fixes more than 15 issues including MCP tool timeouts and Windows network-drive deadlocks.

Why it matters: HKR-H/K/R all pass: this is a small Claude Code release, but the Opus 4.7 Fast-mode default, 8 session flags, and 15+ fixes affect daily dev workflows. Anthropic tool-chain relevance keeps it at the featured floor.

r/LocalLLaMA

I Let a Small Model Train on Its Own Mistakes; It Reached 80% on HumanEval and Beat GPT-3.5 on Math

The author fine-tuned Qwen 2.5 7B base on self-mined mistake-correction pairs, raising HumanEval from 25/164 to 112/164; Qwen 2.5 14B used 100 pairs and a 95-minute H100 run costing $3.50.

Why it matters: HKR-H/K/R pass: the hook is strong and the post gives samples, H100 time, cost, and HumanEval deltas. Kept at 78 because it is a single Reddit post and the 80% claim differs from 112/164.

AI HOT (Curated Pool)

Codex adds automation hooks and programmatic tokens

Codex added hooks and programmatic access tokens: hooks run scripts at key task stages for validation, secret scanning, logging, or repo-specific behavior, while scoped tokens for Business and Enterprise teams support CI/CD, release workflows, and internal automation with expiration or revocation.

Why it matters: HKR-H/K/R all pass: Codex gains concrete automation hooks and programmatic tokens for CI/CD. Score stays in the 72–77 band because the post discloses workflow fit, not pricing, permission detail, or impact data.

Bloomberg Technology

Musk’s xAI Unveils First Coding Agent in Bid to Rival Anthropic

xAI is rolling out its first AI coding agent, Grok Build, for software development workflows; the RSS snippet names Anthropic’s Claude as the rival but does not disclose pricing, availability, benchmarks, or supported IDEs.

Why it matters: HKR-H and HKR-R pass: xAI entering coding agents is a strong competitive hook for developers. HKR-K fails because pricing, availability, and benchmarks are not disclosed, so this stays at the low end of a mid-weight product update.

The Verge · AI

OpenAI’s Codex is now in the ChatGPT mobile app

OpenAI will let users access Codex from the ChatGPT mobile app; the RSS snippet says Codex can write code and use apps on a computer, but the post does not disclose launch timing, pricing, or the full mobile feature scope.

Why it matters: OpenAI added Codex access to ChatGPT mobile, a mid-weight product update. HKR-H/K/R pass through the mobile coding-agent hook, concrete app-control claim, and developer workflow nerve; missing timing, pricing, and support scope keep it at the featured floor.

The Verge · AI

Microsoft starts canceling Claude Code licenses

Microsoft plans to remove most Claude Code licenses and push many developers toward Copilot CLI; the snippet says Microsoft opened access in December to thousands of internal developers, but the post does not disclose the exact license count, pricing, or migration schedule.

Why it matters: HKR-H comes from Microsoft dropping a rival coding tool; HKR-K adds the Dec rollout to thousands of internal devs; HKR-R hits Claude Code vs. Copilot competition. Strong featured, not a major release.

AI HOT (Curated Pool)

The Founder's Playbook: Building an AI-Native Startup

Anthropic published an AI-native startup playbook covering four stages—ideation, MVP, launch, and scaling—with goals, exit criteria, failure modes, and Claude-based exercises for validation, customer discovery, technical debt control, product-market fit checks, and workflow automation.

Why it matters: HKR-H/K/R all pass, but this is an Anthropic playbook rather than a model or product capability release. The concrete value is the 4-stage framework, exit criteria, and Claude-driven exercises, so it lands at the featured floor.

AI HOT (Curated Pool)

Using Claude Code Effectively in Large Codebases: Best Practices and Where to Start

Claude Code is used in million-line monorepos, legacy systems, and distributed architectures, and the post says its large-codebase workflow relies on five extension points: CLAUDE.md, hooks, skills, plugins, and MCP servers for agentic search on local codebases.

Why it matters: HKR-H/K/R pass: official Claude Code guidance, five concrete extension points, and a direct coding-agent workflow nerve. It is a high-quality tutorial, not a new model or major capability release, so it stays in the 72–77 band.

May 14Thursday

AI HOT (Curated Pool)

Use Codex from anywhere

OpenAI added Codex to the ChatGPT mobile app, letting users monitor, guide, and approve remote coding tasks across devices.

Why it matters: OpenAI added Codex controls to ChatGPT mobile for monitoring, guiding, and approving remote coding tasks. This clears HKR-H/K/R as a mid-weight product update, but pricing, permission details, and task limits are not disclosed, so it stays below a major release.

AI HOT (Curated Pool)

MiMo V2.5 Pro Places Third on DesignArena

MiMo V2.5 Pro placed third on the DesignArena overall leaderboard; its Thinking version rose 8 spots over MiMo-V2.5 and matched Claude Sonnet 4.6 performance on frontend coding tasks.

Why it matters: HKR-H/K/R all pass, but the facts come from one official X post with no methodology, access, or pricing. This fits a mid-weight benchmark/product update, not a same-day must-write.

Xinzhiyuan · WeChat

Anthropic Overtakes OpenAI in Enterprise AI Adoption After Three Years

Ramp says Anthropic reached 34.4% enterprise adoption, surpassing OpenAI at 32.3% for the first time; the index is based on credit-card and invoice spending from more than 50,000 companies.

Why it matters: HKR-H/K/R all pass: a reversal hook, concrete 34.4%/32.3% figures, and a strong enterprise-AI rivalry angle. Score stays at 80 because Ramp spending data is not global market share.

QbitAI · WeChat

Chinese GPU Vendor Hosts Open Source Meetup With SGLang Core Developers

Moore Threads said at the SGLang × MUSA Meetup that the MUSA backend has been merged into SGLang mainline, with 47 PRs submitted and 41 merged as of May 12.

Why it matters: HKR-H/K/R all pass, but this is an inference-backend ecosystem update rather than a model launch or platform shift. The 47 PRs and 41 merges make it concrete enough for featured, not P1.

Latent Space

[AINews] Codex Rises, Claude Meters Programmatic Usage

Anthropic changed paid Claude plans to include monthly API credits equal to the subscription price, so a $200 plan includes $200 for programmatic usage outside Anthropic-owned harnesses, while OpenAI promoted Codex enterprise switching incentives in the same news cycle.

Why it matters: HKR-H/K/R all pass: the story ties Claude metering to Codex competition and gives a concrete $200 credit detail. This is a meaningful developer-cost update, not a major model or capability launch, so it sits in mid featured.

AI HOT (Curated Pool)

Moonshot AI founder Yang Zhilin releases a 40-minute video

Yang Zhilin explains Kimi K2 training in a 40-minute video, saying the model cost $4.6 million and beat GPT-5.5 and other competitors on coding tasks.

Why it matters: HKR-H/K/R all pass: the founder-led Kimi K2 training breakdown adds a $4.6M cost figure and GPT-5.5 coding comparison. Single-source X relay and missing benchmark names keep it in 78-84, not P1.

AI HOT (Curated Pool)

xAI launches early beta of Grok Build

xAI launched an early beta of Grok Build for SuperGrok Heavy subscribers, offering a terminal-based coding agent with plan review, parallel subagents for large tasks, and a headless mode for scripting and automation.

Why it matters: HKR-H/K/R all pass: xAI enters terminal coding agents with plan mode, parallel subagents, and headless mode. Early beta access for SuperGrok Heavy keeps it below the 85 same-day must-write band.

r/LocalLLaMA

2x RTX 3090 setup for local Qwen 3.6 27B inference

A Reddit user ran Qwen 3.6 27B on a dual RTX 3090 Ubuntu setup, reporting 48GB VRAM, a 262k context window, no NVLink, about 4000 pp/s prompt processing, and 113 tk/s generation.

Why it matters: All HKR axes pass, and this is a first-person local-inference run with concrete numbers. Source is a single Reddit post with limited reproducibility detail, so it sits at the low featured threshold.

AI HOT (Curated Pool)

Claude paid plans will offer monthly coding usage credits

Claude paid plans can claim monthly coding usage credits from June 15, covering Claude Agent SDK, claude -p, Claude Code GitHub Actions, and third-party apps built on the Agent SDK.

Why it matters: HKR-H/K/R all pass: the update names a date, quota mechanism, and covered Claude coding surfaces. Importance stays in the low featured band because this is a billing/access change, not a model release.

May 13Wednesday

AI HOT (Curated Pool)

Configuring Development Environments for Agents

Cursor released tools for cloud agent development environments, adding multi-repository support, Dockerfile-based configuration, audit logs, and environment-level network and secret controls; the post says cache hits improve build speed by 70%.

Why it matters: HKR-K and HKR-R pass: Cursor adds concrete cloud-agent environment controls, including Dockerfile setup, audit logs, permissions, and 70% faster cached builds. HKR-H is weaker, so this sits at the lower featured band.

OpenAI News

Building a Safe, Effective Sandbox for Codex on Windows

OpenAI built a secure sandbox for Codex on Windows. The RSS snippet discloses controlled file access and network restrictions, but the post does not disclose implementation details, performance data, or rollout conditions.

Why it matters: OpenAI details a Windows sandbox for Codex with file-access and network controls. It is not a major model release, but HKR-H/K/R all pass because the safety boundary matters for coding-agent adoption.

AI HOT (Curated Pool)

Miaoda App and Enterprise Edition launch with 90% self-generated code

Baidu launched the Miaoda app and Miaoda Enterprise Edition, saying 90% of the Miaoda app’s code was generated by Miaoda itself; Miaoda-generated apps have served over 10 million users and reached a total value of RMB 5 billion.

Why it matters: HKR-H/K/R all pass via the 90% dogfooding hook, concrete adoption/value figures, and coding-tool resonance. Company-source metrics lack independent context, so this stays in the lower featured band.