Skip to content

#编码

10 today

May 13Wednesday

QbitAI · WeChat

An 8-Year-Old Turns Ideas into Apps as Baidu Launches Miaoda 3.0

Baidu launched Miaoda 3.0 at its 2026 Create conference, adding iOS and Android app generation, Android packaging, online hot updates, and an enterprise edition with three-level permissions, environment isolation, and SLA commitments.

Why it matters: HKR-H/K/R pass: Baidu’s Miaoda 3.0 adds mobile app generation, Android packaging, hot updates, and enterprise controls. This is a solid product update, not a flagship model release or must-write event.

r/LocalLLaMA

The Trillion-Parameter Dilemma: MiMo-V2.5-Pro Open-Sourced at 1.02T Parameters

Xiaomi open-sourced MiMo-V2.5-Pro with 1.02T parameters, 42B active parameters, a 1M context window, and an MIT license; the author ran 125 Claude Code sessions through the API, spending $70.12 for 387,380,436 tokens with a 96.3% cache hit rate.

Why it matters: HKR-H/K/R all pass: a Xiaomi 1.02T open model plus a concrete Claude Code API cost experiment. Reddit sourcing keeps it at the low end of the 85+ band, but the domestic flagship-model signal clears p1.

New York Times Chinese

China Sought Access to Anthropic’s Latest Technology but Was Rejected

Chinese think-tank representatives asked Anthropic in Singapore last month to give Beijing access to Mythos, and Anthropic refused; the company has limited the vulnerability-finding model to the U.S. government and more than 40 organizations.

Why it matters: HKR-H/K/R all pass: the NYT report gives the Singapore request, Mythos’s bug-finding use, and its US-government-plus-40 access scope. This is a same-day security and US-China AI access story.

AI HOT (Curated Pool)

Claude Code adds /goal feature to keep tasks running until completion

Claude Code introduced a /goal feature that keeps Claude working until a task is completed; the post does not disclose the trigger mechanism, supported versions, pricing, or failure conditions.

Why it matters: HKR-H/K/R pass because /goal targets a real Claude Code reliability pain. It is a single-feature Anthropic update with sparse mechanics, so it lands at the lower featured band, not same-day major news.

r/LocalLLaMA

A real transformer language model running locally on a stock Game Boy Color

maddiedreese ran Andrej Karpathy’s TinyStories-260K on a stock Game Boy Color with INT8 weights, fixed-point math, an MBC5 ROM, bank-switched cartridge storage, and KV cache in cartridge SRAM; the demo uses no phone, PC, Wi‑Fi, link cable, or cloud inference, but output is extremely slow and gibberish.

Why it matters: HKR-H/K/R all pass: a named first-person experiment with concrete model and memory details. Impact stays low-featured because it is a Reddit hardware hack with slow, garbled output, not a usable product or model release.

AI HOT (Curated Pool)

90% of People Are Wasting Tokens

Andrej Karpathy says 90% of AI coding bills is wasted on unnecessary context, including repeated full-repository sends, expensive models for simple tasks, and missing prompt caching.

Why it matters: HKR-H/K/R all pass via the 90% claim, named waste mechanisms, and practitioner cost pain. It reaches featured, but stays at 72 because the post gives no billing sample or reproducible test.

AI HOT (Curated Pool)

Codex Enables Background Multitasking Across Apps

OpenAI Devs says computer use lets Codex click, type, and keep working across Mac apps in the background; the post does not disclose release timing, permission design, or availability scope.

Why it matters: HKR-H/K/R all pass: OpenAI Devs gives a concrete Codex mechanism for background Mac cross-app actions, with clear developer relevance. Release timing, permissions, and availability are missing, so it stays at the lower featured band.

AI HOT (Curated Pool)

How Anthropic's Cybersecurity Team Uses Claude Code to Build a Threat Detection Platform

Anthropic’s detection platform engineering team used Claude Code to build the CLUE threat detection and response platform, completing a proof of concept in one day and delivery in one week while reducing analyst log investigation from hours to minutes.

Why it matters: HKR-H/K/R all pass, but this is an Anthropic internal dogfooding case rather than a Claude Code capability launch. CLUE and the timing metrics keep it just above the featured threshold.

AI HOT (Curated Pool)

Claude Opus 4.7 Fast Mode Opens Research Preview

Claude Opus 4.7 Fast Mode is now available as a research preview in the API and Claude Code. The post does not disclose model parameters, pricing, rate limits, or a general availability date.

Why it matters: HKR-H/K/R pass because this is a Claude fast-mode preview in API and Claude Code, directly tied to developer latency and workflows. Thin disclosure on pricing, limits, parameters, and GA timing keeps it at the featured threshold, not 78+.

AI HOT (Curated Pool)

GitHub Copilot Individual Plans Add Flex Allotments and a New Max Plan

GitHub will update Copilot individual plans on June 1 by adding flex allotments to Pro and Pro+ and introducing a new Max plan; the post does not disclose pricing, quota limits, or the exact allocation rules in the provided snippet.

Why it matters: HKR-H/K/R all land lightly because Copilot plan quotas affect many developers. Missing price, caps, and allocation rules keep it at the low featured threshold, not a major capability update.

Hacker News front page

Show HN: Agentic Interface for Mainframes and COBOL

Hypercubic launched Hopper, an agentic development environment that combines a real TN3270 terminal, z/OS-aware panels for datasets, jobs, and spool output, and an AI agent; sensitive operations require approval, and the terminal remains visible during agent actions.

Why it matters: HKR-H/K/R all pass: the mainframe-agent angle is novel, with concrete TN3270, z/OS, and approval mechanics. Small-vendor Show HN status and missing customer/pricing/results data keep it at the featured floor.

TechCrunch · AI

Everything Google announced at its Android Show, from Googlebooks to vibe-coded widgets

Google announced AI-first Googlebooks laptops, more agentic Gemini features, vibe-coded Android widgets, Gemini in Chrome, and refreshed Android Auto ahead of I/O; the RSS snippet does not disclose specs, pricing, availability, or rollout timelines.

Why it matters: HKR-H/K/R all pass because Google bundled several Gemini/Android AI entry points with named product hooks. Missing parameters, pricing, rollout dates, and testable performance keeps it in the mid-weight product-update band.

AI HOT (Curated Pool)

Code w/ Claude SF 2026: Building on Exponential AI Growth

Anthropic expanded developer tooling at Code w/ Claude SF 2026: Claude Code rate limits doubled, Claude Opus API limits increased, and hosted agents on the Claude platform added four functions, including memory review, multi-agent delegation, output criteria, and webhooks.

Why it matters: Anthropic ships a substantive Claude Code update with concrete numbers and feature additions; HKR-H/K/R all pass. This is strong dev-tool news, not a flagship model release, so it fits the 78–84 band.

May 12Tuesday

AI HOT (Curated Pool)

Dungeons & Desktops: Building a Procedurally Generated Roguelike with GitHub Copilot CLI

A GitHub employee used GitHub Copilot CLI to build an extension that parses any codebase into one Roguelike-style dungeon layout, with procedural level generation used as the core mechanism for a creative coding and game prototyping demo.

Why it matters: HKR-H and HKR-K pass: an official GitHub tutorial has a novel demo and a clear mechanism. It is not a major Copilot capability release, and lacks production metrics, pricing, or benchmark data, so it sits at the tutorial-featured floor.

r/LocalLLaMA

Local LLM Autocomplete and Agentic Coding on a Single 16GB GPU + 64GB RAM

Reddit user grumd runs Qwen2.5-Coder-7B Q6 for autocomplete and Qwen3.6-35B-A3B Q8 for agentic coding on one RTX 5080 with RAM offloading; the post reports about 145k context, 56GB RAM used with other apps open, and Qwen3.6-35B-A3B speed of tg128 at 35.29 tokens/s.

Why it matters: HKR-H/K/R all pass: a named first-person local coding experiment with concrete model, quantization, context, and throughput data. Source is a single Reddit post without replication or comparisons, so it stays in the low featured band.

Hacker News front page

Show HN: Statewright – Visual State Machines for More Reliable AI Agents

Statewright uses a Rust state-machine engine to constrain Claude Code tool access, iterations, transitions, and guards; the post says 13–20B models improved consistently on real SWE-bench tasks, but it does not disclose benchmark scores, sample size, or the exact evaluation protocol.

Why it matters: HKR-H/K/R all pass: the state-machine constraint is a clear agent-reliability hook with a testable SWE-bench claim. Exact scores and reproduction details are not disclosed, so it stays just above the featured threshold.

QbitAI · WeChat

Markdown Is Fading? Karpathy Also Backs HTML

Anthropic engineer Thariq argued for using HTML instead of Markdown and gave 5 reasons; the post says HTML generation takes about 2 to 4 times longer than Markdown.

Why it matters: HKR-H/K/R all pass, but this is a developer format debate rather than a model or product launch. Named Anthropic/Karpathy context and the 2-4x time figure clear the featured threshold at the low end.

AI HOT (Curated Pool)

Large npm Supply-Chain Attack Hits TanStack, Mistral AI, UiPath, and Others

Socket identified the Mini Shai-Hulud supply-chain attack, where attackers used three GitHub Actions flaws to publish nearly 373 malicious versions across more than 160 npm package names, affecting projects including TanStack, Mistral AI, and UiPath and stealing AWS, GCP, Kubernetes, GitHub tokens, and SSH private keys during installation.

Why it matters: HKR-H/K/R all pass: named projects create the hook, Socket provides concrete counts and mechanisms, and credential theft matters to AI engineering teams. It is a strong security incident, not a core model or product release, so it stays in the 78–84 band.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export app in about 5 minutes and run multivariate regression, finding that the last post-dinner AI usage time correlated negatively with sleep duration; after avoiding AI at night, average sleep increased by 1 hour and 40 minutes.

Why it matters: HKR-H/K/R all pass: a first-person quantified experiment links post-dinner AI use to shorter sleep, then reports +1h40m after stopping. Personal-blog scope keeps it below major industry-update territory.

AI HOT (Curated Pool)

What Parameter Golf Taught Us About AI-Assisted Research

OpenAI’s Parameter Golf brought together over 1,000 participants and more than 2,000 submissions to test AI-assisted machine learning research, coding agents, model quantization, and model design under strict parameter constraints.

Why it matters: OpenAI’s Parameter Golf recap clears HKR-H/K/R with a concrete contest, 1,000+ participants, and 2,000+ submissions. It is research/benchmark signal, not a model or product launch, so 78 fits the lower featured band.

The Verge · AI

OpenAI just released its answer to Claude Mythos

OpenAI launched Daybreak, a security initiative that uses the Codex Security AI agent released in March to model an organization’s code, validate likely vulnerabilities, and automate detection of higher-risk issues before attackers find them.

Why it matters: HKR-H/K/R all pass: Daybreak has a rivalry hook, concrete agent workflow, and code-security resonance. It is narrower than a model or ChatGPT capability release, so it stays in the 78–84 band.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export tool and run multivariate regression, found that post-dinner AI use correlated negatively with sleep duration, and added 1 hour 40 minutes of average nightly sleep after avoiding AI for several weeks.

Why it matters: HKR-H/K/R all pass: the personal reversal is clickable, the HealthKit/regression setup adds testable detail, and sleep loss hits AI practitioners directly. Scope is anecdotal, so it stays at the featured floor.

Bloomberg Technology

GitLab Says It Will Cut Jobs to Spend on Growth in the “Agentic Era”

GitLab said it will cut jobs to free up money for the market opportunity around AI agents; the RSS snippet does not disclose the number of roles, budget size, or execution timeline.

Why it matters: HKR-H and HKR-R pass: Bloomberg reports GitLab tying job cuts directly to agent investment, a strong devtools labor signal. HKR-K is weak because headcount, budget, and timing are missing.

r/LocalLLaMA

I catalogued every way local models break JSON output and built a repair library across 288 model calls

Reddit user kexxty ran 288 structured-output calls through OpenRouter models, including Llama 3, Mistral, Command R, DeepSeek, and Qwen, and found similar JSON failure categories across local and API-only models. The MIT-licensed Python library outputguard validates against JSON Schema, applies 15 ordered repair strategies, includes 2,001 tests, and has no LLM provider dependency.

Why it matters: HKR-H/K/R all pass: 288 tests, the outputguard library, and a 15-step repair chain give practitioners reusable detail. Source is a single Reddit post, so it stays in the 72–77 featured band, not 78+.

AI HOT (Curated Pool)

Introducing Daybreak: Frontier AI for Cyber Defenders

OpenAI introduced Daybreak for cyber defenders, combining OpenAI models, Codex, and security partners; the post does not disclose pricing, launch timing, or concrete defense metrics.

Why it matters: OpenAI’s Daybreak announcement clears HKR-H/R as a security-focused product hook, but HKR-K fails: no defense metrics, access terms, or pricing. That keeps it in the 72–77 product-update band.

AI HOT (Curated Pool)

Using LLMs in Script Shebang Lines

Simon Willison demonstrates using an LLM command in a script shebang line, with fragments generating SVG, the -T option calling llm_time, and a YAML template defining Python tools to compute 2344×5252+134 and return 12,310,822.

Why it matters: HKR-H/K/R all pass: Simon Willison shows a reproducible LLM-in-shebang workflow with concrete flags. Impact stays within CLI/script automation, not a model or platform release, so it sits in the low featured band.

AI HOT (Curated Pool)

Replit launches parallel agents with support for 10 concurrent agents

Replit launched parallel agents that run up to 10 agents concurrently, with each agent holding an independent copy of the app, working on its own machine, and merging the results through an agent workflow.

Why it matters: HKR-H/K/R pass: the post gives a concrete 10-agent parallel workflow with isolated app copies and merge. This is a mid-weight dev-tool update, below a Cursor Agent-mode-scale launch, so it sits at the featured threshold.

AI HOT (Curated Pool)

Anthropic Launches Claude Platform on AWS

Anthropic launched the Claude platform on AWS, letting AWS customers use existing authentication, billing, and committed-spend credits to access the full Claude API feature set, including hosted agents, code execution, and the Files API.

Why it matters: HKR-K and HKR-R pass: Anthropic brings Claude Platform into AWS procurement, billing, and committed spend. HKR-H is weak because this is distribution, not a model or capability launch.

The Verge · AI

Google Stopped a Zero-Day Hack It Says Was Developed With AI

Google says GTIG found and stopped its first AI-developed zero-day exploit, and the attackers planned a mass exploitation event to bypass two-factor authentication on an unnamed open-source web-based system administration tool; the post does not disclose the tool name.

Why it matters: This hits HKR-H/K/R: Google-backed AI zero-day claim, a concrete 2FA bypass target, and clear security resonance. Missing tool name and exploit details keep it in the 78–84 band.

May 11Monday

AI HOT (Curated Pool)

Cog House Opens for the First Time: Scott Wu and the Rise of Cognition AI

Cognition AI disclosed internal footage of Cog House, while Devin reached $445 million in annualized revenue within 18 months of launch and the company is valued at about $25 billion.

Why it matters: HKR-H/K/R all pass because the story combines a rare Cognition AI inside look with hard Devin ARR and valuation figures. It stops below P1 because this is a profile-style reveal, not a funding, product, or model release.

AI HOT (Curated Pool)

Pareto Code Reorders Model Selection Using Market Demand

OpenRouter says Pareto Code observes the Pareto frontier using real market demand; DeepSeek V4 Pro ranks first, followed by GPT 5.4 Mini and Gemini 3.1 Pro, while the post does not disclose the scoring formula or evaluation sample size.

Why it matters: HKR-H/K/R all pass, but the source is a single OpenRouter post with no sample size, time window, or pricing basis disclosed. It clears featured as a model-selection benchmark, not the 78+ band.

r/LocalLLaMA

ExLlamaV3 Major Updates

ExLlamaV3 added DFlash in v0.0.31, raising Coding throughput from 59.21 t/s to 177.67 t/s; v0.0.32 optimized five models, with Trinity-Nano gaining 72.4% on 6000 Pro², while v0.0.33 adds DFlash model quantization plus bug fixes and efficiency work.

Why it matters: HKR-H/K/R all pass, but the blast radius is mostly LocalLLaMA and ExLlama users. This fits a mid-weight open-source inference update, not a same-day industry-wide story.

Xinzhiyuan · WeChat

Largest IPO Nears, Topping SpaceX; 2028 AI Self-Iteration Countdown

Xinzhiyuan says Anthropic is considering a near-$1 trillion valuation, with ARR rising to $45 billion in five months; Jack Clark predicts a greater than 50% chance that AI systems can autonomously build better versions of themselves by the end of 2028, while the article cites a 72% Kalshi probability of an IPO announcement before November 1.

Why it matters: HKR-H/K/R all pass: the hook is sharp and the post gives valuation, ARR, and 2028 odds. Source is secondary and IPO/ARR claims lack official confirmation, so it stays in 78-84.

Xinzhiyuan · WeChat

Claude Mythos Hits 50% Success on 16-Hour Tasks in METR Time Horizons

Claude Mythos Preview reached a 50% success rate on METR Time Horizons tasks that take humans 16 hours, while only 5 of 228 tasks exceeded the 16-hour range, so the article says METR lacks enough samples to quantify longer-horizon performance.

Why it matters: HKR-H/K/R all pass: the 16-hour task result is a strong hook, and the METR sample caveat adds substance. Capped at 82 because only 5 tasks exceed 16 hours, so the 2027 extrapolation is not same-day P1 material.

Synced · WeChat

ICML 2026: PRISM Brings Efficient Test-Time Scaling to dLLMs

PRISM raises LLaDA-8B-Instruct on GSM8K from 67.58% to 85.30% by combining hierarchical trajectory search, partial remasking, and self-verified feedback, reducing dLLM test-time scaling cost from O(NT) toward O(N+KT) under a final candidate width K.

Why it matters: HKR-H/K/R all pass: the hook rejects brute-force scaling, the post gives GSM8K and complexity numbers, and it speaks to inference cost. Still an ICML framework paper, not a mainstream product release, so it sits in 78–84.

QbitAI · WeChat

Math Majors in Trouble: Fields Medalist Tests ChatGPT 5.5 Pro, Gets Paper-Level Result in 17 Minutes

Timothy Gowers tested ChatGPT 5.5 Pro on additive number theory problems, where it produced an optimal quadratic upper-bound construction in 17 minutes 5 seconds, then generated a LaTeX preprint in 47 minutes; the article says arXiv rejects AI-generated content, so the result remains on Gowers’s blog.

Why it matters: All three HKR axes pass: Gowers’ first-person test, 17m05s, and a 47-minute preprint are concrete and discussable. It is not a model release, but the named experiment and math-reasoning impact put it in the must-write band.

AI HOT (Curated Pool)

Codex autonomously completes a security audit and earns a bounty

A user instructed Codex to earn $5; Codex spent about 22 hours finding an open-source security audit bounty, submitting a valid PR, communicating with maintainers, passing GitHub verification, and ultimately receiving a $16.88 payment.

Why it matters: HKR-H/K/R all pass: a Codex agent allegedly closed a bounty loop in 22 hours with concrete money and workflow details. Single social-post evidence lacks reproducible logs, so it stays below P1.

r/LocalLLaMA

MTP benchmark results: task type determines speculative inference speedups or slowdowns

A Reddit LocalLLaMA user ran 300+ tests on Qwen 3.6 27B MTP quants, finding coding draft acceptance at 79-89% and F16 coding speed up 171%, while Q4_K_M creative writing slowed down 9%.

Why it matters: HKR-H/K/R all pass: this is a single Reddit experiment, not a market event, but 300+ Qwen 3.6 27B MTP quantization tests give practical numbers for local inference tuning.

May 10Sunday

r/LocalLLaMA

We tried vectors, ASTs, and brute-force context stuffing for code retrieval; LLM semantic graphs worked best

ByteBell open-sourced a code indexing system that stores per-file LLM-generated purpose, summary, business context, entities, classes, functions, keywords, and imports in a Neo4j graph, then uses full-text search instead of vector similarity, with SHA-256 diffing to reindex only changed files and keep LLM calls proportional to churn.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, and the post gives a concrete Neo4j semantic-graph mechanism with SHA-256 incremental rebuilds. Reddit sourcing and missing metrics keep it at the 72–77 featured threshold.

r/LocalLLaMA

I have DeepSeek V4 Pro at home

Reddit user fairydreaming ran DeepSeek V4 Pro Q4_K_M with a modified llama.cpp CUDA repo on one RTX PRO 6000 Blackwell Max-Q workstation GPU, using an 859GB model file; the shared log reports a 1M context window and 8.6 tokens per second generation speed.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU local inference, with concrete file size, context, speed, and runtime path. Reddit single-source sourcing keeps it below must-write model-release territory.