Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

901–920 of 1,196

May 13Wednesday

r/LocalLLaMA

The Trillion-Parameter Dilemma: MiMo-V2.5-Pro Open-Sourced at 1.02T Parameters

Xiaomi open-sourced MiMo-V2.5-Pro with 1.02T parameters, 42B active parameters, a 1M context window, and an MIT license; the author ran 125 Claude Code sessions through the API, spending $70.12 for 387,380,436 tokens with a 96.3% cache hit rate.

Why it matters: HKR-H/K/R all pass: a Xiaomi 1.02T open model plus a concrete Claude Code API cost experiment. Reddit sourcing keeps it at the low end of the 85+ band, but the domestic flagship-model signal clears p1.

New York Times Chinese

China Sought Access to Anthropic’s Latest Technology but Was Rejected

Chinese think-tank representatives asked Anthropic in Singapore last month to give Beijing access to Mythos, and Anthropic refused; the company has limited the vulnerability-finding model to the U.S. government and more than 40 organizations.

Why it matters: HKR-H/K/R all pass: the NYT report gives the Singapore request, Mythos’s bug-finding use, and its US-government-plus-40 access scope. This is a same-day security and US-China AI access story.

AI HOT (Curated Pool)

Claude Code adds /goal feature to keep tasks running until completion

Claude Code introduced a /goal feature that keeps Claude working until a task is completed; the post does not disclose the trigger mechanism, supported versions, pricing, or failure conditions.

Why it matters: HKR-H/K/R pass because /goal targets a real Claude Code reliability pain. It is a single-feature Anthropic update with sparse mechanics, so it lands at the lower featured band, not same-day major news.

r/LocalLLaMA

A real transformer language model running locally on a stock Game Boy Color

maddiedreese ran Andrej Karpathy’s TinyStories-260K on a stock Game Boy Color with INT8 weights, fixed-point math, an MBC5 ROM, bank-switched cartridge storage, and KV cache in cartridge SRAM; the demo uses no phone, PC, Wi‑Fi, link cable, or cloud inference, but output is extremely slow and gibberish.

Why it matters: HKR-H/K/R all pass: a named first-person experiment with concrete model and memory details. Impact stays low-featured because it is a Reddit hardware hack with slow, garbled output, not a usable product or model release.

AI HOT (Curated Pool)

90% of People Are Wasting Tokens

Andrej Karpathy says 90% of AI coding bills is wasted on unnecessary context, including repeated full-repository sends, expensive models for simple tasks, and missing prompt caching.

Why it matters: HKR-H/K/R all pass via the 90% claim, named waste mechanisms, and practitioner cost pain. It reaches featured, but stays at 72 because the post gives no billing sample or reproducible test.

AI HOT (Curated Pool)

Codex Enables Background Multitasking Across Apps

OpenAI Devs says computer use lets Codex click, type, and keep working across Mac apps in the background; the post does not disclose release timing, permission design, or availability scope.

Why it matters: HKR-H/K/R all pass: OpenAI Devs gives a concrete Codex mechanism for background Mac cross-app actions, with clear developer relevance. Release timing, permissions, and availability are missing, so it stays at the lower featured band.

AI HOT (Curated Pool)

How Anthropic's Cybersecurity Team Uses Claude Code to Build a Threat Detection Platform

Anthropic’s detection platform engineering team used Claude Code to build the CLUE threat detection and response platform, completing a proof of concept in one day and delivery in one week while reducing analyst log investigation from hours to minutes.

Why it matters: HKR-H/K/R all pass, but this is an Anthropic internal dogfooding case rather than a Claude Code capability launch. CLUE and the timing metrics keep it just above the featured threshold.

AI HOT (Curated Pool)

Claude Opus 4.7 Fast Mode Opens Research Preview

Claude Opus 4.7 Fast Mode is now available as a research preview in the API and Claude Code. The post does not disclose model parameters, pricing, rate limits, or a general availability date.

Why it matters: HKR-H/K/R pass because this is a Claude fast-mode preview in API and Claude Code, directly tied to developer latency and workflows. Thin disclosure on pricing, limits, parameters, and GA timing keeps it at the featured threshold, not 78+.

AI HOT (Curated Pool)

GitHub Copilot Individual Plans Add Flex Allotments and a New Max Plan

GitHub will update Copilot individual plans on June 1 by adding flex allotments to Pro and Pro+ and introducing a new Max plan; the post does not disclose pricing, quota limits, or the exact allocation rules in the provided snippet.

Why it matters: HKR-H/K/R all land lightly because Copilot plan quotas affect many developers. Missing price, caps, and allocation rules keep it at the low featured threshold, not a major capability update.

Hacker News front page

Show HN: Agentic Interface for Mainframes and COBOL

Hypercubic launched Hopper, an agentic development environment that combines a real TN3270 terminal, z/OS-aware panels for datasets, jobs, and spool output, and an AI agent; sensitive operations require approval, and the terminal remains visible during agent actions.

Why it matters: HKR-H/K/R all pass: the mainframe-agent angle is novel, with concrete TN3270, z/OS, and approval mechanics. Small-vendor Show HN status and missing customer/pricing/results data keep it at the featured floor.

TechCrunch · AI

Everything Google announced at its Android Show, from Googlebooks to vibe-coded widgets

Google announced AI-first Googlebooks laptops, more agentic Gemini features, vibe-coded Android widgets, Gemini in Chrome, and refreshed Android Auto ahead of I/O; the RSS snippet does not disclose specs, pricing, availability, or rollout timelines.

Why it matters: HKR-H/K/R all pass because Google bundled several Gemini/Android AI entry points with named product hooks. Missing parameters, pricing, rollout dates, and testable performance keeps it in the mid-weight product-update band.

AI HOT (Curated Pool)

Code w/ Claude SF 2026: Building on Exponential AI Growth

Anthropic expanded developer tooling at Code w/ Claude SF 2026: Claude Code rate limits doubled, Claude Opus API limits increased, and hosted agents on the Claude platform added four functions, including memory review, multi-agent delegation, output criteria, and webhooks.

Why it matters: Anthropic ships a substantive Claude Code update with concrete numbers and feature additions; HKR-H/K/R all pass. This is strong dev-tool news, not a flagship model release, so it fits the 78–84 band.

May 12Tuesday

AI HOT (Curated Pool)

Dungeons & Desktops: Building a Procedurally Generated Roguelike with GitHub Copilot CLI

A GitHub employee used GitHub Copilot CLI to build an extension that parses any codebase into one Roguelike-style dungeon layout, with procedural level generation used as the core mechanism for a creative coding and game prototyping demo.

Why it matters: HKR-H and HKR-K pass: an official GitHub tutorial has a novel demo and a clear mechanism. It is not a major Copilot capability release, and lacks production metrics, pricing, or benchmark data, so it sits at the tutorial-featured floor.

r/LocalLLaMA

Local LLM Autocomplete and Agentic Coding on a Single 16GB GPU + 64GB RAM

Reddit user grumd runs Qwen2.5-Coder-7B Q6 for autocomplete and Qwen3.6-35B-A3B Q8 for agentic coding on one RTX 5080 with RAM offloading; the post reports about 145k context, 56GB RAM used with other apps open, and Qwen3.6-35B-A3B speed of tg128 at 35.29 tokens/s.

Why it matters: HKR-H/K/R all pass: a named first-person local coding experiment with concrete model, quantization, context, and throughput data. Source is a single Reddit post without replication or comparisons, so it stays in the low featured band.

Hacker News front page

Show HN: Statewright – Visual State Machines for More Reliable AI Agents

Statewright uses a Rust state-machine engine to constrain Claude Code tool access, iterations, transitions, and guards; the post says 13–20B models improved consistently on real SWE-bench tasks, but it does not disclose benchmark scores, sample size, or the exact evaluation protocol.

Why it matters: HKR-H/K/R all pass: the state-machine constraint is a clear agent-reliability hook with a testable SWE-bench claim. Exact scores and reproduction details are not disclosed, so it stays just above the featured threshold.

QbitAI · WeChat

Markdown Is Fading? Karpathy Also Backs HTML

Anthropic engineer Thariq argued for using HTML instead of Markdown and gave 5 reasons; the post says HTML generation takes about 2 to 4 times longer than Markdown.

Why it matters: HKR-H/K/R all pass, but this is a developer format debate rather than a model or product launch. Named Anthropic/Karpathy context and the 2-4x time figure clear the featured threshold at the low end.

AI HOT (Curated Pool)

Large npm Supply-Chain Attack Hits TanStack, Mistral AI, UiPath, and Others

Socket identified the Mini Shai-Hulud supply-chain attack, where attackers used three GitHub Actions flaws to publish nearly 373 malicious versions across more than 160 npm package names, affecting projects including TanStack, Mistral AI, and UiPath and stealing AWS, GCP, Kubernetes, GitHub tokens, and SSH private keys during installation.

Why it matters: HKR-H/K/R all pass: named projects create the hook, Socket provides concrete counts and mechanisms, and credential theft matters to AI engineering teams. It is a strong security incident, not a core model or product release, so it stays in the 78–84 band.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export app in about 5 minutes and run multivariate regression, finding that the last post-dinner AI usage time correlated negatively with sleep duration; after avoiding AI at night, average sleep increased by 1 hour and 40 minutes.

Why it matters: HKR-H/K/R all pass: a first-person quantified experiment links post-dinner AI use to shorter sleep, then reports +1h40m after stopping. Personal-blog scope keeps it below major industry-update territory.

AI HOT (Curated Pool)

What Parameter Golf Taught Us About AI-Assisted Research

OpenAI’s Parameter Golf brought together over 1,000 participants and more than 2,000 submissions to test AI-assisted machine learning research, coding agents, model quantization, and model design under strict parameter constraints.

Why it matters: OpenAI’s Parameter Golf recap clears HKR-H/K/R with a concrete contest, 1,000+ participants, and 2,000+ submissions. It is research/benchmark signal, not a model or product launch, so 78 fits the lower featured band.

The Verge · AI

OpenAI just released its answer to Claude Mythos

OpenAI launched Daybreak, a security initiative that uses the Codex Security AI agent released in March to model an organization’s code, validate likely vulnerabilities, and automate detection of higher-risk issues before attackers find them.

Why it matters: HKR-H/K/R all pass: Daybreak has a rivalry hook, concrete agent workflow, and code-security resonance. It is narrower than a model or ChatGPT capability release, so it stays in the 78–84 band.