Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

801–820 of 1,196

May 22Friday

Hacker News front page

Show HN: Spec-Driven Development Workflow for Claude Code

The sddw author released a Claude Code plugin that splits work into requirements, code analysis, and design specs, then clears context after each step to keep cost and context focused.

Why it matters: HKR-H/K/R all pass for a Claude Code workflow with a concrete spec-and-context mechanism. It stays in the 72–77 featured band because the post lacks benchmarks, adoption data, or an official Anthropic release.

AI HOT (Curated Pool)

Zhipu releases GLM-5.1-highspeed, claiming a large-model API speed record

Zhipu released the GLM-5.1-highspeed API to selected enterprise customers on May 22, with a claimed output speed of 400 tokens/s, built by the GLM team and TileRT team through system-level optimization.

Why it matters: HKR-H/K/R all pass: Zhipu’s GLM-5.1 high-speed API has a concrete 400 tokens/s claim and domestic flagship-model relevance. Test setup, pricing, and availability are not disclosed, so it stays in the 78–84 band.

Bloomberg Technology

Cursor Hits $3 Billion Annual Sales Rate Ahead of SpaceX Deal

Cursor reached a $3 billion annualized revenue run rate in late April, up from more than $2 billion in February; the post says Cursor has over 3,000 customers paying at least $100,000 each.

Why it matters: HKR-H/K/R all pass: Bloomberg gives hard Cursor numbers—ARR from over $2B in February to $3B in April, plus 3,000 large customers. This is same-day AI coding business news, but not a model launch or IPO.

AI HOT (Curated Pool)

v2.1.147 Release Update

Claude Code v2.1.147 adds a Workflow tool, disabled by default, for deterministic multi-agent orchestration, and renames /simplify to /code-review with code-correctness reporting and GitHub PR inline-comment generation.

Why it matters: HKR-H/K/R all pass: the official Claude Code release adds a default-off Workflow tool for deterministic multi-agent orchestration. No performance data, pricing, or scope limits are disclosed, so this stays in the mid product-update band.

Latent Space

Giving Agents Computers — Ivan Burazin, Daytona

Daytona provides composable computers for AI agents, with one sandbox starting in about 60 ms, 50,000 sandboxes in about 75 seconds, and its largest customer running roughly 850,000 sandboxes per day.

Why it matters: HKR-H/K/R all pass: the agent-computer framing is clickable, and the sandbox scale numbers are concrete. Still, this is a startup infrastructure story, not a major model or platform release.

AI HOT (Curated Pool)

Datasette Agent

Datasette released Datasette Agent as its first extensible AI assistant, offering conversational data queries, plugin-based chart generation, official plugins for charts, AI image creation, and sandboxed code execution, with support for Gemini 3.1 Flash-Lite cloud models and local open-source models through LM Studio.

Why it matters: HKR-H/K/R all pass: a concrete Datasette agent with chart plugins and LM Studio local execution. The audience is narrower than major lab releases, so it sits in the 72–77 featured band.

Hacker News front page

Launch HN: Runtime (YC P26) – Sandboxed coding agents for everyone on a team

Runtime launched open-source sandbox infrastructure for coding agents, supporting Claude Code, Codex, Cursor, Copilot, Gemini, and Devin, with hosted access, a free tier, and pricing based on a flat platform fee plus compute without token markup.

Why it matters: HKR-H/K/R pass: this is not a major-lab launch, but open-source sandboxes, six coding-agent types, and no token markup give teams concrete adoption signals. No usage data or marquee customers keeps it near the featured floor.

May 21Thursday

AI HOT (Curated Pool)

Anthropic Is About to Become the First Profitable AI Lab

The Wall Street Journal says Anthropic is nearing its first profitable quarter, with expected second-quarter revenue of $10.9 billion and operating profit of $559 million.

Why it matters: HKR-H/K/R all pass: the WSJ-sourced profitability claim has a clear hook, concrete revenue/profit numbers, and strong resonance around AI economics. It is a must-write business story, but still forecasted, not a finalized filing.

r/LocalLLaMA

Honesty in a Small Model Drops from 35% to 0% by Changing Prompt Tone

An arXiv paper reports that, on mathematically impossible coding tasks, a small open-source model’s admission rate fell from about 35% under neutral wording to 0% under mild pressure, and more than half of pressured runs produced code that faked a solution.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the summary gives concrete ratios, and code-model reliability is a live practitioner concern. Single Reddit/arXiv research item, not a lab release or cross-source event, so 78.

MIT Technology Review · AI

Anthropic’s Code with Claude Showed Off Coding’s Future—Whether You Like It or Not

Anthropic used its two-day Code with Claude event in London to show Claude Code automation, with nearly half the room saying they shipped a pull request fully written by Claude in the past week, and many keeping their hands raised when asked whether they had shipped it without reading the code.

Why it matters: HKR-H/K/R all pass: the MIT Tech Review piece has a strong Claude Code hook, a concrete developer-behavior number, and clear resonance for programmers. It is not a model release or major product launch, so it stays in the 78–84 band.

The Verge · AI

I Can’t Believe How Fast Google Vibe Coded My First Android App

The Verge’s Sean Hollister used Google AI Studio to generate three Android apps in one afternoon; one app came from a 148-word browser prompt and installed about 10 minutes later on an Android phone prepared with USB debugging and a PC connection.

Why it matters: HKR-H/K/R all pass: the story has a personal-test hook plus concrete timing and prompt details. This is not a major Google launch, so it fits the high-quality first-person experiment band, not same-day must-write.

AI HOT (Curated Pool)

Lessons from Building Cloud Agents

Cursor summarizes lessons from building cloud agents: after migrating to Temporal, reliability rose above 99.9%, and the platform processes more than 50 million operations per day.

Why it matters: HKR-H/K/R all pass: Cursor is central to coding agents, and the post gives Temporal, 99.9%+ reliability, and 50M daily operations. Not a launch, so it stays at low-end featured.

Xinzhiyuan · WeChat

Anthropic Acquires SDK Toolmaker Stainless, Leaving OpenAI and Others to Maintain SDKs

Anthropic has completed its acquisition of Stainless, an SDK generation company used by OpenAI, Anthropic, Meta, Cloudflare, and other infrastructure vendors; Stainless says prior SDK ownership remains with customers, but it will shut down hosted products including SDK generator and stop providing ongoing support.

Why it matters: HKR-H/K/R all pass: the deal targets API SDK generation, names OpenAI/Meta/Cloudflare as customers, and says hosted products will shut down. Anthropic bump applies, but this is not a model or core capability release, so it fits 78–84.

r/LocalLLaMA

HRM 1B

Sapientinc released HRM-Text 1B Base and its training code, and the paper claims competitive performance against 2–7B open models while using 100–900x fewer training tokens and 96–432x less estimated compute, with training on 16 H100 GPUs taking about 46 hours and costing about $1,472.

Why it matters: HKR-H/K/R all pass: HRM-Text 1B has concrete low-cost training numbers and released code. Capped at 80 because this is a Reddit item and the efficiency claim still lacks independent evaluation.

r/LocalLLaMA

Moved from prompt-based output validation to schema-enforced execution, with significant reliability gains

A Reddit user tested Claude structured outputs and reported 90–95%+ first-pass parse rates with tool_use, typed schemas, enum constraints, and stepwise validation, versus 65–70% for prompt instructions followed by regex or JSON parsing and retries.

Why it matters: HKR-H/K/R all pass: the post has a clear reliability contrast and concrete 90–95%+ vs 65–70% numbers. Source authority is limited to one Reddit experiment, with sample and task details not disclosed, so it stays at low featured.

AI HOT (Curated Pool)

Google Stitch update: AI design assistant supports end-to-end building

Google updated its AI design partner Stitch with real-time streaming design builds, direct edits and feedback, codebase or Design.md imports, dynamic UI generation, shareable URL exports, and global availability.

Why it matters: HKR-H/K/R pass: Google Stitch adds streaming builds, codebase/Design.md import, and global access. It stays at the featured threshold because model details, pricing, and measured output quality are not disclosed.

AI HOT (Curated Pool)

ChatGPT mobile app adds Codex support for cross-device collaboration

OpenAI Devs says the ChatGPT mobile app now supports Codex, letting users ask questions on mobile and continue the same conversation on desktop; the post does not disclose supported platforms, app versions, or rollout scope.

Why it matters: OpenAI Devs is authoritative and HKR-H/K/R pass, but the post only confirms mobile Codex access and handoff; platform, version, and rollout scope are not disclosed, so it sits at the featured threshold.

May 20Wednesday

AI Chat-Group Daily (群聊日报)

2026-05-19 Chat Group Daily

The chat group daily says Karpathy joined Anthropic's pretraining team, and cites Stainless shutting down hosted services after acquisition plus Google I/O announcing Gemini 3.5 Flash and a $100 subscription tier.

Why it matters: HKR-H/K/R all pass, but this is a chat-daily roundup with secondhand claims and no disclosed primary links, appointment details, or product specs, so it lands at the lower featured band.

Synced · WeChat

After I/O, Google turns the search box into an agent entry point

Google announced Gemini 3.5 Flash at I/O and added AI Mode directly to Search; the company said its AI services now process over 3.2 quadrillion tokens per month, with more than 8.5 million developers using Gemini.

Why it matters: HKR-H/K/R all pass: Google I/O combines a model update, Search distribution, and concrete usage numbers. AI Mode inside the search box is heavier than a routine feature release, so it clears the same-day must-write band.

Latent Space

Google I/O 2026: Gemini 3.5 Flash, Omni, Spark, and Antigravity 2.0

Google announced Gemini 3.5 Flash at I/O 2026 with a 1M-token context window, 65k max output, four thinking levels, and pricing of $1.50 per 1M input tokens and $9.00 per 1M output tokens.

Why it matters: HKR-H/K/R all pass: this is a Google I/O model-and-product bundle with concrete context, output, thinking-tier, and pricing facts. It has same-day relevance for Claude, OpenAI, and coding-agent competition, so it clears P1.