Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,195 picksRelated topicsAgentsCursorTutorials

Latest picks

1001–1020 of 1,195

Apr 29Wednesday

X · @dotey

AI terminal tool Warp open-sources client code with OpenAI as founding sponsor

Warp open-sourced its client code under AGPL; only the client is open, while server code stays closed. The Rust terminal has 700,000+ developers, and its Oz cloud AI handles coding, planning, and tests. The key signal is its AI-first contribution workflow.

Why it matters: HKR-H/K/R all pass: the OpenAI sponsorship hook, AGPL/client-only detail, and 700K-developer signal are concrete. This is a strong dev-tool open-source update, not a major model or capability release.

Apr 28Tuesday

Ben's Bites

Builders

Ben’s Bites published one newsletter on AI builders. It says OpenAI released GPT-5.5 at 2x GPT-5.4 pricing, with a claimed 40% token-efficiency gain. Claude Managed Agents memory entered public beta, and Cursor’s SpaceX/xAI deal includes a $60B 2026 purchase option.

Why it matters: HKR-H/K/R all pass: GPT-5.5 cost/efficiency figures, Claude Managed Agents Memory beta, and a Cursor deal term. It stays in 85–94 because this is a newsletter roundup, not a primary release.

QbitAI · WeChat

Xiaomi open-sources MiMo-V2.5 series; Pro builds a macOS-like desktop in 4 hours

Xiaomi open-sourced MiMo-V2.5 weights, covering Pro Agent, multimodal base, TTS, and ASR models. MiMo-V2.5-Pro built a 54-app macOS-like desktop in 4 hours without human takeover; it scored 233/233 on SysY with 672 tool calls in 4.3 hours. Key details for practitioners are the 1M context, 100T-token program, and free Agent-framework access.

Why it matters: HKR-H/K/R all pass: Xiaomi open-sourced MiMo-V2.5 weights with concrete agent and coding-task numbers. Domestic flagship model release bump puts it in the must-write same-day band.

Hacker News front page

Xiaomi releases MiMo-v2.5 weights with strong coding and agent benchmarks

Xiaomi released MiMo-v2.5 family weights; the title cites strong coding and agent benchmarks. The RSS body only lists URLs, 13 HN points and 2 comments; the post does not disclose size, license, or scores.

Why it matters: HKR-H/K/R pass because a Xiaomi coding/agent weights release is concrete and practitioner-relevant. Sparse sourcing holds it near the featured floor: no parameters, license, or benchmark numbers are disclosed.

The Verge · AI

Attack of the Killer Script Kiddies

The Verge discusses Claude Mythos and AI bug finding, citing DARPA AIxCC scans over 54 million code lines. Teams found most seeded flaws plus over a dozen unseeded bugs; the RSS snippet does not disclose Mythos benchmarks, pricing, or access terms.

Why it matters: HKR-H/K/R all pass: the hook is strong, DARPA AIxCC supplies concrete numbers, and the security angle resonates. No Claude Mythos benchmark, pricing, or access terms are disclosed, so it stays in the featured-threshold band.

Hacker News front page

GitHub Copilot code review will start consuming GitHub Actions minutes

GitHub will make Copilot code reviews consume GitHub Actions minutes starting June 1, 2026. Private-repo reviews use plan entitlements, with overages billed at standard Actions rates; public repos stay free. The change covers Copilot Pro, Pro+, Business, and Enterprise, including direct org billing for unlicensed users.

Why it matters: Official GitHub billing change for Copilot code review hits CI quotas and org invoices; HKR-H/K/R all pass, but it is a pricing rule, not a capability release, so it sits low in 72–77.

Xinzhiyuan · WeChat

Claude bans hit 110-person firm; Cursor incident deletes database in 9 seconds

Anthropic allegedly suspended 110 Claude accounts at a US agtech firm, while API billing continued. The post says appeals went unanswered for 36 hours, and PocketOS says Claude Opus 4.6 via Cursor deleted production data and volume backups in 9 seconds. The key issue is access control: no RBAC, no environment isolation, and no delete confirmation.

Why it matters: HKR-H/K/R all pass: the incident has a strong hook and concrete details: 110 accounts, 36 hours, 9 seconds, and no RBAC. Kept at 82 because it is still a single-source allegation without an Anthropic postmortem.

X · @op7418

Xiaomi open-sources the MiMo-V2.5 model series

Xiaomi open-sourced the MiMo-V2.5 model series under the MIT license for commercial use, retraining, and fine-tuning. It also launched Orbit 100T Token, offering approved AI builders up to 1.6B credits worth 659 yuan. Agent framework teams can apply for free MiMo token access; the post does not disclose model size or benchmark results.

Why it matters: HKR-H/K/R all pass: Xiaomi MiMo-V2.5 open source, MIT terms, and Orbit 100T credits matter to builders. Missing params and benchmarks keep it in the 78–84 band, below P1.

r/LocalLLaMA

Local coding models have reached a threshold for real work

Antigma tested 27B–32B open-weight models; Qwen 3.6-27B scored 38.2% on Terminal-Bench 2.0. The run used 89 tasks and the default per-task timeout, while verified SOTA is about 80%. The key claim is deployment lag: offline coding is about 6–8 months behind hosted frontier models.

Why it matters: HKR-H/K/R all pass: the post gives a real-work threshold claim, a 38.2%/89-task Terminal-Bench result, and a 6–8 month offline gap. Reddit single-post sourcing keeps it in the low featured band.

Hacker News front page

Claude Pro: Opus Requires Extra Usage in Claude Code

Anthropic lists 6 Claude Code models, and Pro users need extra usage enabled and purchased to use Opus. The guide gives 3 configuration paths: /model, --model, and ANTHROPIC_MODEL in zsh or bash. The post does not disclose extra usage pricing or quotas.

Why it matters: HKR-H/K/R all pass, but the facts come from a help doc and cover Claude Code access/configuration, not a new model or major capability. Anthropic relevance lifts it to the lower featured band.

X · @dotey

The West forgot how to build things, and may forget how to write code

Denis Stetskov compares Western defense production gaps with AI coding, citing Stinger orders placed in 2022 for 2026 delivery. He says Europe’s 1M-shell target was 9 months late, and METR found senior developers 19% slower with AI. The key risk is the junior-engineer pipeline, not code generation speed.

Why it matters: HKR-H/K/R all pass: the analogy is clickable, the post gives concrete defense and METR numbers, and the junior-engineer pipeline resonates. X translation/commentary limits authority, so it sits just above the featured threshold.

X · @dotey

GitHub Copilot switches to usage-based billing on June 1

GitHub Copilot will switch to AI Credits billing on June 1 while keeping subscription prices unchanged. Credits count input, output, and cached tokens; Pro includes $10 monthly credits and Pro+ includes $39. Watch Copilot Agent long-task costs.

Why it matters: HKR-H/K/R all pass: Copilot billing moves from subscription expectations to token/cache consumption with date and credit amounts. Single-source X context lacks enterprise details and overage rates, so it stays in the 78–84 band.

X · @dotey

Cursor 3 feedback: users want a reliable AI development workspace

Eric Zakariasson’s Cursor 3 feedback thread summarizes 431 replies, with users asking for a stable AI development workspace. Requests center on Agent Window retaining LSP, debugging, Git, terminal and diff workflows, plus multi-agent worktrees and model-cost transparency. The key issue is workflow reliability, not a flashier IDE.

Why it matters: All HKR axes pass: 431 user replies, concrete workflow requests, and strong resonance for Cursor users. Kept in the low featured band because this is feedback synthesis, not an official Cursor release or roadmap.

Hacker News front page

GitHub Copilot is moving to usage-based billing

GitHub said on 2026-04-27 that GitHub Copilot will move to usage-based billing. The captured post only shows the title, time, and navigation. It does not disclose the launch date, usage metric, prices, or overage rules.

Why it matters: GitHub Copilot billing affects a large developer base. HKR-H and HKR-R are strong, while HKR-K is limited to the usage-based mechanism with no date, metering unit, or price details disclosed.

Apr 27Monday

Dwarkesh Patel podcast

What I've been Thinking About This Weekend: Open Questions, Intelligence vs Power, Verification in Science

Dwarkesh lists open AI questions, including that five hyperscalers own over 70% of global AI compute. He asks about coding agents, KV cache costs, merging training with inference, and online learning; the post gives questions, not experimental answers.

Why it matters: HKR-H/K/R all pass: Dwarkesh adds a concrete compute-concentration claim and practitioner-relevant questions. No experiment, release, or policy change, so it stays in the 72–77 commentary band.

Hacker News front page

Running Local LLMs Offline on a Ten-Hour Flight

Dmitri Lerko ran Gemma 4 31B and Qwen 4.6 36B locally during a 10-hour flight with no Wi‑Fi. The MacBook Pro M5 Max had 128GB unified memory and a 40-core GPU; sustained load used about 1% battery per minute, and performance degraded past 100k tokens. The sharp finding is instrumentation: an iPhone cable delivered 60W, while a MacBook cable delivered 94W under the same load.

Why it matters: HKR-H/K/R all pass: this is a named first-person local-inference test with concrete hardware, model, battery, and power numbers. Scope stays practical rather than industry-shaking, so it lands in the 72–77 band.

Hacker News front page

Show HN: OSS Agent Dirac topped TerminalBench on Gemini-3-flash-preview

Dirac-run released Dirac and says it topped TerminalBench using Gemini-3-flash-preview. The repo claims 50-80% lower API costs via Hash Anchored edits, parallel operations, and AST manipulation; the post does not disclose full scores.

Why it matters: HKR-H/K/R all pass: an OSS coding agent claims a TerminalBench lead with cost and mechanism details. Held to 78 because the post relies on repo claims and lacks full leaderboard scores or reproduction logs.

Hacker News front page

AI can cost more than human workers now

Axios says some firms now spend more on AI than salaries; Nvidia's Bryan Catanzaro says compute costs exceed employee costs. Gartner forecasts 2026 IT spending at $6.31T, up 13.5%, driven by AI infrastructure, software, and cloud. Watch token costs: Uber's CTO has already exhausted the 2026 AI budget.

Why it matters: HKR-H/K/R all pass: the piece turns AI cost anxiety into budget facts, including Nvidia compute costs and Uber’s token-budget issue. It stays in the 72–77 band because this is trend reporting, not a launch or hard news event.

QbitAI · WeChat

DeepSeek V4 Cuts Prices Permanently; Cached Inputs Get 90% Off, Coding Test Costs Drop 83%

DeepSeek V4 cut prices twice in two days: input/output pricing is 75% lower, with cached inputs getting another 90% off. QbitAI’s coding test fell from 31.73 yuan for 35M tokens to 5.34 yuan under new pricing, an 83% drop. The key case is high cache-hit workloads, with V4-Pro at about 95–96% cache hits.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 pricing has a sharp cost hook, concrete test numbers, and strong cost resonance. It is still a pricing update, not a new model release, so it stays below the 85 P1 band.

Synced · WeChat

ACL 2026: Sending AI “~” May Cause It to Delete Your Home Directory

ACL 2026 accepted an LLM safety paper on emoticon semantic confusion. The team tested 6 models with 3,757 cases; average confusion was 38.6%, with over 90% silent failures. The key risk is agent execution, where “ignore emoticons” prompts had limited effect.

Why it matters: ACL 2026 safety research clears HKR-H/K/R: a sharp file-deletion hook, concrete test numbers, and direct agent-execution risk. It is strong research, not a model launch or platform incident, so it stays in the 78–84 band.