Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

1041–1060 of 1,196

Apr 24Friday

Hacker News front page

GPT-5.5: Mythos-Like Hacking, Open to All

XBOW says GPT-5.5 cut miss rate to 10% on its real-vulnerability benchmark, versus 40% for GPT-5 and 18% for Opus 4.6. It scored 97.5% on visual acuity and used about half the login iterations of the next-best model. The key point is black-box testing: GPT-5.5 without source beat GPT-5 with source.

Why it matters: HKR-H/K/R all pass: a major OpenAI model claim, concrete security benchmark numbers, and a clear practitioner safety nerve. The source is XBOW rather than an OpenAI launch post, so it stays below 95.

The Verge · AI

OpenAI says its new GPT-5.5 model is more efficient and better at coding

OpenAI announced GPT-5.5 and says it is more efficient and stronger at coding than GPT-5.4, which shipped last month. The RSS snippet says it handles coding, debugging, online research, and cross-tool work on spreadsheets and documents; the post does not disclose pricing, context window, or benchmark scores.

Why it matters: An OpenAI model release is same-day coverage, and the angle ties efficiency, coding, and tool use into one clear upgrade, so HKR-H/K/R all pass. The post does not disclose price, context window, or benchmark scores, which keeps it in the high 80s instead of 90+.

Hacker News front page

An update on recent Claude Code quality reports

Anthropic said three product-layer changes degraded Claude Code quality for Sonnet 4.6, Opus 4.6, and Opus 4.7, while the API was unaffected; all were fixed on April 20 in v2.1.116. The changes were lowering default reasoning effort on March 4, a March 26 bug that cleared prior thinking every turn after sessions sat idle for over an hour, and an April 16 prompt tweak to reduce verbosity that hurt coding quality. The signal for practitioners is sharp: product and prompt changes can degrade code performance even when model and inference evals do not reproduce it early.

Apr 23Thursday

QbitAI · WeChat

Qwen3.6-27B open-weights, beats its 397B flagship predecessor on agentic coding

Qwen released Qwen3.6-27B and says it beats Qwen3.5-397B on 4 agentic coding benchmarks with about 1/15 the parameters. The post cites SkillsBench rising from 30.0 to 48.2, GPQA Diamond at 87.8, and AIME26 at 94.1; it uses a dense architecture, Thinking Preservation, and Gated DeltaNet, with weights on Hugging Face and ModelScope.

Why it matters: This is a substantive Qwen open-source model release with concrete agent-coding and reasoning scores, so HKR-H/K/R all pass. I keep it at 84, not higher, because the post gives strong benchmarks but no pricing, context window, or independent reproduction yet.

OpenAI News

Introducing GPT-5.5

OpenAI introduced GPT-5.5 and says it targets complex cross-tool tasks such as coding, research, and data analysis. The RSS snippet only confirms “faster” and “more capable”; the post does not disclose benchmarks, context window, pricing, release timing, or availability, which are the details practitioners should watch.

Why it matters: An OpenAI flagship-model release is same-day news, so HKR-H and HKR-R are clear. HKR-K fails because the post discloses the name and use cases but not benchmarks, context window, price, or availability, so this stays featured rather than p1.

Xinzhiyuan · WeChat

Historic moment: Anthropic nears $1 trillion on private secondary markets, surpassing OpenAI for the first time

Anthropic was quoted at $1.05T-$1.15T on private secondary markets, above OpenAI’s roughly $880B quotes on similar platforms. The post attributes the rerating to scarce float, a sharp rise from a $380B funding valuation three months earlier, and momentum around Claude Code and revenue growth; it does not disclose trade volume, revenue figures, or company confirmation. Do not confuse this with a new funding valuation: these are secondary-market quotes on platforms such as Forge Global.

Why it matters: The signal is a private-secondary quote of $1.05T-$1.15T for Anthropic, above OpenAI's quoted ~$880B, not a new financing round. HKR-H/K/R all pass, but missing volume, revenue detail, and company confirmation keep it in the good-quality band, not must-write.

New York Times Chinese

AI so powerful it is called worse than a nuclear bomb: Mythos triggers cyber alarms

Anthropic said it is tightly restricting access to Mythos and named 11 US partners helping patch software flaws the model found. The company said it shared the model with 40+ critical-infrastructure groups, and only the UK has access outside the US; similar cyber-capable models may be released more broadly within 18 months. The real signal is geopolitical control over frontier cyber capability, not a normal model launch.

Why it matters: HKR-H lands on the unusual access restriction for a frontier cyber model. HKR-K lands on 11 partners, 40+ institutions, and the 18-month spread claim; HKR-R lands on the security and export-control nerve. Kept at 84 because benchmark details and eval methods are not disclosed.

Bloomberg Technology

Andreessen, Thrive Poised for Windfall From SpaceX's Bid for Cursor

If SpaceX acquires AI coding startup Cursor for $60 billion, Andreessen Horowitz and Thrive Capital stand to gain billions. The RSS snippet says Andreessen holds about 10%, worth roughly $6 billion at that price; the post does not disclose whether a deal is signed or Thrive's exact stake.

Why it matters: This is a real AI-devtools story, not just a finance sidebar: a reported $60B SpaceX bid for Cursor is novel, concrete, and highly discussable. I stop at featured, not p1, because the article does not disclose whether a deal is signed or what changes for Cursor’s product and go‑t

Latent Space

Shopify’s AI Phase Transition: 2026 Usage Explosion, Unlimited Opus-4.6 Budget, Tangle

Shopify CTO Mikhail Parakhin detailed its AI stack across 3 projects: Tangle, Tangent, and SimGym. The post says Shopify is a 20-year, $200B software company, but does not disclose exact 2026 usage figures. The key shift is from code generation to review, CI/CD, and deployment stability.

Why it matters: HKR-H/K/R all pass: the CTO interview has a clear hook, names internal tools, and maps the coding-agent bottleneck to review and CI/CD. Missing usage numbers keep it in 78–84, not P1.

Hacker News front page

Coding Models Are Doing Too Much

The author programmatically corrupts 400 BigCodeBench problems with single-point bugs to test whether coding models over-edit code during fixes. The post defines the minimal fix as exactly reversing the corruption and measures excess changes with token-level Python Levenshtein distance. The provided body does not disclose final results, model rankings, or training gains.

Why it matters: Strong HKR-K from a concrete 400-task bug-injection eval and a clear minimal-patch metric. HKR-R also lands because over-editing is a daily pain point for Copilot/Cursor/Claude Code users, but the excerpt omits results, model rankings, and effect sizes, so this sits near the low

Hacker News front page

Introducing Parallel Agents in Zed

Zed released Parallel Agents on April 22, 2026, letting multiple agents run in parallel in one window. The new Threads Sidebar sets per-thread folder and repo access, and supports stop, archive, and new-thread actions; the new default layout is opt-in for existing users. The key detail is permission scoping and thread orchestration, not just “multiple agents.”

Why it matters: First-party product update with clear HKR-H/K/R: parallel agents in one window plus thread-level repo and folder access control. It stays in the mid-70s because the post gives no performance delta, pricing impact, adoption data, or external validation; this is still a single-tool

Hacker News front page

Startups Brag They Spend More Money on AI Than Human Employees

Swan AI CEO Amos Bar-Joseph said his 4-person startup spent $113,000 on Claude in one month and treated that bill as headcount budget spent on AI instead of hires. The post says Swan targets $10M ARR with fewer than 10 people and cites Fundable AI claiming AI can replace a 15-person document team; the real signal is that token spend is being used as a growth metric, not proven ROI.

Why it matters: HKR-H lands on the payroll-vs-AI-bill inversion; HKR-K lands on the $113k/month Claude spend from a 4-person team. HKR-R is strong because it speaks to hiring, burn, and replacement anxiety, but this is still a trend piece with a thin sample, not a market-moving event.

Apr 22Wednesday

Hacker News front page

Show HN submissions tripled and are now mostly vibe-coded

Adrian Krebs scored 500 recent Show HN landing pages and says submissions have tripled, with 67% of pages triggering at least 2 AI design patterns. The method used Playwright plus an in-page script to check DOM and computed styles across 15 deterministic CSS/DOM signals; manual QA found about 5% to 10% false positives. The real signal is not model quality, but fast homogenization from AI default frontend templates.

Why it matters: This clears HKR-H/K/R: a sharp hook, a concrete 500-page method, and a real nerve for AI builders. I keep it at 78, not higher, because it is a single-author experiment rather than a product launch or a cross-source industry event.

Hacker News front page

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Qwen released the open-weight 27B dense model Qwen3.6-27B and made it available in Qwen Studio. It scores 77.2 on SWE-bench Verified vs. 76.2 for Qwen3.5-397B-A17B, and 59.3 on Terminal-Bench 2.0 under a 256K context and 3-hour timeout. The real takeaway is deployment: this is not a larger MoE, but a denser 27B model with stronger coding results.

Why it matters: Qwen3.6-27B is a substantive flagship-model release with open weights, concrete coding benchmarks, and a practical dense-deployment angle. HKR-H/K/R all pass, and per policy a major Chinese model launch should score on par with an equivalent US-lab release.

OpenAI News

Introducing workspace agents in ChatGPT

OpenAI introduced workspace agents in ChatGPT, describing them as Codex-powered agents that automate complex workflows in the cloud. The RSS snippet confirms secure work across tools for teams, but the post does not disclose pricing, availability, supported tools, or performance metrics.

Why it matters: This is a substantive OpenAI product update inside ChatGPT. HKR-H lands on the jump from chat to workspace agents, HKR-K on Codex-powered cloud execution across tools, and HKR-R on team workflow automation; the score stops at 86 because pricing, rollout, tool support, and metrics

The Verge · AI

Anthropic’s most dangerous AI model just fell into the wrong hands

Anthropic’s Claude Mythos Preview was accessed by a small group of unauthorized users through a contractor’s access plus common internet sleuthing tools. The snippet says the model can identify and exploit flaws in major operating systems and browsers; the post does not disclose the group size, dwell time, or remediation status. The key issue is access control failure, not the headline’s danger framing.

Why it matters: This is a real Anthropic security incident with a concrete access path, so HKR-H/K/R all pass: strong hook, new mechanism, and clear governance resonance. It stays below 85 because user count, exposure window, and remediation status are not disclosed.

Xinzhiyuan · WeChat

Musk sets Cursor deal terms: SpaceX can buy it for $60B or pay a $10B collaboration fee

SpaceX disclosed terms to work with Cursor: it can acquire the startup this year for $60B or pay a $10B collaboration fee. The post says xAI already supplies Colossus compute to train Cursor's Composer, while Cursor had over $2B annualized revenue and 1M+ daily users by Feb. 2026. The key point is the bundle: SpaceX gets a coding product, and Cursor gets model and compute support.

Why it matters: The reported deal structure alone is material: a $60B buy option or a $10B price to extend the partnership. HKR-H/K/R all clear on novelty, numbers, and industry resonance, but without primary docs or clear multi-source confirmation, it stays high featured rather than p1.

Financial Times · Technology

SpaceX obtains right to buy AI start-up Cursor for $60bn

The headline says SpaceX obtained the right to buy AI start-up Cursor for $60bn. The RSS snippet adds one fact: Elon Musk’s rocket and AI group is trying to catch up with OpenAI and Anthropic; the post does not disclose the trigger, timeline, deal structure, or regulatory terms. The key issue is whether this is tied to future financing or control rights, not a completed acquisition.

X · @dotey

Anthropic quietly removed Claude Code from the $20 Pro plan on its pricing page without an announcement

Anthropic was spotted removing Claude Code from the $20 Pro plan on its pricing comparison page without an announcement. The snippet says help docs also removed the inclusion, while the Claude Code product page and support bot still say it is included, and some Pro users report access still works; the post does not disclose Anthropic’s formal explanation or effective date. The key issue is price floor: if confirmed, entry cost for Claude Code rises from $20 to $100 per month.

Why it matters: The story matters because it may raise Claude Code’s entry price from $20 to $100, giving it HKR-H, HKR-K, and HKR-R. I keep it in featured, not higher, because Anthropic has not confirmed scope, timing, or treatment of existing Pro users.

TechCrunch · AI

SpaceX is working with Cursor and has an option to buy the startup for $60B

SpaceX is working with Cursor and holds an option to acquire the startup for $60B. The RSS snippet discloses the collaboration and purchase option, but not the term, trigger conditions, ownership impact, or payment mix. The sharper signal is strategic weakness: the snippet says neither Cursor nor xAI has proprietary models matching leading offerings from Anthropic and OpenAI.

Why it matters: HKR-H lands on the surprise combo and the $60B option. HKR-K lands on the concrete number and deal structure; HKR-R lands because Cursor is a daily tool for AI builders. Missing term details keep it at the low end of p1, not higher.