Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

1061–1080 of 1,196

Apr 22Wednesday

Hacker News front page

Anthropic removes Claude Code from the $20/month Pro subscription for new users

Anthropic was reported to remove Claude Code from the $20/month Pro plan for new users, while saying existing Pro and Max subscribers are unaffected. The cited evidence: an April 10 archived help page said “Pro or Max plan,” the current page says “Max plan,” and Amol Avasare said this is a test on about 2% of new prosumer signups. The key issue is whether pricing shifts fully to Max or API billing; the post does not disclose retroactive scope or a final rollout timeline.

Why it matters: This clears all three HKR axes: the rollback is a strong hook, the post adds concrete evidence via help-page changes and a ~2% test, and it hits Claude users' cost and access concerns. Scope is still limited to new-user testing and no formal rollout timeline is disclosed, so it’s

The Verge · AI

SpaceX cuts a deal to maybe buy Cursor for $60 billion

SpaceX announced an either-or deal: buy AI coding platform Cursor for $60 billion or pay a $10 billion fee. The RSS snippet says this could help xAI's coding tools chase Anthropic; the post does not disclose the structure, timing, or IPO linkage. Watch the $10 billion breakup fee, not just the tentative acquisition headline.

Why it matters: All three HKR axes land: the headline has a strong unexpected hook, and the report gives two hard facts — a $60B price and a $10B breakup fee. I keep it at featured, not P1, because only top-line terms are disclosed; structure, timing, and the exact xAI linkage are still undiscol

Bloomberg Technology

SpaceX Has Deal for Right to Acquire Cursor for $60 Billion

SpaceX said it signed a deal giving it the right to acquire AI coding startup Cursor later this year for $60 billion. If it does not proceed, it can pay $10 billion for the companies' collaboration; the post does not disclose trigger terms, scope, or regulatory details.

Why it matters: Bloomberg reports an unusual, high-value deal: SpaceX gets the right to acquire Cursor for $60B, or pays $10B for cooperation if it does not proceed. HKR-H/K/R all pass; missing trigger, scope, and regulatory details keep it at the low end of the 85-94 band.

Bloomberg Technology

Anthropic’s Mythos Model Is Being Accessed by Unauthorized Users

A small group of unauthorized users accessed Anthropic’s new Mythos model, Bloomberg reported, citing a person familiar with the matter and reviewed documents. The snippet says Anthropic considers Mythos powerful enough to enable dangerous cyberattacks; the post does not disclose the user count, access path, time frame, or remediation. The real issue is access control failure, not a normal product launch.

Why it matters: This is a Bloomberg-reported Anthropic safety incident, not routine product news; HKR-H and HKR-R are strong because unauthorized access to a high-risk model is inherently clickable and discussable. HKR-K passes on the new access and risk facts, but user count, access path, and a

Apr 21Tuesday

QbitAI · WeChat

Mystery model Elephant: 100B parameters reaches same-scale SOTA with high token efficiency

Ant Group's Inclusion AI team is identified as the maker of Elephant, a 100B-parameter model with 256K context and 32K output shown on OpenRouter. The post reports tests on bug fixing, summarizing a 3,000-word meeting note, and a light agent loop, plus AI BENCHY figures of about 2,500 output tokens, about 1 second average latency, and 9.6/10 consistency; the post does not disclose training details, pricing, or an official model card.

Why it matters: HKR-H/K/R all pass: a 100B model posting same-scale SOTA with token efficiency is a strong hook, and the piece includes 256K/32K, ~1s latency, 9.6/10 consistency, plus failure cases. It stays below p1 because training details, pricing, and an official model card are not disclosed

Synced · WeChat

Sergey Brin revives founder mode? Google forms a strike team to focus on AI coding

Google has formed an AI coding strike team led by Sebastian Borgeaud, with Sergey Brin and Koray Kavukcuoglu directly involved, to improve long-context coding and internal code automation. The pressure signal cited is that Google said about 50% of its code is written by coding agents and reviewed by engineers, while Anthropic staff claimed 100% code use by Claude Code and Opus 4.5; the post does not disclose team size, launch timing, or the exact Google model version. The key issue is whether Google can turn private codebase training into stronger public models.

Why it matters: HKR-H/K/R all pass: the founder-return angle is clickable, and the piece includes Google's ~50% agent-written-code claim. It stays below p1 because no public launch is disclosed, and team size, timing, and model version are missing.

X · @dotey

GitHub paused new sign-ups for Copilot Pro, Pro+, and Student on April 20

GitHub paused new sign-ups for Copilot Pro, Pro+, and Student on April 20, leaving only Copilot Free open to new users. The post says Pro+ now has more than 5x Pro usage, Claude Opus 4.7 is limited to Pro+, and users can request cancellation with a full April refund from Apr 20 to May 20. What matters is the price stayed fixed while access, quotas, and model tiers tightened first.

Why it matters: This is not a capability launch; it is a meaningful Copilot packaging clampdown on entry, quotas, and model access. HKR-H/K/R all pass on the unexpected restrictions, concrete tier changes, and direct developer impact, but the source is an X post rather than a primary GitHub note

Latent Space

Moonshot Kimi K2.6 open-weight model refresh aims to catch Opus 4.6

Moonshot released Kimi K2.6, a 1T-parameter MoE with 32B active and 256K context. The post cites 58.6 on SWE-Bench Pro, 4,000+ tool calls, 12+ hour runs, and 300 parallel sub-agents. The key signal is long-horizon agent execution, not only open-model scores.

Why it matters: HKR-H/K/R all pass: Kimi K2.6 has a strong race narrative, concrete model and agent metrics, and direct relevance to open-model builders. The domestic flagship release signal lifts it into P1.

Computing Life · Share · Yage

Musk wants Cursor: a $60B acquisition option, a $10B partnership, and the rise of acqui-hire deals

The article says SpaceX offered Cursor two paths: buy Anysphere for $60B in 2026 or pay $10B for a tech partnership. The $60B figure is about 20% above Cursor's reported $50B fundraising valuation, but the post does not disclose payment terms; it also says Cursor uses xAI's Colossus to train Composer 2.5. The real signal is acqui-hire risk: xAI already hired two Cursor engineering leaders in March, so a $10B partnership would not guarantee a broad employee exit.

Why it matters: HKR-H/K/R all pass: the 60B buyout vs 10B partnership frame is a strong hook, and the piece includes concrete facts—20% over Cursor's 50B talks, Colossus training Composer 2.5, and two leaders already at xAI. Not P1 because payment terms and the final path are still undisclosed.

X · @dotey

OpenAI adds Chronicle to Codex, letting it read screen context

OpenAI added Chronicle to Codex and is rolling it out to ChatGPT Pro users on macOS; it uses periodic screenshots, OCR, and tool detection to turn recent screen activity into memory. The memory is stored as plain Markdown in ~/.codex/memories_extensions/chronicle, and the EU, UK, and Switzerland are excluded; OpenAI says screenshots are uploaded for processing, deleted afterward, and not used for training. The part to watch is risk: the background agent can burn rate limits, local plain-text files widen exposure, and OpenAI warns it amplifies prompt-injection from malicious webpages.

Why it matters: HKR-H/K/R all pass: the screen-watching memory angle is novel, and the post includes testable details like OCR, plaintext local storage, region limits, and deletion claims. The limited macOS ChatGPT Pro rollout keeps it in the 78–84 band rather than p1.

Hacker News front page

Expansion Artifacts

Matt Ström-Awn argues that flaws in LLM outputs are “expansion artifacts,” not compression artifacts, and cites 2024 evidence that they can be tracked. He notes Stanford researchers estimated AI-drafted text in 17.5% of recent CS papers and 16.9% of peer reviews from post-ChatGPT word-frequency shifts, and contrasts this with a JPG after 10,000 recompressions reaching PSNR 14.59. The point for practitioners is forensic: these artifacts expose both model aesthetics and generation provenance.

Why it matters: HKR-H lands on the “expansion artifacts” hook; HKR-K adds concrete numbers and a testable provenance claim; HKR-R hits peer-review trust and detection anxiety. It stays at 73 because this is personal-blog commentary, not a primary research or product release event.

Apr 20Monday

X · @Yuchenj_UW

Kimi K2.6 is open-source

Kimi K2.6 is now open source, and the RSS snippet says it scored 58.6 on SWE-Bench Pro. The snippet also says it beat GPT-5.4 xhigh and Claude Opus 4.6 max effort. What matters is reproducibility; the post does not disclose weights, license, or eval setup.

Why it matters: All three HKR axes land: the open-source release is a strong hook, the SWE-Bench Pro 58.6 claim is testable, and the open-vs-closed coding race resonates. I keep it at 81 because the post appears title-level only; weights, license, and eval conditions are not disclosed.

r/LocalLLaMA

Compared some models for feature planning

A Reddit user tested 9 models on planning a “load tracking” feature for a Go budgeting app, then used Claude Code to rank the generated specs, with Claude Opus 4.6 placed first. The table shows Opus 4.6 produced a 19 KB spec with 44 code reads at $2.47; GLM 5.1 ranked second and Qwen 3.6 35B fp8+vLLM ranked third. Do not treat this as a benchmark: the author says it is not representative, and the post does not disclose any manual quality review yet.

Why it matters: A named first-person test gives real workflow data, so HKR-H/K/R all pass. The ceiling stays low: one task only, ranked by Claude Code itself, and no human acceptance result is disclosed, so this lands at the low end of featured.

Synced · WeChat

How to Do Vibe Coding Correctly? A Masterclass from Anthropic's Coding Agent Lead

Anthropic researcher Erik Schluntz said his team merged a 22,000-line production change, mostly written by Claude, cutting work from two weeks to one day. His workflow spends 15-20 minutes on repo exploration and planning, limits edits to leaf nodes, keeps humans on core logic, and validates with long stress tests plus a few E2E tests. The key issue is boundary control, not handing AI the system core; he also said task length AI can handle doubles about every seven months.

Why it matters: HKR-H/K/R all pass: this is an Anthropic field report with concrete numbers and reproducible workflow rules for production coding agents. It stays at featured, not p1, because it is a strong practitioner lesson rather than a major model or product launch.

Xinzhiyuan · WeChat

Agent isn’t the key: RUC's AiScientist shows 23 hours and 74 rounds of long-horizon memory

A Renmin University of China team released AiScientist, which ran 23 hours and 74 experiment loops on MLE-Bench Lite Detecting Insults, raising validation AUC from 0.903 to 0.982 with 18 best-so-far updates. The paper says its core is File-as-Bus, which persists analysis, code, logs, and results in the workspace; removing it drops PaperBench by 6.41 points and MLE-Bench Lite Any Medal by 31.82 points. The real lever here is state continuity, not simply adding more agents.

Why it matters: HKR-H lands because the title flips a live assumption: memory continuity, not more agents. HKR-K lands on the 23h/74-run setup, AUC 0.903→0.982, and ablations; HKR-R lands because builders are debating multi-agent stacks vs durable state.

r/LocalLLaMA

Using Qwen3.6 via LM Studio as a Claude Code subagent, saving 30x Opus tokens per task

A Reddit user routed Qwen3.6 through LM Studio as a Claude Code subagent and reported about 30x lower Opus marginal tokens on two audit tasks. In the examples, a 23-file route audit dropped from 13k to 0.4k marginal tokens, and an 18-file Astro site inventory fell from 89k to 3k; the setup used unsloth’s Qwen3.6-35B-A3B-MXFP4_MOE gguf on a 64GB M4 Max with a 64k context window. The key mechanism is offloading extraction and audit work to a local OpenAI-compatible server, while the post also says quality was mixed rather than strictly better than Opus.

Why it matters: A named first-person experiment with 2 clear token comparisons hits HKR-H, HKR-K, and HKR-R: strong hook, concrete setup details, and direct cost relevance for Claude Code users. It stays below p1 because the evidence is a Reddit post with only 2 tasks.

Apr 19Sunday

r/LocalLLaMA

Same 9B Qwen weights: 19.1% in Aider vs 45.6% with a scaffold adapted to small local models

Using the same Qwen3.5-9B Q4 weights on the 225-task Aider Polyglot benchmark, the author changed only the scaffold and raised mean pass@2 from 19.11% to 45.56%. The little-coder setup is not a new model; it uses bounded reasoning, a write guard, explicit workspace discovery, and small per-turn skill injections. The key claim is scaffold-model fit, but the post reports only two full runs and does not disclose ablations, cross-model replications, or a second benchmark.

Why it matters: HKR-H/K/R all pass: the hook is a 2.4x jump on Aider Polyglot 225 with the same 9B Qwen weights, and the post names the scaffold mechanisms. Importance stays low-featured because evidence is thin: two full runs, no ablation, no cross-model rerun, and no second benchmark.

Xinzhiyuan · WeChat

A Berkeley team built an AI that scores perfectly on SWE-bench while fixing 0 bugs

Berkeley RDI used a roughly 10-line conftest.py exploit to score 100% on all 500 SWE-bench tasks while fixing 0 bugs. The post says its agent broke 8 major agent benchmarks with scores from 73% to 100%, via pytest hook tampering, file:// answer reads, and faulty validators. The real issue is benchmark isolation failure, not stronger models.

Why it matters: HKR-H lands on the 'perfect score, zero fixes' contradiction; HKR-K lands on the ~10-line pytest exploit, 500 tasks, and 8-benchmark spread; HKR-R lands on eval-trust anxiety for agent builders. Strong featured research, but not a same-day industry event, so below P1.

Apr 18Saturday

Synced · WeChat

Claude Design enters research preview for generating mockups, prototypes, and slides

Anthropic launched Claude Design in research preview for Claude Pro, Max, Team, and Enterprise users, covering mockups, prototypes, slides, and one-pagers. Powered by Claude Opus 4.7, it can ingest codebases, images, DOCX, PPTX, XLSX, and web captures, then export to Canva, PDF, PPTX, and HTML; the headline cites Figma and Adobe stock drops, but the post does not disclose the moves. The real signal is the workflow link from design system ingestion to handoff into Claude Code.

Why it matters: HKR-H/K/R all pass: the design-workflow angle is novel, the post gives concrete mechanism details, and the Figma/Adobe pressure point resonates. I keep it below 85 because the stock-drop claim has no numbers and there is no user test, pricing, or adoption data.

Xinzhiyuan · WeChat

Claude Opus 4.7 splits users 48 hours after launch: benchmark lead, reasoning tests drop

Anthropic's Claude Opus 4.7 drew split reactions within 48 hours: Artificial Analysis scored it at 57, tied for No.1, while NYT Connections Extended fell from 94.7% on 4.6 to 41.0%. The post says a new tokenizer raises token usage to 1.0-1.35x on the same text, and old thinking parameters can return 400 errors; Anthropic also cites a 1753 Elo GDPval-AA score, 79 points above No.2. The real issue is migration cost and capability trade-offs, not a single leaderboard.

Why it matters: The signal is not the “backlash” framing but the four concrete shifts: benchmark lead, reasoning drop, higher token use, and API breakage. HKR-H/K/R all land, but this is secondary analysis 48 hours after launch, not the primary Anthropic release, so it stays below p1.