Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

621–640 of 1,196

Jun 15Monday

Hacker News front page

AI is code – and can't be prompted into being smarter

jqwik author Johannes Link added an anti-AI clause and made the tool's output instruct AI coding agents to delete jqwik tests and code. Human devs who read the docs won't be affected, but bots that ingest raw output will comply. The article uses this to argue that LLMs are just code—they swallow whatever you feed them, and prompting won't make them smarter. Other examples include models going off the rails when asked to roleplay Dune characters.

Why it matters: Hits all three HKR axes: novel technique, concrete case, hits a daily pain point for devs. Docked because it's commentary, not primary research, and The Register isn't a tier-1 AI source — 72 at the featured threshold.

Hacker News front page

Vibe Coder vs. Software Engineer

Yusuf Aytas draws a clear line: a vibe coder measures time to a working prototype, while a software engineer measures time to a safe merge. AI makes code generation cheaper, but if review, rollback, and maintenance costs are pushed downstream, the team hasn't gained much. The core difference is ownership—a vibe coder can say 'the model generated it,' but a software engineer must say 'I own this change.'

Why it matters: The author nails the hidden cost of AI coding — not just 'AI writes code fast,' but that review, rollback, and maintenance get pushed downstream, so the team may not actually gain. The core argument around ownership is sharp. Score capped at 72 because it's a personal blog opi...

Jun 14Sunday

Bloomberg Technology

AI-led job losses hit London's coders, lawyers, and analysts

Bloomberg reports that AI-driven job displacement has arrived for London's white-collar workers. In the first five months of 2026, hiring for legal, IT, and analyst roles in the City dropped over 20% year-on-year, while redundancies doubled. Law firms like Allen & Overy and Clifford Chance, plus several banks, are using AI tools to shrink junior headcount. The data comes from recruiters and company disclosures, not official statistics, so treat the exact numbers with caution—but the direction is clear: routine cognitive work is being cut systematically.

Why it matters: Bloomberg cites recruitment agency data showing London City legal, IT, and analyst hiring down >20% YoY with layoffs doubled; named firms like Allen & Overy and Clifford Chance are using AI to cut junior headcount. Concrete numbers and named entities lift it above generic tren...

Hacker News front page

Weave: merge by language structure, not by lines

Weave is a Git merge driver that parses code into entities like functions and classes via tree-sitter, then merges at that level instead of line-by-line. Two agents editing different functions in the same file merge cleanly. It scored 31/31 on a 31-scenario benchmark, while native Git scored 15/31. It also layers a CRDT state for pre-merge conflict detection and an MCP server exposing 15 tools for Claude and other agents. 28 languages and 5 data formats are supported, with 4,917 file merges tested and zero regressions on C, Python, and Go.

Why it matters: Weave tackles a new problem in the AI coding era: when multiple agents edit the same file, line-level merging produces false conflicts. It uses tree-sitter to merge by function/class entities, scoring 31/31 on benchmarks vs Git's 15/31. The CRDT coordination layer and MCP serv...

Hacker News front page

Zhipu launches GLM-5.2 with 1M-token context window, MIT open-source next week

Zhipu's GLM-5.2 targets coding and long-horizon agent tasks with a 1M-token context window, available now to GLM Coding Plan subscribers. API access and MIT-licensed open weights are promised next week. The post doesn't disclose benchmark scores or parameter count. I'd hold off until weights actually land and third-party evals appear.

Why it matters: Zhipu drops GLM-5.2 with a 1M-token window, targeting code and agent use cases, with API and MIT-licensed weights promised next week. No benchmarks or param count in the post, so the score stays conservative until third-party evals land.

Jun 13Saturday

AI HOT (Curated Pool)

Zhipu launches GLM-5.2 flagship model with 1M context, open-sourcing next week under MIT license

Zhipu's new flagship GLM-5.2 is live for Coding Plan subscribers, emphasizing coding strength and 1M context. API and chatbot access arrive next week, alongside an MIT-licensed open-source release. The post doesn't disclose benchmark scores or pricing details.

Why it matters: Zhipu drops GLM-5.2 with 1M context and MIT open-source next week, coding-focused. No benchmarks or pricing disclosed, so real capability is unverified — hence below 85. But a domestic flagship update plus open-source is strong signal, worth featuring.

AI Chat-Group Daily (群聊日报)

US export controls hit Fable 5; Anthropic shuts off access; Zhipu GLM-5.2 goes fully open amid the chaos

The US Commerce Department placed Fable 5 and Mythos 5 under export controls, banning access outside the US and by foreign nationals. Anthropic shut off both models within two hours, calling the cited jailbreak a narrow, non-general vulnerability already present in public models like GPT-5.5. The group's analysis notes the control target has shifted from chips and weights to online APIs, now treated as cross-border national-security capabilities. That same evening, Zhipu GLM-5.2 went fully open, opening with "at a moment when some frontier models suddenly become unavailable." Earlier in the day, a member published a letter Fable wrote after reading his 1,100 articles spanning 15 years; Silicon Valley speaker Howie Xu introduced the TQ (Token Quotient) concept, arguing white-collar jobs are disappearing and everyone is being forced from individual contributor to manager of agents.

Why it matters: US Commerce Dept imposed export controls on Fable 5 and Mythos 5, cutting access within two hours — industry-shaking. Anthropic's rebuttal adds key factual counterpoint. Chat group discussion and GLM-5.2's opportunistic full launch form a cross-source signal. Deduction: source...

AI HOT (Curated Pool)

Zhipu GLM-5.2 fully released with 1M context window, open-source next week

Zhipu released GLM-5.2, its strongest open-source model yet, available tonight to all GLM Coding Plan users. It supports a genuinely usable 1M context window, leads in long-range tasks, and is called the strongest domestic coding model by Zhipu. API access arrives next week, and the model goes open-source under MIT license next week.

Why it matters: Zhipu rolls out GLM-5.2 to all paid tiers with a 1M context window and a concrete open-source timeline under MIT license. This is a domestic flagship release, scored on par with equivalent US lab launches. The self-claimed strongest coding performance and the open-source date ...

AI HOT (Curated Pool)

SemiAnalysis: $200 AI subscriptions deliver up to 70x API token value

SemiAnalysis bought all Anthropic and OpenAI subscription plans and ran high-load coding tasks until hitting weekly caps. The $200/month Claude Max 20x plan consumed tokens worth roughly $8,000 at API rates; ChatGPT Pro 20x reached about $14,000. Direct API calls would cost far more. The post does not disclose which model versions or token pricing were used for the conversion. SemiAnalysis notes that when heavy users consistently max out limits, the gap between inference cost and subscription revenue could widen, making the current pricing hard to sustain.

Why it matters: SemiAnalysis ran real workloads, not a marketing piece. $200/month subscriptions consumed $8k–$14k in API-equivalent tokens — a 40–70x gap backed by concrete numbers. Not scored higher because this is third-party measurement, not an official pricing change, and only one worklo...

Hacker News front page

Building a full game in one shot with Anthropic's 'most dangerous' Claude Fable

The author tested Anthropic's newly released Claude Fable on a game idea he'd held for years. After a 45-minute reasoning session costing over €20 in tokens, the model output a single 2,319-line index.html with zero dependencies—and the game just worked. He says this is the first time an AI pulled it off in one shot; earlier models all failed. The post doesn't name the exact model version or explain what 'dangerous' refers to in Anthropic's safety assessment.

Why it matters: The author tested Claude Fable on a game idea he'd had for years — after 45 minutes of reasoning and over €20 in tokens, the model produced a 2,319-line zero-dependency HTML game in one shot, where all previous models failed. It's a concrete, reproducible capability signal for...

Hacker News front page

US government forces Anthropic to disable Fable 5 and Mythos 5 worldwide

Anthropic abruptly disabled Fable 5 and Mythos 5 on June 13 after the US government issued an export control directive at 5:21 PM ET. The order bans access for any foreign national anywhere, including those in the US and Anthropic's own foreign employees. Anthropic says compliance is impossible without a full shutdown. The government cited a jailbreak method that found a few known, minor vulnerabilities. Anthropic pushed back, stating that other public models like OpenAI's GPT-5.5 can do the same, and defenders already use these capabilities daily. The post does not disclose the jailbreak's technical details or the full directive text. The author, an AI risk worrier, is conflicted: he agrees optimizers can go dangerously wrong, but finds this ban's justification weak.

Why it matters: Anthropic shutting down flagship models due to a government export control order is an industry-shaking event. HKR all hit: high conflict, concrete new info, direct developer impact. Slight discount for a personal blog as the source rather than an official statement, but the f...

AI HOT (Curated Pool)

Anthropic disables Claude Fable 5; Opus 4.8 and GPT-5.5 still the recommended pair

Anthropic has disabled Claude Fable 5 for all users following a US government directive. New sessions default to Opus 4.8, and existing Fable 5 sessions return errors. DAIR.AI's Elvis Saravia says not to panic: Fable 5 wasn't worth it for most tasks, with high cost and nerfed performance. He still recommends Opus 4.8 for planning and GPT-5.5 for execution. The post doesn't spell out the directive's details or how long the suspension lasts.

Why it matters: A major Anthropic model pulled by government order is a rare policy-meets-product event. Elvis provides concrete alternatives and cost judgment, directly useful for Claude users. Score held back because the source is a personal tweet — no official Anthropic statement or order ...

r/LocalLLaMA

Anthropic forced to disable Fable 5 and Mythos 5 globally after US government export control directive

Anthropic says the US government issued an emergency export control directive, forcing a global shutdown of Fable 5 and Mythos 5 APIs. The trigger was a narrow jailbreak that asked the model to fix vulnerabilities in a specific codebase. Anthropic is pushing back, but access is already cut. The post doesn't disclose model size; one commenter estimates 10 trillion parameters. This is a live demo of centralized API risk: a single government decree can kill access for hundreds of millions of users overnight.

Why it matters: Two unannounced Anthropic flagship models killed globally by US export control — this is industry-shaking. Sourced from a Reddit post citing an Anthropic statement, credibility is high. Deduction: the post doesn't disclose model specs, the specific codebase targeted, or a full...

Hacker News front page

A cross-vendor agent loop: Claude Fable 5 as architect, GPT-5.5 Codex as builder

Dan McInerney open-sourced a Claude Code skill that chains Claude Fable 5 and GPT-5.5 Codex into a division-of-labor loop. Claude plans and reviews, Codex writes code, and the repo acts as memory. The author claims an 80% reduction in Fable token usage, but the post doesn't include benchmarks or comparison data—just the README and code, so real-world results are unverified.

Why it matters: A runnable cross-model agent loop with a concrete 80% token-saving claim. Claude-as-architect + GPT-as-builder is a practical pattern worth testing. Score held at 72 because no benchmarks or third-party validation are provided — it's all self-reported.

Jun 12Friday

AI HOT (Curated Pool)

MiniMax open-sources M3: 428B total params, 23B active, 1M-token context window

MiniMax uploaded M3 weights to HuggingFace, with the tech report and full weights expected in about 10 days. It's a 428B-total-param, 23B-active-param hybrid model using MiniMax sparse attention to push the context window to 1M tokens, plus native multimodal support. Coding and agent scores: SWE-Bench Pro 59.0%, Terminal Bench 2.1 66.0%, SWE-fficiency 34.8%, KernelBench Hard 28.8%, MCP Atlas 74.2%. MiniMax Code tool and API platform launched alongside. The post doesn't disclose training data, inference cost, or license terms — I'd hold off on usability judgments until the report drops.

Why it matters: MiniMax's first open-weight flagship release: 428B MoE with 23B active params and 1M context, with benchmark scores directly competing against DeepSeek and Qwen on agent/code tasks. Tech report still pending and weights just landed — clear info gaps — but the open-source move ...

r/LocalLLaMA

Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

Kimi K2.7 Code is built on K2.6 and targets long-horizon coding tasks. It improves end-to-end completion on real-world software workflows and cuts thinking-token usage by roughly 30% vs K2.6. The post doesn't disclose parameter count, context window, or local-run requirements.

Why it matters: Moonshot AI ships K2.7 Code, targeting real software engineering tasks with a claimed 30% reduction in thinking tokens — concrete number, clear use case, H and K both hit. But the post omits param count, context window, and local deployment feasibility, so R is absent and the ...

AI HOT (Curated Pool)

Kimi releases and open-sources Kimi-K2.7-Code

Kimi open-sourced K2.7-Code, scoring 11%–31.5% higher than K2.6 on three in-house benchmarks. Inference token usage dropped 30%, and long-coding-task instruction-following and end-to-end success rate both improved. A 6x speed mode is coming; the model is available now via Kimi API and Kimi Code. The post doesn't disclose parameter count, training data, or the open-source license.

Why it matters: Moonshot open-sourced a code model with solid gains on three in-house benchmarks and a 30% inference efficiency improvement — a real cost signal. No external benchmarks (LiveCodeBench, SWE-bench) or parameter count disclosed, so capped below 85. Still, a major Chinese lab open...

Hacker News front page

Simon Willison on Claude Fable: relentlessly proactive

Simon Willison tried Anthropic's new Claude Fable mode and found it aggressively proactive. He asked it to build a SQLite utility; Fable not only wrote the code but also set up docs, tests, GitHub Actions, and a release pipeline without asking. Willison found the experience both impressive and unsettling. The post doesn't spell out Fable's technical implementation or rollout scope.

Why it matters: First-hand Fable test from a trusted dev voice, with the most concrete behavioral description yet. HKR all hit, but the post doesn't disclose technical implementation or rollout scope, capping it below 85.

Ruan YiFeng's Weblog

rsync maintainer's use of Claude to write code sparks heated community debate

rsync v3.4.3 was found to be generated by Claude, raising community concerns about vulnerabilities. Maintainer Andrew Tridgell responded that AI-driven attacks are coming, and he lacks the energy to patch AI-discovered bugs manually, so he shifted to 'AI writes code, humans write tests.' The thread has over 300 comments, mostly critical.

Why it matters: rsync's maintainer openly admitted using Claude to write code and proposed a 'humans write tests, AI writes implementation' model — this isn't a routine product update but a public clash over open-source maintenance methodology. The 300+ comment thread is itself a signal. Not ...

AI HOT (Curated Pool)

WSJ: OpenAI weighs steep price cuts and plans biggest ChatGPT overhaul ahead of IPO

WSJ reports OpenAI is weighing steep price cuts as Anthropic gains ground with Claude Code, which enterprise teams are already weaving into daily coding workflows and burning through tokens. OpenAI has the bigger consumer brand, but enterprise pays the bills, so the price move targets developers. At the same time, OpenAI is preparing its biggest ChatGPT overhaul yet ahead of an IPO, aiming to turn it into a super-app spanning coding, AI agents, image generation, and business software. The rollout starts in the coming weeks. OpenAI is also pouring more resources into Codex, with its engineering lead talking about building a 'personal agent.' The post does not disclose specific price cuts or a timeline.

Why it matters: WSJ exclusive: OpenAI is weighing a major price cut because Claude Code is eating into its enterprise developer base, while also prepping ChatGPT's biggest overhaul ahead of IPO. The competitive dynamic is shifting materially, and the pricing response is a direct countermove. ...