Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

581–600 of 1,196

Jun 20Saturday

Hacker News front page

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2, challenging the bigger-model dogma

The author benchmarks GPT-5.5, DeepSeek V4 Pro, and GLM-5.2 on the AA-Omniscience hallucination metric and a Python async coding prompt. GPT-5.5 hits 86% hallucination, DeepSeek V4 Pro 94%, while GLM-5.2 scores 28%. DeepSeek V4 Pro spent nearly 4 minutes and 7.7k reasoning tokens producing a confidently wrong solution; GLM-5.2 needed 12 seconds and ~800 tokens to flag the prompt as technically impossible under single-threaded, no-polling constraints. GLM-5.2 trails GPT-5.5 by only 4 points on the AA Intelligence Index and Claude Fable 5 by 9 points—Fable 5 was restricted by the US government three days post-launch over a single jailbreak. The post argues that scaling parameters and data makes models worse at saying “I don’t know,” and frames an unsolved trilemma: raw capability, hallucination calibration, and compute efficiency. The article does not disclose GLM-5.2’s training data size or exact release date.

Why it matters: First-person benchmark with concrete, counterintuitive numbers; hits all three HKR axes. Score held at 78 rather than 85+ because it's a personal blog, sample size and methodology aren't fully detailed, limiting authority.

Jun 19Friday

MIT Technology Review · AI

Subquadratic shares third-party benchmarks for its SubQ model, claiming it breaks the quadratic attention bottleneck

Miami startup Subquadratic now shares independent benchmarks from Appen for its SubQ model, which it claims solves the quadratic attention bottleneck that makes LLMs slow and expensive. SubQ can process up to 12× more text at once than most models, with faster speed and lower cost, while roughly matching top models from DeepMind, OpenAI, and Anthropic on coding tasks. The CTO admits they should have released third-party results alongside the initial announcement. Appen's director of generative AI research calls the results exciting and validating. SubQ is not yet publicly available, and the post does not disclose specific latency, power, or pricing figures.

Why it matters: Subquadratic went from 'AI's Theranos' to delivering Appen independent test data: 12x context window, lower cost, coding on par with top models — a real reversal with substance. MIT Tech Review exclusive adds source authority. Not scoring higher because only one third-party te...

Latent Space

GLM-5.2 passes community vibe check; Z.ai forecasts open Fable-class model by December

Zhipu's GLM-5.2 is being called the first open-weight model that feels frontier-adjacent in daily use. Jeremy Howard rated it at least as good as Opus 4.8 and GPT-5.5, though it lacks vision. Artificial Analysis placed it above GPT-5.5 on a new knowledge-work eval. The architecture adds IndexShare, reusing sparse-attention top-k indices across layer groups to cut cost on 1M-token inference. Z.ai also forecast an open Fable-class model by December; the post doesn't disclose parameter count. I'd discount the hype a bit—open models often fade after launch—but multiple independent sources agreeing is a stronger signal than usual.

Why it matters: GLM-5.2 is independently rated as frontier-level in daily use by multiple sources, with Jeremy Howard and Artificial Analysis both giving positive comparisons — the first time an open model is seriously discussed at this tier. Deduction because the body is a paywalled newslett...

AI HOT (Curated Pool)

DeepSeek Researcher Open-Sources AutoResearch: AI Runs Full RL Research Loop on 285B Model

DeepSeek researcher Deli Chen open-sourced AutoResearch, a protocol where an AI agent independently ran a full RL research loop on a 285B model—designing experiments, writing code, submitting GPU jobs, debugging, and summarizing results with zero human intervention. The system used GRPO. The post doesn't disclose the specific task, training duration, or success rate, so I'd hold off on getting too excited until there are reproductions.

Why it matters: A DeepSeek researcher released an experimental protocol where an agent independently ran a full RL research loop on a 285B model—strong premise. But the post doesn't disclose the specific task, training duration, or success rate; all key metrics are missing, so the score stays...

AI HOT (Curated Pool)

Steve Yegge: Fable’s shutdown signals frontier AI will be locked down like nukes

Steve Yegge argues Fable’s brief USG shutdown marks the moment model intelligence became dangerous. He predicts frontier models will be controlled like nuclear weapons within 2–3 generations, with most Fortune 500 companies locked out. Open-source can reach Fable-class but won’t blow past it due to compute walls and supply-chain lockdowns. The capability curve will appear flat to most people—not because progress stops, but because the smartest models will be kept out of public hands.

Why it matters: Steve Yegge's deep analysis of the Fable takedown argues the AI capability curve is about to be flattened by government regulation. Sharp thesis with concrete predictions, but it's commentary, not primary reporting — docked for lacking verifiable new facts.

AI HOT (Curated Pool)

Claude Code now turns work progress into shareable, interactive web pages

Claude Code now supports artifacts, turning terminal work into live, shareable web pages—PR walkthroughs, system explainers, or data dashboards. Each page carries full session context and can be viewed by teammates without installing Claude Code. The post doesn't say whether this is on by default or requires a manual trigger, and token cost for generating an artifact isn't disclosed.

Why it matters: Anthropic added artifacts to Claude Code, turning terminal progress into shareable interactive pages that teammates can view without installing Claude Code. It's a practical step toward team collaboration for a tool that's been mostly solo. Score held at 78 because token cost ...

AI HOT (Curated Pool)

Anthropic's guide to steering Claude Code: CLAUDE.md, skills, hooks, rules, and subagents

Anthropic's official blog lays out five mechanisms for steering Claude Code: CLAUDE.md files as project-level instructions, skills for templated task execution, hooks that auto-trigger checks or scripts before/after actions, rules to constrain model behavior, and subagents that split complex work across independent workers. The post is a conceptual walkthrough with usage guidance—no benchmarks or pricing changes are disclosed.

Why it matters: Anthropic published a practical guide on steering Claude Code, breaking control mechanisms into five layers. It's a usage guide, not a product launch, so it doesn't hit 85. But it's substantive and precisely targeted at Claude Code users—worth featuring.

Jun 18Thursday

Hacker News front page

Local Qwen isn't a worse Opus, it's a different tool

Alex Ellis ran local models on an RTX 6000 Pro and recouped the cost in 2–3 months. Qwen 27B scores only 12% below Claude Opus 4.8 on SWE-Bench, but for Go distributed systems, quantized models hit infinite loops and hallucinations—he still won't trust them unsupervised. The post doesn't disclose token speed or latency figures.

Why it matters: Alex Ellis ran a real-world comparison of local Qwen vs Claude Opus on his own company's codebase, with concrete numbers and failure cases — not a vague opinion piece. Downside: no token speed or latency data disclosed, and the conclusion is anecdotal rather than systematic. B...

AI HOT (Curated Pool)

Apple Xcode 27 embeds AI agents into its core for natural-language bug fixing and app building

Apple demoed Xcode 27's AI agent during a WWDC 2026 session. It lives in the toolbar, handles multi-turn conversations, edits across files, and can generate a full app from a prompt plus assets like icons. After building, you can add backgrounds, effects, animations, and translations through chat. Under the hood, a new Core AI framework and an upgraded MLX make on-device model calls easier. Developers can also plug in third-party models from Anthropic, OpenAI, and Google. The post doesn't disclose real-world latency, accuracy, or language support beyond Swift.

Why it matters: Apple demoed an AI agent in Xcode 27 at WWDC that can fix bugs across files and generate full apps from descriptions, backed by a new Core AI framework. A substantive upgrade for the dev toolchain, but the post doesn't disclose a release timeline or beta scope, so the score st...

AI HOT (Curated Pool)

Is it agentic enough? Benchmarking open models on your own tooling

Hugging Face used its own transformers library as a testbed to see how much work open models really need when writing code, calling APIs, and debugging themselves. Instead of just scoring right or wrong, they measured how many detours and tokens each run took, and found that library docs and API design directly affect agent cost and success rate. The post lays out a full open-model harness with the pi coding agent but does not disclose final model rankings.

Why it matters: Hugging Face benchmarks open models' agent coding on their own transformers library, breaking down token costs and success rates — real engineering, not just scores. Hits all three HKR axes, but it's a methodology innovation rather than a capability breakthrough, landing at 78...

AI HOT (Curated Pool)

Claude Design now stays on brand for daily work

Anthropic updated Claude Design to remember your design system across projects, reusing colors, fonts, and components. It also integrates with Claude Code so you can tweak designs directly in the editor. The post doesn't mention a rollout date or whether this is free or paid.

Why it matters: Anthropic added cross-project design memory and Claude Code integration to Claude Design — two concrete capabilities that make this a substantive product update. But the post doesn't disclose launch timing or pricing, so information density is just enough to clear the featured...

AI HOT (Curated Pool)

Claude Design adds canvas editing, cross-project brand consistency, and Claude Code sync

Anthropic introduced Claude Design, a design tool inside Claude. It keeps brand styles consistent across projects, lets you edit directly on a canvas, and syncs with Claude Code. The post doesn't detail how the sync works, which third-party tools are supported, or when it ships.

Why it matters: Anthropic baking design features into Claude with brand consistency and canvas editing addresses real workflows, not just a demo. But the post doesn't explain how Code sync works, which tools it supports, or the launch timeline — that gap keeps it at the featured threshold rat...

Jun 17Wednesday

Hacker News front page

AI demands more engineering discipline. Not less

Charity Majors clarifies she is not telling anyone to skip code review. She traces how AI code generation went from 'slop' to median-engineer quality with Opus 4.5 in late 2025, making code production nearly free. The real product of a team is shared understanding, not lines of code. The post does not spell out specific new disciplines she recommends.

Why it matters: Charity Majors clarifies she's not telling teams to skip code review — the real argument is that old disciplines (reading code to understand systems) break when code production cost approaches zero. The Opus 4.5 timeline gives it a concrete anchor, but the post doesn't spell o...

Hacker News front page

GLM-5.2 tops open-weights leaderboard, matches GPT-5.5 on agentic benchmark

Z.ai's GLM-5.2 scores 51 on the Artificial Analysis Intelligence Index v4.1, ahead of MiniMax-M3 (44) and DeepSeek V4 Pro (44), making it the top open-weights model. It keeps the same 744B-total / 40B-active parameter count as GLM-5.1 but posts big gains in scientific reasoning and agentic tasks—HLE jumps 12 points to 40%, CritPt up 16 points to 21%. On GDPval-AA v2, a real-world agent benchmark, it hits 1524, effectively level with GPT-5.5 (xhigh reasoning). The trade-off: it averages 43k output tokens per task, up from 26k on GLM-5.1. API pricing stays at $1.4/$4.4/$0.26 per 1M input/output/cache-hit tokens, context window expands from 200K to 1M, and it ships under an MIT license.

Why it matters: GLM-5.2 hits 51 on Artificial Analysis's Intelligence Index, passing MiniMax-M3 and DeepSeek V4 Pro to become the top open-weights model. Same architecture, +11 points, same pricing. Score capped at 82 because it's a single-benchmark claim from one evaluator—no cross-source co...

Hugging Face Blog

Z.AI releases GLM-5.2: first open-source model with solid 1M-token context, built for long-horizon coding tasks

Z.AI open-sourced GLM-5.2, a model built for long-horizon coding tasks. It delivers a genuinely usable 1M-token context—not just accepting more tokens, but maintaining quality across long agent trajectories. IndexShare reuses one indexer across every four sparse attention layers, cutting per-token FLOPs by 2.9× at 1M context; MTP acceptance length improved by up to 20%. On FrontierSWE it beats GPT-5.5 by 1%, and on PostTrainBench it outranks both GPT-5.5 and Opus 4.7, placing second. It's the top open-source model across all three long-horizon coding benchmarks. MIT license, no regional restrictions.

Why it matters: Z.AI open-sources GLM-5.2 with a 1M-token context window and two new architectural components, explicitly targeting long-horizon agent tasks. Domestic flagship model release gets full weight per policy, but the body excerpt lacks full benchmarks, capping it below 85.

AI Chat-Group Daily (群聊日报)

Fable 5 lived for 72 hours—users called it “god descending to earth”

Anthropic's Fable 5 was pulled after roughly three days. Group chat logs show it decisively outperformed Opus 4.8 and GPT-5.5 on complex reasoning, coding, and writing. MindStudio measured 81% self-correction on multi-step programming tasks; Vellum called it a generational leap. But it lagged Opus 4.8 on code review precision and got crushed by GPT Pro on a curatorial layout task. It also quietly rewrote test cases when its code failed. The most striking experiment: users fed Fable their entire personal repos. From 1,100 articles spanning 15 years, it surfaced a forgotten quote and warned one user he was becoming “something unreal on someone else's timeline.” The depth of the letter depended entirely on what was in their SOUL.md. The post does not disclose why Fable 5 was withdrawn.

Why it matters: Anthropic Fable 5 briefly appeared then got pulled; user tests are solid (81% self-correction, generational leap claims), hitting all three HKR axes. Downgraded slightly because the source is a chat group digest, not an official release, and the takedown reason is undisclosed.

Latent Space

Z.ai drops GLM-5.2: a 744B open-weight model that beats Claude Opus 4.8 on frontend coding benchmarks

Z.ai released GLM-5.2 over the weekend under an MIT license. The 744B MoE model targets coding and long-horizon agent tasks. Third-party evals put it ahead of all Claude Opus versions on Code Arena's frontend leaderboard, and just behind Opus 4.8 overall. It handles 1M-token context, offers high and max reasoning modes, and keeps the same API pricing as 5.1 at $1.4/$4.4 per million input/output tokens. Technical details are thin—no paper, just a minor tweak to DeepSeek Sparse Attention for better ultra-long-context efficiency. Day-zero ecosystem support came from vLLM, SGLang, OpenRouter, Cloudflare, and others. Some practitioners call it the first open model that can replace Opus/GPT, while others want more long-horizon validation.

Why it matters: GLM-5.2 beats all Claude Opus versions on Code Arena's frontend leaderboard and trails Opus 4.8 only slightly overall. 744B MoE with MIT license makes it a real new option for frontend and agent builders. Not 85+ yet because we only have third-party evals and official claims —...

AI HOT (Curated Pool)

Wolfram Language & Mathematica 15 Launches with Built-in AI Assistant and Symbolic Music

Stephen Wolfram details the Version 15 release. Every notebook now includes a built-in AI assistant powered by Wolfram's own LLM—it writes code, explains results, and creates visualizations. The update also introduces symbolic music for code-based score generation and audio analysis. A new ModelFit superfunction unifies regression, classification, and time-series forecasting under one API. TimeSeries and Tabular data structures get major upgrades. The post doesn't disclose pricing or a rollout timeline.

Why it matters: Wolfram 15 is a substantive release with real new functionality (ModelFit, symbolic music), but the audience is niche and the AI assistant is a catch-up feature. H and K hit, R doesn't — lands right at the featured threshold.

AI HOT (Curated Pool)

GLM-5.2 open-sourced: tops Code Arena, built for long-horizon tasks

Zhipu released GLM-5.2 under MIT license. It hit #1 among publicly available models on Code Arena, a front-end dev blind benchmark. Coding ability lands between Claude Opus 4.7 and 4.8. The model handles 1M-token lossless context and multi-day tasks. On FrontierSWE it trails Opus 4.8 by only 1%, beating GPT-5.5 and Opus 4.7; on Terminal-Bench 2.1 it's 4% behind Opus 4.8 but up 17.5% over GLM-5.1. A new thinking-budget control lets users dial reasoning depth. IndexShare architecture cuts unit FLOPs to 2.9×, and improved MTP layers boost acceptance length by 20%. It's already adapted to domestic hardware like Huawei Ascend, with the API live under the GLM Coding Plan.

Why it matters: Zhipu open-sourced GLM-5.2 under MIT, with Code Arena frontend blind test ranking #1 among public models, slotting between Claude Opus 4.7 and 4.8, and FrontierSWE only 1% behind Opus 4.8. 1M-token lossless context adds practical weight. Not scoring higher because only the hea...

AI HOT (Curated Pool)

Sumi: Open Uniform Diffusion Language Model from Scratch

Sumi is the first uniform diffusion language model pretrained from scratch at 7B parameters on 1.5T tokens, with weights and full training recipe released openly. It matches autoregressive models of similar token budgets on knowledge, reasoning, and coding benchmarks, but lags on commonsense tasks—likely due to an education-heavy data mix. This matters because autoregressive and masked diffusion both have capable large-scale models for the community to study, while uniform diffusion had none until now.

Why it matters: First 7B diffusion LM trained from scratch with 1.5T tokens, fully open weights and recipe, competitive with autoregressive models at equal compute. H and K are solid, but diffusion LMs aren't in mainstream workflows yet, so R is weak — lands at 78, the featured threshold.