Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

261–280 of 1,196

Aug 14Friday

Hacker News front page

GLM-5.3: Post-training-only gains push open-weight coding and exploit capability to the top

Z.ai released GLM-5.3 with the same base model as 5.2 — every gain is from post-training. Coding jumped 50% on their internal Z.ai Code Bench, and Terminal Bench 3.0 went from 4.6 to 28.3. The bigger surprise: exploit capability grew far faster than expected. ExploitGym 2h score rose from 29 to 105, 6h from 39 to 130. The team credits training environments that mirror real expert workflows, pushing the model to chain full exploit sequences. Weights will be open-sourced in two weeks after safety hardening.

Why it matters: Zhipu releases GLM-5.3 — same base model as 5.2, all gains from post-training. Code bench up 50%, Terminal Bench from 4.6 to 28.3, 2-hour exploit score from 29 to 105. The lab admits cyber capability emerged faster than expected. Domestic flagship model launch with concrete nu...

AI HOT (Curated Pool)

DeepSeek V4 Pro lands on SiliconFlow with 1M context and three inference tiers

DeepSeek V4 Pro is now available on SiliconFlow with Day-0 support, a 1M context window, and three inference intensity levels. It targets coding, tool use, and agent workflows under the MIT license. Pricing: $1.32/M input, $3.96/M output, $0.44/M cache hit. A Flash variant is also live for cost-sensitive production use. The post does not disclose parameter count or architecture details.

Why it matters: DeepSeek V4 Pro lands on SiliconFlow day one with 1M context, tiered reasoning, MIT license, and clear pricing — solid signal density. Held below 85 because this is a platform availability announcement without benchmarks or user reports yet; sits right at the featured threshold.

AI HOT (Curated Pool)

Claude takes over app maintenance, opens 388 PRs in weeks

Boris Cherny had Claude handle routine app maintenance via Slack—fuzz testing, deduplicating code, removing dead code. It opened 388 PRs in weeks; 180 were merged after Claude code review and human approval. Claude usually got it right in one shot; when it didn't, tweaking the routine fixed it the next day.

Why it matters: First-person experiment by Boris Cherny with concrete numbers and a reproducible workflow — not marketing fluff. Claude handling maintenance isn't industry-shaking, but the 388-PR scale makes it stand out among similar experiments. Not scored higher because detailed failure br...

Hacker News front page

Understanding is the new bottleneck: why you still need to read your agent's code

Geoffrey Litt argues that as agents write more code, human understanding shifts from verification to participation—you need a rich mental model to drive the next creative iteration. He borrows three techniques from education: auto-generated explainer docs that teach background and intuition before code, self-quizzes to check real understanding, and interactive micro-worlds for hands-on exploration. The post doesn't quantify how much these techniques improve outcomes, but frames the cost of skipping them as 'cognitive debt' that compounds over time.

Why it matters: Geoffrey Litt's AI Engineer talk introduces 'cognitive debt' as a framework, which resonates directly with developers using coding agents. It's a sharp concept, not generic commentary. The cap at 78 reflects that this is a personal blog transcript, not a product launch or rese...

AI HOT (Curated Pool)

Google DeepMind launches Gemini 3.7 Flash, a work model built for coding and agents

Gemini 3.7 Flash is a lightweight model from Google DeepMind, positioned as a workhorse for coding and agentic tasks. The official post claims clear gains over its predecessor in code generation, tool use, and long-context work, with better latency and cost. Specific benchmarks and pricing aren't disclosed in the body—only that it will be available via Google AI Studio and Vertex AI. I'd wait for third-party evals before drawing conclusions, but the direction is clear: it's aimed squarely at developer workflows and agent deployment.

Why it matters: Google DeepMind drops a new lightweight model with a clear positioning, but the announcement lacks benchmarks and pricing. Solid product update, but missing key data keeps it from a higher score.

Aug 13Thursday

AI HOT (Curated Pool)

Alibaba open-sources Qwen3.8-2.4T-A95B, SiliconFlow provides Day-0 support

Alibaba released a 2.4T total / 95B active parameter MoE model targeting autonomous coding, deep research, and end-to-end agent execution. SiliconFlow launched API support on day zero: $2.00/M input tokens, $6.00/M output, $0.25/M cached input. The post doesn't disclose benchmarks or architecture details, so I'd wait for third-party evals before getting excited.

Why it matters: Alibaba open-sources Qwen3.8-2.4T-A95B with same-day API availability on SiliconFlow. The 2.4T total / 95B active MoE architecture puts it in DeepSeek V4 Pro territory, and the explicit agentic positioning plus disclosed pricing make this a strong signal. HKR all hit: launch c...

AI HOT (Curated Pool)

Cloud agents start 3x faster with builds

Cursor now pre-builds cloud agent environments in the background every hour—repos cloned, dependencies installed—so agents skip cold setup and respond up to 3x faster. Failed builds are automatically quarantined; agents keep using the last good snapshot. Faire runs 2,000+ automated agent jobs a week on builds, with large repos booting in seconds. Builds become the default for all environments on August 17 at no extra cost.

Why it matters: Cursor cuts cloud agent cold starts from minutes to seconds via background pre-builds and automatic rollback — not just marketing fluff. Faire's 2,000 weekly tasks give the claim a concrete anchor. Not p1 because this is an experience optimization, not a model capability leap,...

AI Chat-Group Daily (群聊日报)

Closed-source reasoning chains extracted at scale; Coze CLI hijacks AI tools

The big one today: researchers extracted hidden reasoning chains from Anthropic, OpenAI, and Google models at scale. The trick is absurdly simple—take Opus 4.8's encrypted CoT and feed it to Haiku 4.5, which decodes it verbatim. All three API families were broken, and decoding 10K trajectories costs about $720. A separate paper shows you can even reverse-engineer reasoning from public outputs alone using a 1.5B-param model. Separately, Coze CLI was caught silently scanning local Codex and Claude Code directories and injecting its own skills into workflows. On the engineering side, the group discussed how prompt debt now rivals traditional code debt—old rules pile up, evals lag behind model iterations, and nobody dares delete anything.

Why it matters: Strong cross-source cluster signal (chat digest + original paper + study notes). First systematic validation that encrypted CoT from three major vendors is cross-model decodable, with concrete $720/10k cost. All three HKR axes hit, but the source is a secondary digest rather t...

Latent Space

xAI drops Grok 4.6 and Grok Bot, a strong new entrant in the AI teammate race

xAI launched Grok 4.6 and the Grok Bot early beta. Grok Bot logs into your tools, operates them like a human, and returns finished work—positioned as an AI teammate. The 1.5T-parameter Grok 4.6 scores near GPT-5.6 Sol Max on the AA-Briefcase knowledge-work benchmark but costs far less: $2/M input tokens, $6/M output. Training reused Grok 4.5 to regenerate SFT traces and added agentic RL across coding, web, CAD, and kernel optimization. Elon says Grok 4.7 is already training. The same day, Qwen3.8-Max dropped as open weights: a 2.4T total / 95B active MoE.

Why it matters: Grok 4.6 matches GPT-5.6 Sol Max on a knowledge-work benchmark at an order-of-magnitude lower price, while the simultaneously launched Grok Bot enters the AI teammate race built by the ex-Cursor team with positive early feedback. Score isn't higher because the Bot is still in ...

Computing Life · Share · Yage

DeepSeek open-sources DSH: agent loop as a hot-swappable plugin, paving the way for self-evolving agents

DeepSeek released its first agent harness, DSH, as open source on August 13. Unlike Codex or Claude Code, DSH treats the agent loop itself as a plugin that can be swapped at runtime. The Cordis runtime handles hot reloads, dependency notifications, and transactional rollbacks. For everyday coding, declarative plugins plus a quick restart are enough—DSH's imperative model adds complexity. But if you want an agent that can generate new tools or replace its own control flow mid-run, DSH is the only option with the infrastructure in place. The post does not disclose performance benchmarks or production-scale data.

Why it matters: DSH makes the agent loop itself a hot-swappable plugin — a real architectural difference, not marketing. But this is a third-party analysis, not an official launch, and DSH has zero production track record yet. Defaulted to the lower band per policy.

AI HOT (Curated Pool)

AutoGPT uses AGENTS.md and skill gating to manage AI-generated pull requests

Over 60% of AutoGPT pull requests now come from AI tools. Maintainer Reinier van der Leer uses an AGENTS.md file to set rules for AI contributors and adds skill gating so only agents that pass linting and unit tests can submit code. Spam PRs dropped sharply, though the post doesn't say how many human contributors were wrongly blocked.

Why it matters: First-hand maintainer account from AutoGPT with hard numbers (60% AI PRs) and two reproducible mechanisms. Downside: the post doesn't disclose how many human contributors got blocked, and it's a single-project case study — generalizability is unproven.

AI HOT (Curated Pool)

Alibaba open-sources Qwen3.8-2.4T-A95B: 2.4T MoE, 95B active, native 256K context

Alibaba's Qwen team open-sourced its first Qwen-Max-level weights. Qwen3.8-2.4T-A95B has 2.4T total parameters with 95B active per token, native 262K context expandable to 1.01M tokens. It uses a 512-expert MoE, routing 10 experts plus one shared expert per token, and includes multi-token prediction training. The model targets coding, office tasks, research, and long-horizon agent workflows. Benchmarks against Opus 4.8, Fable 5, and GPT 5.6 Sol show mixed results, with top scores on PaperBench and IFBench among listed models. Post-training combines combinatorial environment scaling, a unified reward system, and an online data balancer to reduce gradient variance. The post does not disclose the open-source license or inference hardware requirements.

Why it matters: Alibaba's first full open release of a cloud-grade flagship — 2.4T total params, 95B activated, native 256K context — puts it in the top tier. Hits all three HKR axes and triggers the domestic flagship model positive signal. Held back from 90+ because we only have the announce...

TechCrunch · AI

Lovable raises $400M Series C at $13.3B valuation

Lovable confirmed a $400M Series C at a $13.3B valuation, led by Menlo Ventures and Scaleup Europe Fund. It hit $500M ARR in June, hosts 60M projects, and draws 900M monthly visitors. The post doesn't disclose burn rate, but notes Lovable trained its own model and signed a multiyear Google Cloud deal.

Why it matters: Lovable's $400M Series C at $13.3B valuation, with ARR doubling to $500M in six months and a self-trained model, is a solid funding story with real differentiation. Featured tier fits — the numbers are concrete and the self-trained model angle adds substance — but it's not p1 ...

Aug 12Wednesday

Hacker News front page

AI is removing the middle class of software engineering

The author contrasts a 2020 vacation mess with a 2026 Monday morning: 7 PRs, one at 24,506 lines. AI removed the speed limit on bad decisions. Anyone can prompt an agent and ship something that looks functional, but no one knows where the data comes from or why Kafka was added. Reverting one bad call is far harder than generating it, and five more land while you fix it. The bet: AI widens the salary gap—good decision-makers become more valuable, while engineers who only implement become too expensive to hire.

Why it matters: A grounded, first-person engineering observation with concrete scenes and numbers, not generic 'AI will replace devs' fluff. Hits all three HKR axes, but it's a personal blog commentary, not a product launch or research breakthrough, so it lands in the 78-84 band. No cross-sou...

AI HOT (Curated Pool)

Nathan Lambert wrote an AI textbook—models still can't handle long-form nonfiction

Nathan Lambert just finished his post-training textbook *Reinforcement Learning from Human Feedback*. He used LLMs for LaTeX formatting, copyediting, and diagrams, but when he tried to get a model to write a full technical chapter, the output was confusing, poorly organized, and made random conceptual errors. He argues long-form nonfiction writing has stagnated even as models became superhuman at coding and math. The post doesn't cite benchmark scores, but Lambert points to a lack of good training data and notes inference-time scaling hasn't helped writing. His takeaway: if models can't coherently organize established knowledge, autonomous scientific breakthroughs are still far off.

Why it matters: Lambert's first-person experiment delivers concrete failure cases and a data-gap diagnosis — all three HKR axes hit. Deduction: no quantitative benchmark, it's personal experience not systematic research, and the second half drifts into general capability discussion. Sits righ...

AI HOT (Curated Pool)

Meta open-sources Muse Glimmer, a 30B multimodal model for local agents

Meta's Superintelligence Lab released its first open-weight model, Muse Glimmer, now live on OpenRouter. It's a 30B dense text+image model under Apache 2.0, built for reliable local agents. Scores: MCP Atlas 75.5, SWE-Bench Pro 51.2. The post doesn't disclose training data, hardware requirements, or real-world latency—I'd wait before assuming a 30B dense model runs smoothly on consumer hardware.

Why it matters: Meta's first open-weight agent-specific model: 30B dense, Apache 2.0, built for local execution. Scores are cited but SWE-Bench specifics aren't spelled out in the summary, so capped at 78.

TechCrunch · AI

AI code-testing startup Blacksmith's valuation jumps nearly 10x to $550M in under a year

Blacksmith raised a $45M Series B led by Peak XV Partners, with GV and Y Combinator participating. Valuation hit $550M, up from $60M less than a year ago. The startup handles pre-production code testing and validation, growing from 700+ to 5,000+ customers including Mercury, Supabase, Clerk, Ashby, and Expensify. Revenue grew more than tenfold over the past year, per the CEO. The surge reflects a new bottleneck: AI writes code fast, but testing it still needs to catch up.

Why it matters: AI coding has turned testing into the new bottleneck, and Blacksmith's 10x valuation jump nails that trend into a funding headline. Revenue up 10x+, customers from 700 to 5,000+ — the numbers are solid. Docked because the post doesn't disclose actual revenue base, and the topi...

AI HOT (Curated Pool)

xAI releases Grok 4.6, focused on long-running agent capabilities

Grok 4.6 builds on Grok 4.5 with a focus on long-running agents that can research, analyze, code, or turn an idea into a working app across many steps. It matches GPT-5.6 Sol on the AA Intelligence Index at 61, and jumps from 54% to 65.9% on DeepSWE 1.1. xAI reports the model shows more self-testing and verification on longer trajectories. Pricing is $2/M input tokens and $6/M output tokens, with a fast variant at double the price. Available today in Cursor and Grok Build, with 2x included usage for the first week.

Why it matters: xAI releases Grok 4.6 with a focus on long-running agents, matching GPT-5.6 Sol on the AA Intelligence Index and showing a clear jump on DeepSWE. This is a substantive update from a major lab with concrete benchmarks and a direct competitor comparison, earning featured. Not sc...

AI HOT (Curated Pool)

Cursor and SpaceXAI release Grok 4.6, tuned for long-running agents and interactive projects

Grok 4.6 adds a supplemental training run on top of Grok 4.5, using model-generated data to strengthen reasoning and engineering. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The model is better at turning a broad product idea into a working first version and shows more self-verification on long tasks. Pricing starts at $2/M input tokens and $6/M output tokens, with a fast variant at double the price. 2x usage is included in Cursor and Grok Build for the first week.

Why it matters: Matching GPT-5.6 Sol on 9 benchmarks is a hard signal, and the pricing is transparent. But the post only gives a summary — no concrete examples of self-verification or failure modes, so it stays below 85. Cursor's user base and the coding angle make this worth featuring.

Hacker News front page

NVIDIA ships Nemotron 3.5 Lightning and NeMo Switchyard for faster, smarter agent routing

NVIDIA added a 30B-parameter MoE model, Nemotron 3.5 Lightning, to its Nemotron 3 family. It targets specialized tasks inside multi-agent systems, delivering 4x faster output and 30% faster agentic task completion than peers. It runs locally on RTX PCs, DGX workstations, and Jetson. The company also open-sourced NeMo Switchyard, a routing library that directs requests to the best model for each job without app rewrites. CrowdStrike, Harvey, and CodeRabbit are already using customized versions. The post does not disclose pricing or a release timeline.

Why it matters: Nvidia released a 30B MoE model positioned as a specialized worker in multi-agent systems, not a general-purpose model. The 4x output speed and 30% task acceleration claims are useful references, and Switchyard is open-sourced. But this is Nvidia's own blog with no third-party...