Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

341–360 of 1,196

Jul 30Thursday

AI HOT (Curated Pool)

OpenAI cuts GPT-5.6 Luna price by 80%, adds Fast mode for Sol

OpenAI slashed GPT-5.6 Luna's price by 80% and Terra's by 20%. Luna now costs roughly 6% of last year's frontier models per task while running nearly 9× faster. A new Fast mode for GPT-5.6 Sol delivers up to 2.5× speed at 2× price with no intelligence drop. Replit, Notion, Cognition, and others report using Luna for background agent automations, workspace Q&A, and pair programming—citing lower cost, higher speed, and prompt-cache reuse jumping from 24% to 90%.

Why it matters: OpenAI officially announced GPT-5.6 pricing updates: Luna drops 80%, cost falls to 6% of last year's flagship; Sol adds a Fast mode. Concrete numbers, customer quotes (Replit, Notion), substantive product update. Not 85+ because this is pricing/performance optimization of exis...

Hacker News front page

A local merge queue for running parallel Claude Code agents without conflicts

funador open-sourced a local tool that brings GitHub-style merge queuing to Claude Code. You can spin up multiple Claude Code agents in parallel—the tool rebases each agent's changes onto the latest main sequentially, merging them one by one and rolling back on conflicts. The README shows two modes: running directly via the `claude` CLI, or wiring it as an MCP server for Claude Desktop. The post doesn't disclose throughput limits or real team-scale testing, so I'd treat it as a personal experiment for now.

Why it matters: A practical Claude Code utility that ports the CI merge queue pattern to local multi-agent workflows, with clear mechanics and concrete usage. Docked because it's a solo open-source project with no scale validation and a narrow audience (heavy Claude Code users), so it lands r...

Latent Space

AI is eating Finance; AIE NYC now open

OpenAI and Anthropic both held NYC finance AI events, releasing dedicated plugins for equity investing, investment banking, and agent templates for corporate finance workflows. AIE NYC made AI in Finance its mainstage theme, with early bird tickets now open. The post also notes OpenAI's agent security incident expanded beyond Hugging Face to four additional accounts, shifting the discussion toward sandboxing, audit trails, and access controls.

Jul 29Wednesday

AI HOT (Curated Pool)

Why compute might get 10x+ more expensive in coming years

Dwarkesh Patel argues that if a model matches a human software engineer, an H100 should rent for over $250k/year—15x today's spot price. Anthropic may hit $100–150B revenue this year, but training compute only grows 3x annually; sustaining 10x revenue growth would require inference compute to get far more expensive. Google and Anthropic already pay ~2x spot for SpaceX GB200/GB300 clusters, and spot prices are up 40%+ since February. The post doesn't give a timeline, but the logic is clear: smarter models make the same compute more valuable, making it harder for latecomers to compete.

Why it matters: Dwarkesh reverse-engineers compute pricing from engineer salaries, providing a concrete valuation anchor rather than vague trend talk. But it's a personal thought piece, not an industry event, so the score sits at the featured threshold.

Hacker News front page

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

imec's aistack team benchmarked 64 real coding tasks across self-hosted GPUs, rented hardware, and commercial APIs. The newly added Kimi K3, running on an 8×B300 node, costs about 20% more in hardware than the 8×B200 setup used for GLM-5.2, but hits 86.4% task resolution—roughly 24 points above both GLM-5.2 and Claude Opus 4.8 at 62.5%. The trade-off is speed: K3 handles 16 concurrent sessions with a median task time of 38 minutes, about 8× slower than the Claude Code baseline. The post flags that SWEBench Pro tasks may have leaked into K3's training data, so take that resolution number with a grain of salt. The core takeaway: self-hosting doesn't save money—you buy hardware for peak load but pay for it 24/7, and utilization is what makes or breaks the cost case.

Why it matters: imec benchmarked self-hosted, rented, and API setups on 64 real coding tasks. Kimi K3 on 8×B300 hit 86.4% completion—nearly 24pp above GLM-5.2 and Claude Opus 4.8—at a 20% hardware premium. Solid data with named models and numbers; direct signal for teams deciding on self-host...

Hacker News front page

Linux kernel adopts AI for authoring and reviewing patches, Linus insists on technical-only debate

Drew DeVault pushes back on Linus Torvalds' endorsement of AI in kernel development. Over 1,200 commits now carry an 'Assisted-by' tag, mostly from LLM-aided patches. A new tool, Sashiko, uses Google Gemini to auto-generate code reviews, making AI interaction unavoidable even for contributors who opt out. Linus refuses to entertain ethical or political arguments, telling critics to fork the kernel. DeVault calls that disingenuous: Linux is inherently political, the GPL choice was political, and forking is practically impossible. He also flags externalities—rising consumer hardware prices, CO₂ and water costs, and the legitimization of AI firms at the highest political levels.

Why it matters: Linus personally set the tone on AI in kernel development, and Drew DeVault's piece is one of the most substantive counter-voices. The 1,200 commits and Sashiko tool make this more than abstract debate. Downside: it's a single opinion piece, not an official community decision,...

OpenAI News

OpenAI launches ChatGPT for Academic Researchers, giving 100,000 scientists free access to GPT‑5.6

OpenAI is giving 10,000 researchers free access to GPT‑5.6 Sol Pro and Codex this summer, scaling to 100,000 through 2027. Each participant can invite up to four collaborators; data is not used for training by default. The program includes training and hands-on support, and is part of a $250M+ commitment to external research. GPT‑5.6 Sol scores 83% on FrontierMath Tier 4 vs. 72.5% for GPT‑5.5. The post does not spell out eligibility criteria or selection process.

Why it matters: A large-scale free academic rollout with concrete model names and cohort numbers. Capped below 85 because it's a distribution play, not a capability release, and the impact is concentrated in the research community.

AI HOT (Curated Pool)

OpenAI Releases GPT-5.6 Model Family: Sol, Terra, and Luna

OpenAI launched the GPT-5.6 family. Flagship Sol beats Claude Fable 5 on the Artificial Analysis Coding Agent Index at under half the cost. Terra matches GPT-5.5 at half the price, and Luna is 80% cheaper than Sol. Efficiency gains come from inference optimizations and the agentic harness: Sol autonomously rewrote production GPU kernels, cutting end-to-end serving costs by 20%. The post doesn't name the benchmarks for Terra and Luna, nor does it give absolute pricing for Sol.

Why it matters: OpenAI launches GPT-5.6 family: flagship Sol beats Claude Fable 5 on coding agent benchmarks at less than half the cost, with Terra and Luna targeting price-performance tiers. This is a top-tier model refresh with concrete comparisons and disclosed efficiency mechanisms — a sa...

Jul 28Tuesday

Latent Space

OpenAI's Codex and ChatGPT Work hit 10M users, with non-developers making up 20%

OpenAI product engineering lead Akshay Nathan walked through the origin of ChatGPT Work on the Latent Space podcast. Codex started as a coding tool but took off internally among non-engineers; knowledge workers now account for roughly 20% of its user base and are growing over 3x faster than developers. The team extracted Codex's agent harness to build ChatGPT Work for documents, spreadsheets, and slides, launching July 9 and reaching 10M combined users within two weeks. Akshay detailed the shared harness, differing UX and sandboxing defaults, and how Sites, OpenClaw, memory, and sub-agents let non-coders delegate work to AI. He also noted that when anyone can build, ideas and taste become the bottleneck—and LLMs still struggle to generate genuinely grounded new ideas.

Why it matters: OpenAI's product engineering lead gives a first detailed breakdown of ChatGPT Work's origin, with a concrete stat that knowledge worker growth is 3x that of developers. Directly useful for anyone building agent products. Score isn't higher because this is a podcast interview, ...

Ben's Bites

Claude Opus 5 ships at half the price of Fable 5, but early users say it argues and stops early

Anthropic released Claude Opus 5 at half the cost of Fable 5, claiming near-parity. Every's review found it argues, stops early, and fights old prompting habits. Anthropic cut over 80% of Claude Code's system prompt for Opus 5 and Fable 5 with no measurable coding-eval loss. Theo spent hours rewriting CLAUDE.md and skills files and called it worth it. ChatGPT Voice now controls the desktop app inside Work and Codex, spawning new sessions for tasks and reporting back—like a voice-driven OpenClaw. Keshav found it weaker for serious work than manually using 5.6 Sol in Codex, but decent for email, dashboards, and charts. Claude's voice mode quietly added Sonnet and Opus support plus mid-conversation tool calls to Gmail, Calendar, and Slack. Kimi K3 weights and tech report are public, with a 50% discount on Droid until Aug 10. Jensen Huang posted on X for the first time amid rumors of a US ban on Chinese open-weight models.

Why it matters: Anthropic model launch with halved pricing is a substantive update. Every and Theo's hands-on tests provide concrete signal: strong capability but awkward behavior requiring prompt rewrites. Cross-source discussion is forming, but the body is summary-only—missing full review d...

Hacker News front page

Kimi Linear: A Hybrid Linear Attention That Beats Full Attention

Moonshot AI's Kimi team released a tech report on Kimi Linear, a hybrid linear attention architecture. Its core, Kimi Delta Attention (KDA), extends Gated DeltaNet with finer-grained gating to use limited RNN memory more effectively. They trained a 3B-active, 48B-total MoE model mixing KDA and MLA layers. Under the same recipe, it outperforms pure MLA across all benchmarks, cuts KV cache by up to 75%, and boosts 1M-context decoding throughput 6x. The team open-sourced the KDA kernel, vLLM integration, and model checkpoints.

Why it matters: Moonshot AI drops an architecture-level tech report with a concrete hybrid linear attention mechanism and a 48B MoE model. Not scoring higher because it's an arxiv preprint with no product timeline — real-world impact depends on community reproduction and third-party benchmarks.

AI Chat-Group Daily (群聊日报)

Chat Digest: Gowers Says Math Is Dying, Opus 5 Stumbles on Day 3

Fields medalist Gowers refused to sign the Leiden Declaration and wrote a long post arguing math won't die from AI's inability but from an evidence glut—like lake eutrophication, where literature booms but human experts vanish. He's twice seen GPT 5.6 Pro one-shot problems he'd thought hard about. Meanwhile, Anthropic's Claude Opus 5 entered day three of real-world testing: it stalls on execution after one step, and its safeguards falsely flag a dev board query, triggering a double downgrade. Sentiment turned negative.

Why it matters: Fields Medalist Gowers refused to sign the Leiden Declaration and published a long essay arguing AI won't kill math through incompetence but through evidence surplus, backed by two personal encounters with GPT 5.6 Pro. The source is a chat-group digest rather than original rep...

TechCrunch · AI

Cursor launches India-specific $7/month plan ahead of SpaceX acquisition

Cursor launched Cursor Start, a ₹649/month (~$7) India-only plan, well below its standard $20 Pro tier. It's the startup's first country-specific pricing. India is now Cursor's third-largest market globally; the company plans to expand local hiring and enterprise sales. The move comes weeks before SpaceX's expected acquisition closes — the post doesn't say whether India strategy changes post-deal.

Why it matters: Cursor's first country-specific pricing, dropping India to $7/month right before the SpaceX acquisition closes, with concrete price anchors and market data. Score isn't higher because this is a market expansion move, not a product capability update, and the post doesn't clarif...

TechCrunch · AI

Microsoft launches its first cybersecurity model MAI-Cyber-1-Flash and agentic platform Perception

Microsoft unveiled two security products at a small San Francisco event. MAI-Cyber-1-Flash is its first cybersecurity-focused model, built to find hard-to-spot vulnerabilities in complex codebases and power the MDASH vulnerability harness. Perception is a new platform that deploys agent teams to automate security workflows like bug discovery and remediation. The post doesn't disclose model parameters, benchmarks, pricing, or which tools Perception integrates with.

Why it matters: Microsoft's first dedicated cybersecurity model and agentic platform bring real mechanism novelty, but the post omits param count, benchmarks, and pricing — thinning the knowledge signal. H and K hit, R is weak, landing right at the featured threshold.

Jul 27Monday

Import AI (Jack Clark)

AI completes week-long coding tasks and robot chores in 9 minutes

Epoch and METR's MirrorCode benchmark shows Claude Opus 4.7 reimplemented a 2–17 week human coding task in 14 hours for $251, though it still struggles with projects like ruff. Anthropic had Opus 4.7 autonomously finish robot fetch tasks in 9 minutes 35 seconds, 20x faster than last year's human-assisted record. Robot startup Sunday confirmed the same pattern: scale pretraining, then fine-tune on small high-quality data, hitting 99.1% on laundry folding.

Why it matters: MirrorCode is a long-horizon programming benchmark from Epoch and METR, with Claude Opus 4.7 reimplementing a 2-17 week human project in 14 hours — concrete numbers and failure cases included. HKR all hit, but this is a newsletter summary, not the original paper, and complex t...

Hacker News front page

Bun's Rust rewrite: six weeks after merge, still no release tag

Bun announced a Zig-to-Rust rewrite using Anthropic's Claude on July 8, claiming 11 days and $165K in API costs. Tom Lockwood dug into the repo and found no release tag six weeks after the merge—the last release was May 12. Open PRs from robobun (Claude Code) grew from 1,277 to 2,475; merging them all at current CI speed would take 86 continuous days. Anthropic employees are directly contributing PRs, and the pace is accelerating. Lockwood estimates real spending may be approaching $800K and argues the rewrite is far from 'done.' The post does not disclose feature-completeness or test-pass rates.

Why it matters: An independent repo audit with receipts, directly answering Bun's splashy '11-day AI rewrite complete' claim. All three HKR axes hit: suspenseful headline, concrete numbers and release gaps, and the topic sits right on the fault line of AI-replacing-OSS-maintainers. Not scored...

OpenAI News

OpenAI study: 43.5% of occupation-specific ChatGPT use crosses job boundaries

OpenAI Economic Research analyzed 800,000+ ChatGPT messages from US users. 16.8% of work messages and 43.5% of occupation-specific messages involve tasks from another occupation—a pattern they call 'task crossover.' Customer experience (77%), design (75%), and HR (69%) workers borrow the most. Marketing and engineering tasks travel farthest across fields. Crossover is more common in small businesses. The report also notes AI is creating new tasks like prompt engineering and output review that don't fit standard job classifications. This is the first paper in the 'Work at the Frontier' series; the full PDF is available.

Why it matters: OpenAI's own research with 800k conversations as the dataset—credible scale. The 43.5% crossover rate is a fresh signal, far more concrete than generic 'AI changes work' narratives. Not an 85 because it's a report, not a product launch or model release—impact is more diffuse.

Computing Life · Share · Yage

Four AI coding harnesses all claim multi-agent, but their architectures diverge radically

This piece dissects the multi-agent architectures of Claude Code, OpenAI Codex, Cursor, and Antigravity. Claude Code explores tree-based spawning and peer-to-peer Agent Teams with a shared tasks.md ledger. Codex assigns different models and reasoning effort (low/medium/high) per sub-agent to optimize cost and throughput. Cursor binds agent loops directly to IDE state, using Merkle Tree indexing and SQLite for non-blocking background edits. Antigravity enforces explicit planning with a Proceed Gate and isolates sub-agents via Git Worktree. The choice depends on whether you prioritize communication topology, compute efficiency, editing UX, or audit-grade governance.

Why it matters: A cross-sectional deep dive into four major AI coding tools' multi-agent architectures, with source-level details like shared ledgers and reasoning-affinity matching. The density is well above typical reviews. The slight discount is because it's an independent blog rather than...

Computing Life · Share · Yage

Why high SWE-bench scores don't translate to real-world Kotlin projects

JetBrains released the Kotlin Benchmark on July 24, 2026, testing AI coding agents on 105 real-world tasks across 8 open-source projects. With the same Opus 4.7 model, Claude Code hit 85.71% and Junie 81.9%—a nearly 4-point gap driven by how each agent harness handles Gradle build logs. A good harness uses Language Server diagnostics to catch static errors locally, then runs full builds only at key checkpoints and extracts just the blocking lines from noisy output. The article argues that Python-based benchmarks like SWE-bench reward trial-and-error strategies that collapse under Kotlin's heavy build overhead. It recommends teams stop buying off public leaderboards and instead use the Harbor container spec to build private micro-eval matrices from their own historical PRs and issues, measuring Pass@k stability, token cost, and whether patches respect internal architectural constraints.

Why it matters: JetBrains official benchmark with concrete numbers plus an engineering-level breakdown of SWE-bench's limitations. Not just complaining about leaderboard distortion—it traces the root cause to static compilation overhead vs Python's dynamic runtime feedback. Slight ding becaus...

Hacker News front page

AST-grep Rewrote Tree-sitter's Core in Rust, Parsing Is 30% Faster

AST-grep rewrote Tree-sitter's C parsing core in Rust, with AI generating most of the code. Raw parsing throughput is up ~30%, and end-to-end ast-grep outline runs are 22% faster, at the cost of ~8 MiB more memory. The rewrite drops incremental parsing and WASM loading, targeting AI coding agents that analyze full file snapshots. The post notes an early version peaked above 1 GiB on a TypeScript stress corpus; the final build peaks at 91.2 MiB.

Why it matters: Concrete perf numbers and clear engineering tradeoffs hit H and K. But the audience is narrow, R is absent, so it lands at the featured threshold of 72. If later posts in the series deliver quality data on the AI-generated code, the score could go higher.