Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

461–480 of 1,196

Jul 9Thursday

Hacker News front page

Anthropic's Fable is not a useful model for CS research tasks

Rob Patro from COMBINE-lab shares two first-hand failures that make Fable useless for his CS research. First, Fable's safety classifier rejected a prompt to help port the C++ tool salmon to Rust, flagging RNA-seq biological terms. After 15–30 minutes of rephrasing, he gave up and used Opus 4.8 successfully. Second, he asked Fable to tackle a network evolution reconstruction algorithm; the post doesn't disclose the outcome but calls it an 'unforgivable' flop. Patro argues Fable's classifier behaves more like a crude blocklist of terms and users, refusing even 'what is a mitochondrion?'.

Why it matters: Rob Patro tested Fable on two real coding tasks, both killed by safety filters; Opus 4.8 handled them fine. First-hand record of Anthropic's safety model failing in professional use, with concrete comparisons and time costs. Score capped because it's a single blog post, not a ...

TechCrunch · AI

SpaceXAI releases Grok 4.5, which Elon describes as an ‘Opus-class model’

SpaceXAI dropped Grok 4.5 weeks after going public, pitching it for coding, office work, research, and writing. The company claims twice the token efficiency of other leading models, which would cut usage costs if it holds up. Elon Musk calls it an ‘Opus-class model,’ signaling it aims at Anthropic’s top tier. The post doesn’t disclose pricing, parameter count, or third-party benchmarks, so I’d wait for independent evals before buying the efficiency claim.

Why it matters: First model post-IPO with Musk directly calling it 'Opus-class' — strong H and R, but the post lacks params, pricing, and benchmarks, so the 2x efficiency claim is unverified. Hits featured threshold on narrative weight alone.

Hacker News front page

SpaceXAI launches Grok 4.5, built for coding and agentic tasks, co-trained with Cursor

Grok 4.5 is SpaceXAI's strongest model, tuned for coding, agentic tasks, and knowledge work. It scores 62% on DeepSWE 1.0 and 64.7% resolve rate on SWE Bench Pro, though it trails Fable and GPT 5.5 on most listed benchmarks. The standout number is token efficiency: 15,954 output tokens on average per SWE Bench Pro task, 4.2× fewer than Opus 4.8. Inference speed is 80 TPS, priced at $2/$6 per million input/output tokens. The model was trained across tens of thousands of GB300 GPUs, with RL focused on multi-step software engineering. The post doesn't disclose parameter count, context window, or a precise EU launch date beyond mid-July. Available now in Grok Build, Cursor, and via API.

Why it matters: SpaceXAI launches Grok 4.5 targeting coding and agents, co-trained with Cursor — a real differentiator. 64.7% on SWE Bench Pro isn't top, but 16K avg output tokens (4.2x less than Opus 4.8) is a concrete cost edge. Pricing and latency not disclosed — those decide whether this ...

Hacker News front page

Cognition launches SWE-1.7: near GPT-5.5 coding intelligence trained from Kimi K2.7 at a fraction of the cost

Cognition released SWE-1.7, a coding model trained via RL post-training on a Kimi K2.7 base. It scores 42.3% on FrontierCode 1.1, close to GPT-5.5’s 43.0% and a huge jump from SWE-1.6’s 9.4%. It also hits 81.5% on Terminal-Bench 2.1 and 77.8% on SWE-Bench Multilingual, both competitive with GPT-5.5. The gains come from four RL pipeline upgrades: top-p sampling with distribution replay to prevent entropy collapse, multi-continent multi-cluster training with fault tolerance, automated execution-based data filtering, and self-compaction that lets the model summarize long-horizon task state to exceed the context window. SWE-1.7 is live in Devin via Cerebras at 1000 TPS. The post does not disclose specific pricing, only that it advances the cost-performance curve.

Why it matters: Cognition drops SWE-1.7: RL post-training on a Kimi K2.7 base lifts FrontierCode 1.1 pass rate from 9.4% to 42.3%, within a point of GPT-5.5. The numbers are solid and the narrative is sharp, but the post only gives a summary—training details and cost comparisons aren't spelle...

Jul 8Wednesday

OpenAI News

OpenAI audits SWE-Bench Pro, finds ~30% of tasks are broken

OpenAI audited SWE-Bench Pro and estimates ~30% of its tasks are broken. An automated pipeline flagged 286 suspicious tasks; Codex-based investigator agents and five experienced engineers then reviewed them. Engineers identified 249 (34.1%) flawed tasks, mostly due to overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI advises model developers to scrutinize results rather than trust leaderboard scores. The post does not disclose a fix timeline or a revised dataset release.

Why it matters: OpenAI audited SWE-Bench Pro and found ~34% of tasks defective — a ratio that forces the industry to re-examine coding benchmark reliability. The post provides concrete defect categories and a human review pipeline. Not scored higher because this is a benchmark quality report,...

AI HOT (Curated Pool)

Claude team shares two multi-agent patterns: Advisor and Orchestrator

Claude developers shared two multi-agent patterns their team uses heavily. In Advisor mode, Sonnet 5 executes while calling Fable 5 for guidance via tool calls; on SWE-bench Pro the combo hits 84% at $1.40, saving 37% cost vs pure Fable 5 with only an 8-point accuracy drop. In Orchestrator mode, Fable 5 plans and fans out tasks to multiple Sonnet 5 workers; on BrowseComp it reaches 86.8% at $18.53, less than half the cost of all-Fable 5. Both patterns route heavy lifting to cheaper models and reserve expensive ones for key decisions.

Why it matters: Anthropic dev shares two multi-agent patterns with concrete SWE-bench scores and cost breakdowns — directly useful for teams building agents. Score held back because it's an individual share, not an official release, and the Orchestrator mode lacks benchmark numbers.

Hacker News front page

Three senior engineers charge $10k/week to delete AI-generated code

Three Polish engineers launched Slopfix on odra.dev: one week, $10,000, to refactor vibecoded codebases back to maintainability. They offer a free analysis with a committed reduction target—e.g., 100k lines down to 35k, same functionality. Payment is proportional to the target hit; if they promise 50% and deliver 20%, you pay $4,000. Deliverables include the smaller codebase, a QA checklist, guardrails (CLAUDE.md, lint rules, CI checks), and a two-week warranty. They use Claude Code but say 'the agent doesn't get a vote'—the differentiator is 30 years of combined engineering experience. The post does not disclose any specific client cases or number of projects delivered.

Why it matters: Three Polish engineers turned AI code refactoring into a fixed-price service with pay-for-results. Hits all three HKR axes, but as a small-team service page rather than an industry event, importance caps at 78.

AI HOT (Curated Pool)

Liquid AI open-sources Antidoom, a final-token preference optimization method that fixes reasoning model doom loops

Reasoning models can get stuck in doom loops, repeating useless tokens until the context window fills up. Liquid AI open-sourced Antidoom, which uses Final Token Preference Optimization (FTPO) to fix this. The method trains the model on 1,040 preference pairs to learn when to stop at the end of reasoning. On DeepSeek V4 Pro, the doom-loop rate dropped from 3.2% to 0.3% without hurting math or coding scores. The post doesn't disclose training cost or how well it transfers to non-DeepSeek models.

Why it matters: Liquid AI open-sourced a practical fix for reasoning model doom loops, dropping the rate from 3.2% to 0.3% on DeepSeek V4 Pro — solid numbers. Not scoring higher because it's a single blog post with no paper or third-party validation yet; 78 for a strong single-source piece.

Hacker News front page

Liquid AI cuts reasoning-model doom loops from 10.2% to 1.4% with Final Token Preference Optimization

Liquid AI introduces Antidoom, a method that targets the exact first token of a repetitive loop in small reasoning models. Using Final Token Preference Optimization (FTPO), it trains the model to prefer coherent alternatives at that single position while leaving the rest of the distribution mostly untouched. On an early LFM2.5-2.6B checkpoint, the loop rate on hard math and coding prompts dropped from 10.2% to 1.4%, and eval scores improved as a result. The approach adapts Antislop and uses chosen/rejected single-token pairs, making it cheaper than RL. The post does not disclose training compute cost or latency impact.

Why it matters: Liquid AI proposes a lightweight fix for doom loops in small reasoning models: identify the first token of the loop and use preference optimization to swap it. The idea is clever, but it's only validated on an early 2.6B checkpoint—no cross-model or larger-scale comparisons ar...

Jul 7Tuesday

Hacker News front page

Craig Mod built his own accounting software TaxBot2000 in five days with Claude Code

Writer Craig Mod describes a year of obsessive building with Claude Code. He rebuilt a Twitter-like community space with ephemeral posts, then made video search tools and small utilities. Last week he spent five days building TaxBot2000—a local, subscription-free accounting app in Python, Flask, and SQLite. It handles multi-currency, multi-country accounts, pulls daily FX rates, learns categorization habits, and lets him talk to Claude to fix anomalies. He calls it the best accounting software he's ever used, replacing a decade of Quicken and Google Sheets hacks. The post doesn't disclose exact build costs, only that occasional fixes cost a few dollars.

Why it matters: Craig Mod's five-day TaxBot2000 build with Claude Code is a concrete first-person experiment that hits all three HKR axes. Not scored higher because it's a personal productivity tool share, not an industry-level product update or research breakthrough — sits right at the featu...

Hacker News front page

Automating away LLM clumsiness with deterministic tools

The author finds that even brilliant LLMs like Claude remain imprecise and non-deterministic—committing the build/ dir twice, for example. The fix is sandwiching the LLM between fast, deterministic tools and formal workflows: automate repeated actions into scripts, automate verification for recurring failures. Beagle SCM lets LLMs script their own routines in JavaScript, with heavy lifting in C and a malleable JS tooling layer, so the model essentially automates itself away.

Why it matters: A hands-on reflection from a developer building with Claude. Uses a concrete failure (committing build/ twice) to argue for sandwiching LLMs between deterministic tools and workflows. Not scored higher because it's a sharp engineering essay, not a product launch or research re...

AI HOT (Curated Pool)

Claude Code now lets you pick a Claude model and effort level for each task

Anthropic added two controls to Claude Code: model selection and effort level. You can assign Opus to cross-file refactors, Sonnet to routine edits, and Haiku to quick fixes. Effort levels—low, medium, high—adjust how deeply the model thinks and how many tool calls it makes. High effort with Opus triggers multi-step codebase searches and test runs, but burns more tokens. The post doesn't disclose exact pricing deltas, only that high effort plus Opus is the most expensive combo. The update lets developers dial compute up or down per task instead of using one model for everything.

Why it matters: Official Anthropic product guide, not fluff. Effort-level behaviors are concrete (multi-step search, auto test runs), directly useful for daily users. Points off for no pricing comparison—only says high effort burns 'the most' tokens without numbers. Lands at the featured thre...

AI HOT (Curated Pool)

Intelligence is Free, Now What? Data Systems for, of, and by Agents

UC Berkeley's BAIR Lab argues that as inference costs approach zero, data systems face three shifts. First, systems for agents: a single user request can spawn thousands of SQL queries, but 80–90% of sub-queries are duplicates, so reusing results or returning approximate answers can speed things up. Second, systems of agents: thousands of agents need a new substrate to manage state, coordinate, and handle failures. Third, systems by agents: agents can now synthesize entire data systems, but verifying correctness remains an open problem. The post is a research roadmap and does not provide a deployment timeline.

Why it matters: Berkeley BAIR dropped a roadmap with a real thesis and hard numbers, not a vague trend piece. The core insight — when inference is nearly free, database systems get rebuilt for, of, and by agents — is sharp, and the 80-90% duplicate subquery stat gives engineers a concrete tar...

Hacker News front page

YC CEO claims 37K AI LoC/day; a dev finds bloat, test files, and rookie mistakes in production

Garry Tan posted that his AI coding agents ship 37K lines/day across 5 projects, with a 72-day streak. Polish dev Gregorein inspected the front end of Tan's AI blog and found 169 requests totaling 6.42 MB—versus Hacker News's 7 requests and 12 KB. The site ships 28 test files to every visitor, loads 78 JS controllers regardless of use, and serves the logo in 8 formats including a 0-byte file. Gregorein notes this is front-end only; the back end wasn't touched. His take: AI generates code faster than anyone can review, and the response from people like Tan seems to be 'so stop reviewing'—echoing Facebook's 'move fast and break things.'

Why it matters: Garry Tan claims 37K LoC/day via AI coding; a third-party dev finds 169 requests and 6.42 MB for a simple blog frontend, vs. HN's 1 request and 0.02 MB. Numbers, contrast, and identity reversal hit all three HKR axes. Capped below 85 because it's a single report, not a product...

Product Hunt · AI

Meituan releases LongCat-2.0: a 1.6T MoE model, MIT-licensed, trained on custom AI ASICs

Meituan launched LongCat-2.0 on Product Hunt: an MIT-licensed 1.6T-parameter MoE model with ~48B active parameters and 1M context window. It uses LongCat Sparse Attention and is post-trained for coding and agentic workflows. The model was trained entirely on Meituan's own AI ASIC superpods, not NVIDIA GPUs. It integrates with Claude Code, OpenClaw, and Hermes. The post doesn't disclose benchmark scores, API pricing, or throughput — I'd hold off on performance claims until numbers land.

Why it matters: Meituan LongCat-2.0 is a 1.6T-param MoE model, MIT-licensed, trained entirely on in-house AI chips with a 1M-token context window and post-training focused on code and production deployment. Flagship domestic model release with a non-NVIDIA training story — HKR all hit. No ben...

Hacker News front page

GLM 5.2 hands-on: the first open-weights model that feels like Opus and GPT, and why inference margins are next to collapse

The author used GLM 5.2 as a daily driver for two weeks and found it nearly indistinguishable from Claude Opus for most tasks. Switching is trivial—just point the API base URL to a compatible endpoint and it runs inside Claude Code. Two real gaps: no vision support, and the built-in web search is slow and poor, which hurts agentic workflows that rely on images or live lookups. Inference pricing sits around $4.40/MTok, under 20% of Opus’s retail rate; even with heavier token usage, costs drop by more than half. The post argues that frontier labs’ ~90% inference gross margin is unsustainable once open-weights models hit this quality bar.

Why it matters: The author ran GLM 5.2 as a daily driver for two weeks and provides a reproducible swap path plus pricing—this isn't a press release. Two limits keep it at 78: the test covers only coding workflows, and the vision/search gaps narrow the claim's reach. It's a single-blog experi...

Jul 6Monday

Import AI (Jack Clark)

Fable writes first GPU megakernel; AI online work automation quadruples in 8 months

Fable submitted the first genuine GPU megakernel on KernelBench-Mega, achieving an 18.71x speedup over an optimized PyTorch baseline with a single cooperative kernel launch per decoded token. Claude Opus 4.8 reached 14.4x and GPT-5.5 only 4.34x. This benchmark measures AI systems writing their own low-level kernels, a signal for recursive self-improvement. Separately, the Remote Labor Index shows AI end-to-end success on online freelance projects rose from 2.5% in October 2025 to 16.1% in July 2026, with Fable 5 hitting 16.1%. Tasks span 3D modeling, animated ads, and architectural renders, with a median human completion time of ~1.6 hours. The post does not disclose specific model scores on OSWORLD 2.0, only noting poor performance so far.

Why it matters: Fable submitted the first genuine megakernel to KernelBench-Mega, hitting 18.71x speedup with a single cooperative kernel launch — cleaner than Claude Opus 4.8 and GPT-5.5 entries. It's an early signal of AI improving its own low-level kernels, directly relevant to people doin...

Hacker News front page

Does Code Cleanliness Affect Coding Agents?

SonarSource researchers ran 660 trials with Claude Code across 33 tasks in minimal-pair repos. Code cleanliness didn't change pass rates, but cleaner code cut token usage by 7–8% and file revisits by 34%. Clean code doesn't decide success, but it materially lowers compute cost and navigation overhead for coding agents.

Why it matters: Solid experimental design with minimal pairs to isolate the variable. The finding is counterintuitive: cleanliness doesn't affect pass rate but saves tokens and reduces redundant operations. Directly useful for engineers using coding agents daily. Points off for small sample (...

Jul 5Sunday

AI HOT (Curated Pool)

Meituan LongCat-2.0 fully open-sourced under MIT license, releasing 1.6T MoE weights and inference code

Meituan fully open-sourced LongCat-2.0 under MIT license, releasing both weights and inference code. It's a 1.6T-parameter MoE model activating ~48B per token, with 1M-token context. LongCat Sparse Attention handles long sequences, Zero-Compute Experts dynamically activate 33B–56B to avoid wasted compute, and MOPD routes tasks across Agent, Reasoning, and Interaction expert groups. On benchmarks: SWE-bench Pro hits 59.5, edging out GPT-5.5's 58.6; Terminal-Bench 2.1 scores 70.8; multilingual SWE-bench reaches 77.3. It natively integrates with Claude Code, OpenClaw, and Hermes Agent, supports GPU and NPU deployment, and has been validated on large-scale domestic clusters.

Why it matters: Meituan fully open-sources LongCat-2.0, a 1.6T MoE model, under MIT license with weights and inference code — a rare move from a major Chinese tech company. The 1M-token context window and sparse attention design are concrete technical hooks, not just marketing. Score held at ...

Hacker News front page

Simon Willison used Claude Fable to fix critical bugs in sqlite-utils 4.0 for about $149.25

Simon Willison had Claude Fable do a final review before shipping sqlite-utils 4.0 stable. The model found a data-loss bug where delete_where() never commits, silently rolling back all subsequent writes. The fix took 37 prompts, 34 commits across 30 files, costing $149.25. He then had GPT-5.5 cross-review the changes and found two more issues. The new release rewrites transaction docs: all writes auto-commit by default, and you only need to think about transactions when using db.atomic() or manual begin().

Why it matters: Simon Willison used Claude Fable for a pre-release review of sqlite-utils 4.0 and the model caught a silent data-loss bug. Full prompt count, commit count, and cost are disclosed — this is a first-person experiment, not a vendor case study. Not scored higher because it's a sin...