Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

381–400 of 1,196

Jul 22Wednesday

TechCrunch · AI

Menlo Ventures' Matt Murphy: The model was never the moat—platforms win

Menlo Ventures partner Matt Murphy told Equity podcast that Anthropic hit a $47B revenue run rate by May 2026, up from $9B in 2025—growth he hasn't seen in 25 years across internet, mobile, or cloud waves. Menlo led Anthropic's $500M Series D at a $4B pre-revenue valuation. Murphy argues the model was never the real moat; Claude Code, MCP, and Claude Skills turned Anthropic into a platform. He also flagged Lovable and Legora as growing even faster, and pushed back on criticism that Anthropic's Mythos launch was more marketing than safety—though the post doesn't detail his counterarguments.

Why it matters: Anthropic revenue figures are newsworthy, and the investor's cross-cycle perspective has real judgment. HKR all hit. Capped below 85 because this is a podcast recap, not hard news, and TechCrunch's Equity is a regular column.

Hacker News front page

Codeberg bans vibe coded projects via ToU amendment

Codeberg members voted to amend the Terms of Use, banning projects that mostly consist of LLM-generated code without human review. The proposal argues such projects have unclear copyright and lack safeguards. The post doesn't define 'mostly' or specify an enforcement timeline.

Why it matters: Codeberg membership voted to ban unreviewed AI-generated code via ToU amendment — a substantive governance move with conflict, new information, and emotional resonance. Score held back by vague enforcement details and scope limited to Codeberg ecosystem, not industry-wide.

Hacker News front page

Fireworks benchmarks Kimi K3 vs Fable 5: routing the two hits 93% accuracy at far lower cost

Fireworks ran ~1,030 real agent tasks comparing Kimi K3 and Fable 5. Head-to-head on SWE: K3 92.4%, Fable 92.6% — a near tie. Under the hood, K3 excels at symbolic math and dev tooling; Fable wins on web and data viz. Routing between the two lifts accuracy to 93% and cuts cost up to 50x on long agent loops. Oracle routing shows K3 can handle 72–96% of tasks, but a practical router still needs an order of magnitude more routing data.

Why it matters: Fireworks ran ~1,030 real agent tasks comparing K3 and Fable, with concrete numbers and a task-type breakdown. The finding that routing between the two beats either alone is useful and testable. Score stays at 78 rather than higher because this is a single-source vendor post—F...

Hacker News front page

Poolside launches Laguna S 2.1, a 118B MoE coding model that leads its weight class on long-horizon benchmarks

Poolside released Laguna S 2.1 today, a 118B MoE model with 8B active parameters per token and a 1M-token context window. It scores 70.2% on Terminal-Bench 2.1, beating DeepSeek-V4-Pro Max (64.0%) and Inkling (63.8%), and trailing Tencent Hy3 (295B) by only 1.5 points. On DeepSWE long-horizon tasks it hits 40.4% vs DeepSeek-V4-Pro Max's 9.0%. Poolside says training to launch took under nine weeks and published full eval trajectories. The post doesn't disclose training data cutoff or non-coding performance.

Why it matters: Poolside ships a small-activation MoE coding model that beats DeepSeek-V4-Pro Max on Terminal-Bench 2.1 (70.2% vs 64.0%) with a 1M context window. Capped below 85 because Poolside lacks tier-1 market presence and third-party repro — treat as a strong product update with number...

AI HOT (Curated Pool)

Google ships three new Gemini models, but the flagship 3.5 Pro is still missing

Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. 3.6 Flash is the new workhorse—better at coding and multimodal tasks, with up to 17% lower token usage and a cheaper price than 3.5 Flash. 3.5 Flash-Lite targets extreme cost efficiency, and 3.5 Flash Cyber is fine-tuned for finding and fixing security vulnerabilities, available only to governments and trusted partners in a limited pilot. The whole drop is about efficiency, latency, and reliability for customers building AI agents at scale. The real story is what’s absent: the flagship Gemini 3.5 Pro hasn’t been updated since February, while OpenAI shipped GPT-5.5 and started rolling out GPT-5.6, and Anthropic launched Claude Opus 4.8. The post doesn’t explain what’s holding up the Pro line.

Why it matters: Google dropped three Gemini models at once, with 3.6 Flash as the new workhorse showing clear gains in code and multimodal tasks plus a 17% token reduction—a real cost signal for developers. The absence of 3.5 Pro adds discussion value. Score capped slightly because the TechCr...

Jul 21Tuesday

Hacker News front page

Claude Is Not a Compiler — It's Better

Josh Bleecher Snyder now says calling Claude a compiler is a category error — it's better. A compiler only handles source-to-binary decisions, but Claude works vertically across strategy, product, architecture, and machine code. He walks through building exe.dev's distributed DNS server: multiple concurrent agent loops implemented the same design yet made wildly different choices on decisions they never asked about. By answering their questions and reverting bad calls, he slowly converted hard-won knowledge into terse written guidance.

Why it matters: The author uses a real product case to argue Claude is a cross-layer decision-maker, not a compiler, with concrete code behavior comparisons. But it's a personal blog opinion without peer validation or benchmarks, so it lands at the 72 featured threshold.

Ben's Bites

Kimi K3 tops Fable on frontend coding leaderboard, but token inefficiency cancels cost edge

Moonshot AI's Kimi K3 beat Fable and GPT-5.6-Sol on Arena's frontend coding leaderboard and came close on other benchmarks. It's a 2.8T-parameter model with a 1M-token context window and image support; weights will be open-sourced by July 27. Token inefficiency cancels its per-token price advantage: half the cost per token but twice the tokens used. New subscriptions are paused due to GPU shortages. Fable 5 is now a permanent part of Claude Max/Team plans, with Pro users getting a one-time $100 credit. Fable also found a counterexample disproving the 87-year-old Jacobian conjecture. Sierra launched Horizon, outcome-priced long-running agents. NotebookLM rebranded to Gemini Notebook and added Collections.

Why it matters: Moonshot drops Kimi K3, topping Fable and GPT-5.6-Sol on Arena's frontend coding board. 2.8T params, 1M context, open-source on July 27 — all hard signals. The token-efficiency gap is a real weakness but makes the story more substantive. Held at 82 rather than 85+ because only...

AI HOT (Curated Pool)

Anthropic team says Claude Tag lands 65% of product PRs, system prompt cut by 80%

Anthropic's Cat Wu and Thariq Shihipar shared Claude Code team practices at AI Engineer World's Fair with Simon Willison. Claude Tag, their Slack collaboration tool, now lands 65% of the team's product engineering PRs. They found that stuffing system prompts with examples or 'don't do X' rules hurts output quality on Fable 5 and Opus 4.8, so the Claude Code system prompt shrank by 80%. Thariq noted Fable can edit video to meet their brand team's bar, and the team now treats rewrites as a valid move—the Bun-in-Rust Claude Code already shipped to everyone. The transcript cuts off before detailing what non-engineers do with Claude Tag.

Why it matters: First public disclosure of Claude Tag's internal adoption (65% of product PRs) and a concrete prompt-engineering shift (80% reduction) from the Claude Code team. Held at 82 because it's a fireside chat without reproducible evals or scripts — strong signal, not yet a paper.

Jul 20Monday

Hacker News front page

Kimi K3 and Qwen 3.8 go open, squeezing Anthropic from both sides

Moonshot's Kimi K3 and Alibaba's Qwen 3.8 launched this week, both near Anthropic Fable 5 in performance and set to release weights publicly. The piece runs the numbers: Anthropic leases data centers and buys electricity, so inference costs scale with usage. Fable 5 costs nearly 3× per completed task vs. competitors. Open models catching up makes a premium-pricing strategy fragile. Anthropic bets on regulation and recursive self-improvement, but its product moat is thin—open-source harness startups are flooding in. The post doesn't spell out a clear countermove.

Why it matters: The K3 and Qwen 3.8 releases are notable, but the real value is the cost analysis: Fable 5 inference costs 3x competitors, and Anthropic's lack of owned infrastructure means costs scale linearly with usage. This is a concrete economic argument for open-source catching up, not ...

Hacker News front page

Kimi K3 tested on a Rust optimization task: one-shot success and a 2% speedup

The author asked Kimi K3 to optimize a hand-tuned Rust function. The model chose a SWAR approach, treating a 64-bit value as eight u8 values, completed the task in about 15 minutes with zero errors, and delivered a 2% speedup. Token cost was roughly $1 via OpenRouter. The author sees this as near-frontier performance.

Why it matters: A clean first-person experiment: the author threw a hand-tuned Rust function at Kimi K3, the model chose SWAR on its own, delivered zero-error code in 15 minutes, and netted a 2% speedup for ~$1. High signal density with concrete numbers and approach names—not marketing fluff....

Import AI (Jack Clark)

Open-weight cyber gap shrinks, Kimi K3 lands, Hassabis pitches AGI regulation

The UK's AISI found that GLM-5.2 and DeepSeek V4-Pro now trail closed frontier models by only 4–7 months on cyber tasks, down from 6–10 months in 2025. GLM-5.2 matches Claude Opus 4.6 on 70 narrow evals but falls further behind on long-horizon hacking ranges. Kimi released K3, a 2.8T-parameter model that scores near Claude Fable 5 and GPT 5.6 Sol, though the post hints at benchmark overfitting. K3 also wrote a GPU compiler and designed a chip in 48 hours; weights will be released in weeks. DeepMind's Demis Hassabis proposed a FINRA-style US standards body to test frontier models for national security risks, starting with voluntary 30-day pre-release reviews before moving to law.

Why it matters: AISI's first public quantification of the open-vs-closed cyber capability gap, with concrete model names and time deltas. Downside: this is a newsletter summary, not the original report, and the scope is limited to cybersecurity only.

AI HOT (Curated Pool)

Cursor's planner-worker agent swarm rebuilds SQLite in Rust, passing 80% of tests in 4 hours with widely varying costs

Cursor redesigned its agent swarm into a tree-shaped planner-worker split and retested it on rebuilding SQLite in Rust from docs. The new swarm beat the old one in every model config: with Grok 4.5 it hit 80% on a held-out SQL test suite in 4 hours, while the old swarm spiraled before hour two. Quality stayed similar across model mixes, but costs varied enormously—the post shows a comparison chart without exact dollar figures. The design isolates context so planners never see implementation details and workers only focus on narrow tasks, preventing drift on long runs.

Why it matters: First-party engineering experiment from Cursor, not a press release. The planner-executor tree architecture comes with concrete numbers (4 hours, 80% pass rate), hitting all three HKR axes. Not 85+ because this is a single experiment, not a product launch, and the post doesn't...

Computing Life · Share · Yage

Why coding agents need sandboxes beyond command approval

Approval gates only decide whether a command starts, not what happens after package managers load scripts and spawn child processes. The article walks through a bug-fix task to show how OS isolation (Seatbelt/bubblewrap), credential proxying (Docker Sandboxes), and isolated workspaces each address different risks. No performance numbers or latency figures are disclosed.

Why it matters: Hits all three HKR axes: the headline has genuine curiosity pull, the walkthrough of a full bug-fix task makes the sandbox-vs-alternatives comparison concrete, and it directly speaks to Cursor/Claude Code users who click that sandbox button daily. Docked a few points because i...

Jul 19Sunday

Bloomberg Technology

Moonshot AI plans IPO within six months after Kimi model breakthrough

Moonshot AI plans to IPO within six months, riding the momentum of its new Kimi K2 model. K2 matches OpenAI o3 and DeepSeek V4 Pro on math and coding benchmarks. The company is valued at about $3 billion, with roughly $150 million in 2025 revenue from Kimi chatbot subscriptions and API fees. The post doesn't specify the listing venue or underwriters. The six-month timeline hinges on market conditions and regulatory approvals—don't bank on it yet.

Why it matters: Moonshot sets a six-month IPO timeline with Kimi K2 matching o3 and DeepSeek V4 Pro as the trigger, backed by concrete valuation and revenue figures. Bloomberg exclusive, strong source. Capped below 85 because the exchange and underwriters aren't disclosed, and a six-month tim...

Computing Life · Share · Yage

Grok Build open-sourced its client harness, not the model or cloud

xAI released the Rust client harness that handles local files, commands, and permissions for Grok Build under Apache-2.0. The Grok model, cloud services, and the official binary build chain remain closed. The repo doesn't accept external PRs. The commit from the earlier upload controversy isn't in the public history, so the current code can't close that case. The real win: you can now pin a public commit, build it yourself, and compare its behavior against the official binary.

Why it matters: xAI open-sourcing Grok Build's client harness is substantive—Apache-2.0, headless mode, and ACP support go beyond signaling. But the model and build chain remain closed, and the repo rejects PRs, capping it below 85. All three HKR axes hit, so featured.

Hacker News front page

Kimi K3 matches Claude in daily coding, at a fraction of the price

The author ran Kimi K3 alongside Claude for coding and couldn't tell them apart on output quality or token usage. K3's API costs $3/$15 per million input/output tokens vs Claude's $10/$50. Subscriptions are even more lopsided: Kimi's $39 coding tier is far more generous, while Claude's $20 plan quietly dropped Fable access because the economics didn't work. The bigger story is US AI policy failure—restricting American models only constrains American customers, while frontier-quality Chinese models like K3 and GLM 5.2 ship without those limits. Semgrep found GLM 5.2 beating Claude on cyber benchmarks precisely because the restricted model declines work the open one just does. The author expects the US to repeat its auto-industry playbook: subsidies and tariffs propping up domestic models that can't compete internationally.

Why it matters: A hands-on developer comparison with real data: Kimi K3 matches Claude on code quality and token efficiency at a fraction of the price. Also calls out Claude's subscription bait-and-switch (Fable access removed). Score capped below 85 because it's a personal blog, not an offic...

Jul 18Saturday

Hacker News front page

Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal Help?

The author tested Claude Fable 5 and GPT-5.6 Sol on an unpublished fiber-network optimization problem, each with three 30-minute runs, comparing plain mode against /goal. Fable 5's plain mean was 32,386—1,875 points lower than Sol's 34,261—and its three plain runs stayed within a 319-point range, showing remarkable consistency. /goal won four of six trials but made both models' means worse: Fable 5 by 759 points, Sol by 868. The feature occasionally gives a small edge but can also cause large regressions. The post also breaks down how /goal differs under the hood: Claude Code uses Haiku as a transcript-only evaluator, while Codex has persisted state and lifecycle tools. Bottom line: Fable 5 is the real story here; /goal is not a safe default.

Why it matters: First-person experiment with concrete numbers across 3 runs per model. The counterintuitive finding that /goal mode destabilizes Fable 5 is worth surfacing. Docked slightly because the problem domain is narrow and this is a personal blog, not an official release.

Computing Life · Share · Yage

LLMs need a smaller executable world: DSL, FastAPI, and Context Infrastructure

This piece reframes DSLs as building a smaller executable world for LLMs, not just inventing syntax. The author walks through his own abandoned DSL project and argues that AI has lowered the cost of designing languages, writing parsers, and preparing examples—so teams can build domain boundaries first and decide their lifespan later. The core stack: FastAPI provides atomic actions, DSL describes complete plans, and Context Infrastructure stores past judgments and corrections. The article cites Unmesh Joshi's DSL post and the Tickloom demo, but notes Tickloom is an engineering showcase without cross-model benchmarks proving DSLs improve LLM accuracy. It warns that overly narrow boundaries can exclude correct answers—effect surface and composition rules matter more than Turing completeness.

Why it matters: Hits all three HKR axes: fresh angle, concrete architecture, speaks directly to agent engineering pain points. Score held at 72 because it's an opinion piece from a personal blog with no reproducible experiments or production case studies—high-quality thinking but not a hard r...

AI HOT (Curated Pool)

Cursor's eval lead confirms Claude Fable 5 hits 72.9% on CursorBench, targeting the hardest 1% of coding tasks

Cursor's eval lead Nate Schmidt explains on Anthropic's blog how they determined Claude Fable 5 was ready for the hardest 1% of real-world coding problems. The headline number is 72.9% on CursorBench, a significant jump over the prior generation. The post stresses this isn't a generic benchmark grind—it targets long-tail tasks that actually stump developers. The article doesn't disclose the baseline score, test set size, or sample problems, so treat the 72.9% as a directional signal rather than a cross-benchmark comparison point.

Why it matters: Cursor's eval lead publishes on Anthropic's blog with a concrete 72.9% CursorBench score — a substantive first-party eval. The post doesn't disclose the previous-gen baseline or test set size, so score lands at 82 rather than higher.

Jul 17Friday

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash Cyber security model

Google DeepMind released Gemini 3.5 Flash Cyber, fine-tuned from 3.5 Flash to find, verify and patch vulnerabilities quickly. With multiple calls, it approaches larger models on benchmarks such as CyberGym.

Why it matters: It reports how a lightweight security model performs on several benchmarks and inside Google's own codebase, so readers can judge the cost-benefit for vulnerability discovery.