Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

61–80 of 1,196

Sep 23Wednesday

AI HOT (Curated Pool)

Claude Opus 5.5 lands on OpenRouter with better agentic coding and a 20% price cut vs Opus 5

Anthropic released Claude Opus 5.5 on OpenRouter, the first model in the Claude 5.5 series. It beats Opus 5 and Fable 5.1 on agentic coding, knowledge work, and computer use, with a 1M context window. Pricing is $4 per million input tokens and $20 per million output tokens, 20% cheaper than Opus 5. The post doesn't include benchmark scores or latency figures.

Why it matters: Anthropic's flagship Claude Opus 5.5 lands on OpenRouter as the first 5.5-series model, with explicit gains in agentic coding and computer use, plus clear pricing. Hits all three HKR axes — a same-day must-write. Not scoring higher because only the platform announcement is ava...

AI HOT (Curated Pool)

Anthropic releases Claude Opus 5.5, ~30% faster and ~40% cheaper

Claude Opus 5.5 is the first model in the Claude 5.5 family. It matches Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5. Claude Devs adds it's ~30% faster per task. Claude Code's 5-hour session limit increased 20% today; lower pricing means 25% more usage within the cap. Pro, Max, and Team users also get a one-time quota reset. Terminal-Bench 4.0 scores lead across effort tiers.

Why it matters: Anthropic flagship model update with a double jump in speed and cost — a same-day must-write. Score stays below 90 because we only have the official tweet and community notes so far, no third-party benchmarks or cross-model comparisons yet.

AI HOT (Curated Pool)

Anthropic launches Claude Opus 5.5, matches Fable 5.1 performance at 40% lower cost

Anthropic dropped Claude Opus 5.5, the first model in the 5.5 family. It matches Fable 5.1 on most tasks and costs 40% less to run than Opus 5. The author notes clearer communication, better token efficiency, and availability across all effort levels. The 5-hour rate limit is raised and a banked reset feature is added. The post doesn't disclose specific benchmarks or pricing.

Why it matters: Anthropic drops Claude Opus 5.5, claiming Fable 5.1-level performance with 40% lower running cost vs Opus 5, plus a raised rate limit and banked reset. A substantive flagship update that directly addresses long-standing user complaints about cost and limits. Not scoring higher...

AI HOT (Curated Pool)

Anthropic launches Claude Opus 5.5 with lower cost and better token efficiency

Anthropic released Opus 5.5, the first model in the Claude 5.5 family. The company says it matches Claude Fable 5.1 on most tasks, costs 40% less to run than Opus 5, and has lower per-token pricing with more efficient token usage. It supports all effort levels and is already available in Claude Code. The post doesn't disclose exact pricing or benchmark comparisons.

Why it matters: Anthropic's flagship model refresh with 40% cost reduction matching Fable 5.1 is a direct win for Claude ecosystem users. Score held back because the post doesn't disclose actual pricing or benchmark numbers — real savings need real tests.

TechCrunch · AI

Anthropic releases Opus 5.5 with lower prices and Fable-level performance

Anthropic launched Opus 5.5 on Tuesday, calling it “the strongest-performing model we've tested to date.” The company claims new state-of-the-art results in coding and knowledge work, with lower prices than previous Opus models. The post doesn't disclose specific pricing, benchmark scores, or a direct comparison with Fable, so I'd hold off on the “strongest” claim until third-party evals land.

Why it matters: Anthropic's flagship model update with a price cut and Fable-level performance claim is a real signal. But the post doesn't disclose actual pricing or benchmark numbers — the two most critical pieces — so the score stays below 85.

Sep 22Tuesday

Hacker News front page

JetBrains Air: A product system for agentic software development

JetBrains consolidates six months of agentic development experiments into Air, an open system spanning developers, teams, and orgs. It goes beyond the IDE with multi-surface, multi-service design, betting on a multi-vendor future. The post confirms Central CLI, shared context, cloud agents, automations, and AI cost controls are already rolling out, but pricing and GA dates aren't disclosed.

Why it matters: JetBrains officially launched Air, a product system that upgrades AI coding from an IDE plugin to a cross-tool, multi-model platform, with named components like Central CLI, shared context, and cloud agents. It's a heavyweight response to agentic coding from a legacy tool vend...

Latent Space

Xiaomi MiMo-V2.6-Pro tops open weights leaderboard, trained for $3M

Xiaomi released the MiMo-V2.6 series. The Pro version ranks #1 among open weights models on the Artificial Analysis Intelligence Index with a score of 46, at a training cost of $3M. A Flash variant targets efficiency, and an UltraSpeed variant offers 20x faster output. The technical report details RL scaling across three axes: larger batches and throughput, richer multi-task environments, and more grader compute. Code and training recipes are open-sourced, but the 7k+ task datasets are not yet released. Former DeepSeek engineer Fuli Luo, now at Xiaomi, previously live-streamed the training runs.

Why it matters: Xiaomi's MiMo-V2.6-Pro hit #1 on the Artificial Analysis open-weights leaderboard with a $3M training budget — price-performance right at the frontier. Flash and UltraSpeed variants cover efficiency and speed use cases, and the tech report details an async RL architecture. Not...

Computing Life · Share · Yage

Three AI Coding Stories, Three Numbers to Read

ZCode was caught silently packaging entire Git histories for cloud upload, with .git objects making up 86.6% of snapshots. HarnessTax benchmarked Claude Fable 5 across frameworks: Claude Code and minimal Pi achieved near-identical success rates but a 2x cost gap. A Microsoft architect replaced multi-model agent loops with a single model reading skill docs and calling tools directly—halving API calls but increasing total tokens by 22%. Each story unpacks one number and a reminder to check which layer a metric actually measures.

Why it matters: Three distinct AI coding stories bundled into one piece. The ZCode silent snapshot upload is a hard security story backed by ferstar's forensic report and community reproduction. HarnessTax benchmark and Microsoft architecture case add billing and engineering angles. HKR all h...

AI HOT (Curated Pool)

NVIDIA Nemotron 3.5 Lightning: a 30B sparse model built for high-frequency agent execution

NVIDIA positions Nemotron 3.5 Lightning as the execution layer in agent workflows—handling frequent tool calls, file reads, and result checks rather than heavy planning. It's a 30B MoE model that activates only ~3B parameters per token, keeping latency and cost low for high-volume calls. It complements, not replaces, Nemotron 3 Ultra. Weights are open, with tool calling and structured output support. Context goes up to 1M tokens, though OpenRouter's standard tier caps at 262K. Worth a look if your agent makes many model calls per run.

Why it matters: OpenRouter's breakdown of NVIDIA's new model is substantive, clearly explaining high-frequency agent calls and MoE architecture choices, but the topic is engineering-focused and lacks an emotional hook — R missed.

AI HOT (Curated Pool)

Xiaomi MiMo-V2.6-Pro hits ~10th on Code Arena WebDev, ~3rd among open-weight models

Xiaomi released two omni-modal models: MiMo-V2.6-Pro and Flash. The Pro version scored 1628 on Code Arena WebDev, up 153 points from MiMo-V2.5-Pro's 1475, landing around 10th overall and ~3rd among open-weight models under MIT license. The post doesn't disclose Flash's benchmark numbers or parameter counts.

Why it matters: Xiaomi's multimodal model hits ~10th on Code Arena WebDev and ~3rd among MIT open-weight models, with a 153-point gain for Pro. Flags a domestic flagship release with concrete benchmark data. Flash variant lacks params and scores, capping it below 80.

AI HOT (Curated Pool)

Xiaomi releases MiMo-V2.6 Pro and Flash, two fully multimodal open-source models

Xiaomi MiMo dropped two fully multimodal open-source models. The Pro version matches Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index—the highest among open-source models so far. Capabilities span coding, computer use, 3D reasoning, and creative tasks. The post doesn't disclose parameter counts, training details, or where Flash sits in the lineup, so I'd hold off on direct comparisons for now.

Why it matters: Xiaomi released MiMo-V2.6 Pro, a fully open-source multimodal model that matches GPT-5.6 and Claude Opus 5 on agent benchmarks, scoring 46 on the Artificial Analysis Intelligence Index—the highest for any open model. Domestic flagship launch with concrete numbers and direct co...

Sep 21Monday

AI HOT (Curated Pool)

xAI launches Grok 4.7, twice as fast as Grok 4.6 at the same price

Grok 4.7 uses a larger base model and a longer RL run on harder, multi-hour tasks. It scores 46.3% on CursorBench 4.0, ahead of GPT-5.6 Sol Max (41.7%) but behind Fable 5.1 Max (51.8%). Pricing stays at $2/$6 per million input/output tokens, same as Grok 4.6, with double the speed. Safety stack is new: only 3.3% of risky cyber prompts get through, and it hits 62.4% on LatchBio's biosafety benchmark. Available today in Cursor, Grok Build, and the API.

Why it matters: xAI drops Grok 4.7 targeting coding and knowledge work, hitting 46.3% on CursorBench 4.0 — above GPT-5.6 Sol Max but behind Fable. Concrete benchmark and training details clear all three HKR axes. Held below 85 because the post doesn't disclose model size, architecture changes...

Hacker News front page

Will Larson tries the software factory pattern at Imprint, letting agents own goals, fill gaps, and push PRs

Will Larson pushed his agent setup at Imprint further: give an agent a Linear project, and it first checks for a Notion RFC and Datadog/Snowflake dashboards—prompting you to create them if missing. It then scans task status, adds newly identified work, and picks up unblocked tasks to write PRs, nudge reviews, or ask clarifying questions. After a task completes, if the project description is stale, it re-runs the full loop. Larson says this forced him to hand over goal-state he used to hoard, so agents can now judge direction. He plans to move this loop from local to the company's internal 'Agent Fleet' orchestrator. The post does not disclose performance numbers or cost.

Why it matters: Will Larson's 'software factory' experiment is one of the most grounded first-person accounts of AI coding adoption in 2026. From company-wide Claude Code rollout to Agent Fleet orchestration, every step has concrete decisions and failure modes—directly useful for teams pushin...

Sep 20Sunday

Hacker News front page

Microsoft used AI agents to port Copilot runtime from C# to Rust for $120K

A Microsoft team used an in-house AI agent system called Nachete to rewrite the Copilot runtime from C# to Rust, at a total cost of about $120K. The agents ran 7 iterations—writing code, compiling, and fixing errors—producing 125K lines of Rust that compiled on the first try. Humans only did code review and security audit. The team estimates a manual rewrite would have cost $1.1M and taken 9 months, though the post doesn't detail how that baseline was calculated.

Why it matters: Microsoft used an internal agent to port the Copilot runtime from C# to Rust: 7 iterations, 125K lines, first-try compilation, $120K cost. The human baseline of $1.1M/9 months isn't explained, so I'm discounting that claim. Hits all three HKR axes but it's a single engineering...

Hacker News front page

The Chief of Staff Pattern: One Claude Code session coordinates, others execute

This post describes a pattern for running long Claude Code sessions reliably: separate coordination from execution. One long-lived session assigns work, verifies claims, and records lessons, while short-lived sessions do the actual coding. State lives in a durable external board, not in context. The key discipline is to never trust an agent's self-report—re-run the commands and check exit codes. cmux is used to spawn execution workspaces. The pattern is essentially orchestrator-worker; the author calls it Chief of Staff but notes it's different from Anthropic's calendar-managing agent of the same name.

Why it matters: A practical engineering pattern piece with real substance, not generic advice. The author splits long-running Claude Code work into coordinator + executor layers, uses an external board instead of conversation context for state, and the core discipline is 'don't trust agent se...

Hacker News front page

StepFun launches Step 5 Preview, a 600B MoE flagship model targeting coding and finance

StepFun introduces Step 5 Preview, a 600B-parameter MoE model with 27B active per token, a 1M-token context window, and vision support. It scores 67.7 on DeepSWE v1.1, ahead of Kimi K3 and GLM-5.3 but behind GPT-6 Astra and Claude Opus 5. On the in-house StepCodeBench it hits 49.0, again leading domestic models and trailing the two US labs. On FrontierFinance it reaches 66.4, second only to Claude Opus 5. Artificial Analysis gives it an intelligence index of 44; StepFun claims substantially lower cost per task at comparable intelligence. The post does not disclose API pricing, release timeline, or training details.

Why it matters: StepFun's Step 5 Preview is a 600B MoE model that edges out Kimi K3 and GLM-5.3 on coding benchmarks but still trails GPT-6 Astra and Claude Opus 5 by 6-7 points. Scored 78 because it's a substantive domestic model push in agentic coding with real numbers, but not industry-sha...

Sep 19Saturday

Computing Life · Share · Yage

Anthropic postmortem: when AI writes code too fast, patching test infra stops paying off

Anthropic's test-impact-analysis service saw 25× load growth in six months. Three patches bought 70 days, 29 days, then less than a day of stability. One engineer rewrote it in three weeks—a task the author estimates would have taken a quarter a year ago. The rewrite cost dropped while the hidden cost of patching rose, shifting the break-even point earlier. The post does not disclose the new system's exact running cost, defect rates, or production incident data.

Why it matters: First-person postmortem from an Anthropic engineer with concrete numbers and a decay curve across three patches—not generic AI productivity fluff. Hits all three HKR axes, but as an engineering practice piece rather than a product launch or model breakthrough, it lands in the ...

TechCrunch · AI

A ChatGPT inventor built Jev, a model that runs code instead of chatting, and developers are excited

Diogo Almeida, a former OpenAI researcher who co-invented RLHF, built a model called Jev that optimizes for code execution rather than human language. It's still a transformer, but TypeSafe AI designed it to run programs and call APIs directly, acting more like an automation agent. Developer reception has been enthusiastic. The post doesn't disclose benchmark scores, pricing, or whether weights will be open.

Why it matters: First model from an RLHF inventor's new startup, with a genuinely different training objective. But the TechCrunch piece lacks benchmarks, pricing, or API success rates — strong signal, soft on hard numbers, so 78.

Hacker News front page

There's no point at which turning your brain off will work

Dan Luu notes a growing trend of developers blindly trusting LLM outputs and acting as a 'meat proxy' in a loop. By September 2026, this brain-off approach can produce barely functional software, but Luu argues that if LLMs get good enough to work unsupervised, companies will just run the loop themselves and lay off the human. He shares concrete failures, including an AI bot weaker than a simple heuristic bot and a commercial product trapping users in an infinite loop. Luke Burton adds that high-value tasks still require constant supervision due to too many unknown unknowns.

Why it matters: Dan Luu coins 'meat proxy' to name the core tension in AI coding: the better models get, the more replaceable brain-off devs become. Sharp take with a Sept 2026 timestamp, but lacks hard data — 78.

Sep 18Friday

AI HOT (Curated Pool)

Justin Cormack on AI Agent Evaluation: Start With Evidence, Not Coverage

Justin Cormack built an S3-compatible storage system with AI, reaching 350k lines of Rust. He ran 1,500 tests against real S3 as an oracle, which caught real S3 500 errors. Chasing 100% coverage backfired—agents wrote trivial tests. Docs were often wrong, and AI was bad at finding edge cases from them. His hard rule: fix flaky tests immediately, or the agent learns to ignore failures.

Why it matters: A first-person experiment from Justin Cormack with real numbers and documented pitfalls—not generic commentary. The 350k-line Rust + 1,500 test case scale gives the findings weight. Downside: the post is ultimately Tessl brand content, so it doesn't hit 85+. But the experiment...