Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

161–180 of 1,196

Sep 3Thursday

Hacker News front page

MBZUAI releases K2 Horizon, a six-model fleet with the 0.9B scoring over 48 on AIME 2026

IFM at MBZUAI released K2 Horizon, a six-model fleet from 0.9B to 375B-A23B. The 0.9B, 3.7B, and 7B models set new SOTA in their size classes; the 0.9B scored above 48 on AIME 2026 with reasoning and tool-use capabilities. The 36B-A4B uses a new MoVA attention mechanism, outperforming larger models per active parameter. This is a full open-science release: intermediate checkpoints, data recipes, code, logs, and evals from pretraining through agentic post-training, under Apache 2.0. The post doesn't disclose specific benchmark comparison numbers or latency data, so real-world performance still needs third-party validation.

Why it matters: IFM dropped six fully open models at once, with the 0.9B hitting 48+ on AIME 2026 math and the 36B introducing a new MoVA attention mechanism — high information density. Not scoring 85+ because IFM isn't an OpenAI/Anthropic-tier lab yet and market validation hasn't caught up; ...

Latent Space

Meta's Muse Spark 1.3 matches GPT-5.6-Sol, training at >90% discount

Meta released Muse Spark 1.3, now ranked #3 globally on AAII, directly competing with OpenAI and Anthropic's frontier models. Zuck called it their biggest jump yet on coding and agentic work, and promised open weights. Pricing is aggressive: opt into training and the cost drops by over 90%. Meanwhile, two new Stanford courses are teaching agent engineering from scratch, replacing 85% of old material with agent skills, context engineering, and security. Sebastian Raschka also tempered the Astra hype, pointing out that looped transformers aren't new—Nanbeige 4.2-3B already reused layers, trading ~2x compute for parameter savings without inherently hiding chain-of-thought.

Why it matters: Muse Spark 1.3 hits #3 on AAII, directly matching GPT-5.6-Sol, with Zuck promising open weights and a >90% training discount. This is Meta's first time cracking the top tier on a major benchmark, and it reshuffles the open-source landscape. Not a perfect score because it just ...

Computing Life · Share · Yage

OpenAI Codex's self-wake mechanism: it sets its own alarm to watch CI after fixing code

A system prompt template merged into OpenAI's open-source codex repo in late August reveals how Codex Persistent mode actually works: it's not a 24/7 always-on process, but a wake-check-sleep loop every 1–3 minutes. The template requires the agent to record its goal, latest status, completion condition, and next check time before sleeping, then decide what to do upon waking. One hard rule: persistence does not broaden authorization scope—anything beyond scope requires explicit permission. WIRED reported on this mode earlier, but media headlines saying 'always-on' clash with the code's 'sampled again' language. OpenAI hasn't launched it yet; the backend request still shows 'disabled.' ProAgentBench shows models achieve only 64.4% accuracy in judging when to proactively help, and Anthropic's engineering blog reports a 17% miss rate on real overreach during automated review—two numbers that explain the hold. Tasks suited for it are delivery-type jobs like CI, deployment, and builds that execute for one minute and wait for ten. Open-ended tasks like writing proposals or designs are a bad fit. Three discipline rules from the template can be adopted today: write four-element checkpoints, stay silent when nothing has changed, and prefer deterministic mechanisms.

Why it matters: High information density with concrete sourcing from the open-source repo — reveals the real wake-check-sleep loop and the authorization scope rule. Deduction because this is interpretation of a template, not an official launch; actual product experience is unknown.

Computing Life · Share · Yage

Agent token usage 5× human, but caching discounts cut the real bill to ~2×

OpenRouter data shows agents consume 7.3T tokens weekly, nominally 5.2× human usage. But 70–85% are cached reads; with ~90% discount, the real bill is roughly 2×. GitHub's Knowledge Compressor prototype halves doc length and claims breakeven at 2,000 reuses, but factoring in caching pushes the median to 5,000+. OpenAI's Jalapeño chip beats Nvidia GB200/GB300 on fixed-length benchmarks, yet lacks AgentX scores for real agent workloads. All three stories share one distortion: prompt caching inflates headline numbers.

Why it matters: Three stories bundled, but the core value is the first: someone finally separated nominal agent token consumption from the caching-discounted real cost, landing at ~2x. The OpenAI chip benchmark and GitHub compression prototype are bonuses but less dense. Cross-source cluster ...

Hacker News front page

AI agents don't get lost in messy code, so the refactoring reflex disappears

Rodrigo Rosenfeld Rosas argues that AI coding agents remove a critical safeguard: the moment a human gets lost in tangled code and decides to refactor. Agents never get lost, so they keep adding branches to an unmanageable mess. The short-term speed hides a long-term cost—teams lose the ability to reason about their own systems, reviews become rubber stamps, and messy code burns more tokens per change while increasing hallucination risk. He urges teams to deliberately reinstate the checkpoint the agent won't trigger.

Why it matters: A sharp practitioner observation, not another generic AI-code-quality take. It identifies a neglected mechanism: AI removes the human instinct to call for a refactor when logic gets tangled. Fresh angle, concrete mechanism, strong resonance—but it's a personal blog, not an ind...

Hacker News front page

Meta launches Muse Spark 1.3, tuned for agentic workflows and competitive coding

Meta's Muse Spark 1.3 is built for agentic workflows: it handles long-horizon tasks, calls tools reliably, and asks for clarification on messy inputs. It's tuned for higher first-attempt coding accuracy and competes with frontier models on several coding evals. The model natively perceives video, images, and documents. Pricing: $1.25/M input tokens and $4.25/M output tokens for the standard tier; a contributor tier costs $0.10/M input. Both offer a 1M context window. The post doesn't spell out specific benchmark scores, only a chart.

Why it matters: Meta ships Muse Spark 1.3, targeting long-chain agent tool calling and first-attempt coding accuracy with clear pricing. A substantive model update from a major lab, but the post lacks benchmark data and technical specifics to back the 'competitive with top models' claim, so i...

Hacker News front page

Meta releases Muse Spark 1.3 with better agentic and coding performance

Meta launched Muse Spark 1.3 today on Muse Code and Meta Model API. The model handles longer multi-step tasks by asking clarifying questions, requesting help when stuck, and confirming before taking consequential actions. Benchmarks show it beats Muse Spark 1.2, GPT 5.6 Sol (max), and Opus 5 (max) on agent, coding, instruction-following, and long-context evals. Two demos are included: one generates a CFD simulation report from CAD files and exports it as a PDF, another edits bass guitar mistakes in a multi-track session. The max reasoning mode is still undergoing safety testing and will ship later.

Why it matters: Meta ships Muse Spark 1.3 with agent/coding benchmarks beating GPT 5.6 on several metrics, plus three concrete interaction mechanisms that make agent deployment more practical. Held below 85 because it's an iterative release, not a new architecture, and max reasoning mode is s...

AI HOT (Curated Pool)

GitHub Copilot cuts AI coding costs to one-third with preference-trained small models

GitHub published an engineering blog detailing how they cut Copilot's AI coding costs to roughly one-third without hurting task quality. The key move: training a 1.8B-parameter model on 1,040 preference pairs to act as a router that decides when to use a cheap model and when to call a stronger one. After rollout, strong-model calls dropped 70% and overall latency stayed under 11 seconds. The post also mentions a training method called DV-DPO that uses preference data to teach a small model a specific response style. One caveat: these numbers come from GitHub's own setup, so your mileage may vary.

Why it matters: GitHub shared a concrete cost-optimization engineering post with real numbers and methods, directly useful for teams shipping AI products. Score capped because it's an engineering optimization, not a new model release.

Hacker News front page

Same model, 9 harnesses: cost per pass varies 17× in FrontierHarness Eval

Runta benchmarked 12 harness configs on Kimi K3 with identical cold-start environments across 360 runs. Codex led at 66.7% pass rate and $3.47 per task; Exo Harness was cheapest at $1.05 with 53.3% pass rate; Claude Code hit 63.3% but cost $18.34 per task. Cache hit rate doesn't equal savings—Claude Code had the lowest cache hit rate at 67.8% yet the highest cost per successful task at $0.288. The post doesn't disclose which specific software engineering tasks were used or their difficulty distribution.

Why it matters: 360 cold-start trials, same Kimi K3 model, 12 harness configs, 17x cost spread — the cleanest coding-agent benchmark I've seen. Claude Code at $18.34/task with 63.3% pass rate vs Codex at 66.7%/$3.47 is a sharp contrast. Not scoring higher because it's a single blog post with ...

Sep 2Wednesday

AI HOT (Curated Pool)

Cursor launches Self-Hosted Machines so cloud agents run on your own infrastructure

Cursor cloud agents can now execute tool calls on machines inside your network while inference and planning stay in Cursor's cloud. Teams register their own machines via a worker that maintains an outbound HTTPS connection, giving agents direct access to internal repos, private services, and custom hardware like GPUs or Macs. Cursor says over 60% of its internal PRs are already created by cloud agents, and this targets enterprises that need network isolation or specialized infrastructure.

Why it matters: Cursor decouples cloud agent execution from its own infra, letting enterprises keep code and GPUs on-prem while still using the cloud brain. It's a real architectural shift, not a minor tweak. Score stays at 78 rather than higher because it's launch-day with no user validation...

Hacker News front page

Multiverse Computing releases Quasar 438B, the highest-scoring European model on Artificial Analysis

Multiverse Computing launched Quasar 438B, its first large model, a reasoning model for enterprise agents and coding that supports English and Spanish. It scores 43 on the Artificial Analysis Intelligence Index, the highest among European models, ahead of Mistral Medium 3.5 at 30 and NVIDIA Nemotron 3 Ultra at 38. It outputs 500 tokens in 15.3 seconds, faster and smarter than Mistral Medium 3.5. Long-context reasoning hits 75.0, close to Claude Opus 5 at 75.7. Terminal-Bench v2.1 scores 69.3, leading Mistral by 18.7 points but trailing Claude Opus 5 at 89.1. The model is available via the CompactifAI API. The post does not disclose training data, detailed parameter count, or pricing.

Why it matters: First 438B reasoning model from Europe with concrete benchmark numbers and competitor comparisons — enough signal. But the source is the company's own blog, no third-party testing yet, so score stays at the featured threshold.

Latent Space

Anthropic drops Claude Fable/Mythos 5.1: new SOTA for coding, but 70% more output tokens

Anthropic launched Claude Fable 5.1 and Mythos 5.1 on Sep 1, claiming SOTA on coding and knowledge work. Fable 5.1 hits 55.8% on Terminal-Bench 4.0 and is pitched for autonomous multi-step tasks. Cache read price dropped 75% to $0.25/MTok, but Artificial Analysis found output tokens rose 1.7x, netting a ~20% per-task cost increase. Community speculation suggests Fable and Mythos may share weights with different safety routing—the post doesn't confirm this. Early praise for coding ability is offset by complaints about rate limits, false safeguard triggers, and subscription UX.

Why it matters: Anthropic dropped Claude Fable/Mythos 5.1 with a 55.8% Terminal-Bench 4.0 score, a 75% cache read price cut to $0.25/M tokens, and a 70% increase in output tokens. A capability upgrade plus major pricing shift makes this a same-day must-write. Not a 95 because we only have Lat...

AI Chat-Group Daily (群聊日报)

DeepSeek V4 Flash beats Sol in real-world use; Anthropic drops Fable 5.1

Community members ran two-month SBS comparisons and a week-long 5.1B-token workload on DSH + DeepSeek V4 Flash, concluding it feels better than GPT-5.6 Sol in real tasks. Sol overthinks and produces bloated output; V4 Flash is fast (2.3s first token) and cost ¥362.84 total. A 'subscription gym paradox' theory argues subscription-based harnesses quietly throttle usage while pay-per-token models don't. Anthropic launched Fable 5.1 with 75% cheaper cache reads, but Fable 5 scored below Opus 5. Also: Astra hits Critical cybersecurity tier, Anthropic's $35B compute deal, Qwen 3.8-Max-0902 benchmark run, Microsoft AI secretary setup, and Grok Bot hands-on.

Why it matters: The side-by-side data is solid — 5.1B tokens, ¥362.84 total spend, 2.3s first-token latency — but the source is an anonymized chat log, not an official release or reproducible benchmark. That caps the authority. HKR all hit, so featured is the right tier.

AI HOT (Curated Pool)

Qwen3.8-Max-0902 tops Code Arena and leads the Pareto frontier at $5/MToken

Alibaba Qwen's new Qwen3.8-Max-0902 scored 1,691 on Code Arena's WebDev leaderboard, ranking first overall. At a blended price of $5/MToken, it's the highest-scoring model on the Pareto frontier. Available now on QwenCloud. The post doesn't disclose further technical details or comparison data.

Why it matters: Qwen3.8-Max-0902 tops Code Arena's overall leaderboard with a $5/MToken blended price and a Pareto-frontier claim — a substantive domestic flagship model update that earns the positive-signal bump. HKR all hit: topping the chart creates suspense, concrete score and pricing add...

AI HOT (Curated Pool)

Qwen releases Qwen3.8-Max-0902: 2.4T parameters, 1M token context window

Qwen3.8-Max bumps to the 0902 version with 2.4T parameters and a 1M-token context window. Post-training focuses on coding and cowork, targeting complex enterprise tasks, scientific research, and long workflows. The post doesn't include benchmark comparisons or pricing.

Why it matters: Alibaba Qwen drops a new flagship: 2.4T params, 1M context, post-training aimed at coding and long-chain collaboration. Domestic flagship release gets featured-tier treatment per policy. No benchmarks or pricing disclosed, so real competitiveness is unclear — score held at the...

Hacker News front page

Simon Willison tests Claude Fable 5.1's pelican benchmark across five reasoning levels

Simon Willison ran his classic 'SVG of a pelican riding a bicycle' prompt against Claude Fable 5.1 at five reasoning levels. Low and medium produced near-identical outputs with no visible reasoning, taking ~23 seconds and ~10 cents. At xhigh the model spent 7m51s and $1.83, adding real detail. Max ran for 13m54s and $3.30, delivering his best Anthropic pelican yet—blue hat, basket with a fish, feet on pedals—though he still says it lacks the flair of Gemini 3.7 Flash. Separately, Fable 5.1 hit 52.6% on the new Terminal-Bench-Science 0.1 benchmark, up from 24.7% for Fable 5.

Why it matters: Simon Willison ran a controlled five-tier reasoning comparison on Claude Fable 5.1 with concrete latency and cost numbers, making it more useful than the official announcement. Score stays below 85 because this is a personal evaluation rather than a major capability breakthrou...

The Verge · AI

Anthropic launches Claude Fable 5.1, up to 45% cheaper for agentic work

Anthropic released Fable 5.1 and Mythos 5.1, directly addressing customer complaints about cost, data retention, and overzealous safeguards. Fable 5.1 outperforms Fable 5 while costing ~25% less typically and up to 45% less for complex agentic tasks, driven by lower pricing on cached data. Every CEO Dan Shipper called it the strongest coding model they've used, now fast, token-efficient, and speaking like a normal person. The post doesn't spell out Mythos 5.1 specs or detailed pricing.

Why it matters: Anthropic drops Fable 5.1 and Mythos 5.1 with a clear cost-reduction story for agent workloads — up to 45% cheaper via cached call pricing. Concrete performance and pricing details make this a strong signal. Held at 85 rather than higher because we only have the headline and s...

Hacker News front page

Rewriting 65k lines of Go to Rust with Fable cost $400

The author rewrote a 65k-line terminal editor from Go to Rust using Fable 5 for $400. The method has three steps: extract code into a data representation (state machines, graphs, formulas), operate on that representation, then regenerate code in the target language. Fable's precise data-flow tracing is the key enabler. The post doesn't report compilation pass rate or test coverage, so I'd discount the 'fully autonomous' claim until those numbers surface.

Why it matters: 65k lines Go-to-Rust for $400 with a clever intermediate-representation approach hits H and K. But the post doesn't disclose compile pass rate or test coverage, so 'fully automated rewrite' needs a discount — lands at 72, right at the featured threshold.

AI HOT (Curated Pool)

Claude Fable 5.1 lands on Claude Code and Platform, cache reads 75% cheaper

Anthropic shipped Claude Fable 5.1 and Mythos 5.1 together. Pricing matches Fable 5, but API cache reads are 75% cheaper. The model stays autonomous longer on long tasks, flags when it's stuck more proactively, and writes more naturally. The post doesn't disclose latency, context window, or benchmark scores—I'd discount the 'most advanced' claim until numbers land.

Why it matters: Anthropic shipped Fable 5.1 and Mythos 5.1 together with a 75% cache-read price cut — a real cost improvement that heavy Claude Code users will care about. Missing latency, context window, and benchmark numbers keeps it from scoring higher, but the price drop and tooling updat...

Hacker News front page

Anthropic launches Claude Fable 5.1 and Mythos 5.1, cutting price by 25% and targeting coding and scientific research

Anthropic released two models, Fable 5.1 and Mythos 5.1—same underlying model, different safeguards. Fable 5.1 is generally available; Mythos 5.1 is gated behind trusted access programs for cybersecurity and life sciences. Fable 5.1 beats Fable 5 across coding, knowledge work, and long-horizon tasks, while costing ~25% less on typical workloads and up to ~45% less on highly agentic work. Enterprise Frontier Safeguards (EFS) will let customers keep data in their own cloud infra, rolling out in phases from fall 2026; until then, eligible customers get zero data retention. Cybersecurity false positives dropped 60%, and the model can discover vulnerabilities but not build exploits. The post does not disclose parameter count, context window, or training details.

Why it matters: Anthropic's flagship model refresh with a dual-track release (Fable 5.1 for everyone, Mythos 5.1 gated behind trusted projects) is an industry first. Coding and long-horizon tasks beat the previous gen across the board, and bio capabilities are strong enough to require governm...