Skip to content

#编码

10 today

Sep 26Saturday

Hacker News front page

A developer's confession after one month without AI: dumber, lazier, and losing control

After a month without AI coding tools, the author looks back: it started with asking AI to write a function, then escalated to feeding entire Jira tickets to multiple agents in parallel. He realized he hadn't written a single line of code or even a commit message himself in months. Every AI-generated PR took two days to review, fix style, and add tests—far longer than doing it manually. Code quality kept dropping, requiring repeated prompting. He calls the perceived speedup an illusion that left him exhausted and disconnected from his own code.

Why it matters: A first-person experiment with concrete numbers, not vague complaining. The 'two days per PR to review' cost is a rare quantified pushback against AI coding hype. Not scored higher because it's a personal blog, not an industry event, and the body was truncated, leaving the ful...

Hacker News front page

How to keep enjoying programming in a world of LLMs

A Haskell programmer argues that letting LLMs write all your code kills the joy of programming and turns your codebase into an alien wasteland. His advice: keep writing code yourself, and use LLMs only for boring tasks like planning, testing, or documentation. The post doesn't specify tools or workflows, but the core idea is clear—don't outsource coding, outsource chores.

Hacker News front page

What Even Is an OS Now?

Thomas Ptacek left Fly.io to build a phone designed for AI-generated, single-user apps. He argues AI is dissolving the boundary between programmers and users, so most software will soon be conjured by its own users. When apps aren't from strangers, the OS's core job of isolating them makes less sense. The post does not disclose specs, pricing, or a launch date.

Why it matters: Thomas Ptacek announces his departure from Fly.io to build a phone, arguing that AI-generated ephemeral software undermines the OS's core isolation model. Fresh argument with concrete technical intuition, not hand-waving. Capped at 78 because it's a personal blog departure pos...

Hacker News front page

Benchmarking frontier models by porting Prince of Persia from 6502 assembly to C#

The author fed the original 6502 assembly source of Prince of Persia (Apple II, 1989) to frontier models and asked them to port it to C#. Claude Opus 4.6 produced a tile-grid engine that didn't play like the real game. OpenAI Codex fixed rendering details but left the broken architecture. Claude Opus 5, given DOSBox screenshots, diagnosed the architecture problem, rebuilt the engine, and read real animation frames and all 15 levels directly from the DOS game files. Claude Opus 5.5 ported the community-reverse-engineered room-drawing routine and reduced pixel differences on level 1 from 8,429 to 2. The post does not disclose Opus 5.5's exact release date or per-call cost.

Why it matters: A first-person experiment that stress-tests three Claude generations against the same 6502 assembly source, with failure modes specific enough to learn from. Not a product launch or industry event, so it stays in the good-quality band, but the signal-to-noise ratio is excellen...

Hacker News front page

Meta's Muse coding agent appears to route some tasks to an OpenAI model labeled muse-special

A developer digging through Muse's local files found a model called azure/muse-special that uses OpenAI's GPT Responses API. Nearly all sessions run on Meta's in-house Avocado model, but at least one sub-agent task was routed externally. The shipped daemon also bundles clients and API keys for Claude Opus 4.6/4.7/4.8, Sonnet 4.6, and GPT-5.5/5.6, with a kill switch to disable the external proxy. The author believes muse-special is likely a GPT model on Azure, though the exact version isn't disclosed. External reasoning chains are encrypted and unavailable to Meta, so distillation seems unlikely; Avocado's reasoning is stored in plaintext and usable for RL.

Why it matters: First-hand reverse-engineering find with concrete file names and routing evidence — not speculation. Meta's in-house Avocado handles most tasks but at least one sub-agent routes to OpenAI, plus bundled Claude Opus versions. Docked because it's a single-source blog without Meta...

TechCrunch · AI

OpenAI Astra and Anthropic Opus just cracked unsolved WWII Enigma messages

Two cryptanalysts used OpenAI's Astra and Anthropic's Opus to decode two Enigma messages that had remained unbroken since WWII. Developer Carter Leffen had Astra search archives, find context clues, build an Enigma simulator, and recover the plaintext. The post doesn't spell out Opus's exact role, nor the time taken or accuracy rate.

Why it matters: The story has strong narrative pull and a concrete knowledge hook in Astra's autonomous simulator-building. But Opus's role and key metrics are missing, and historical codebreaking is far from daily AI workflows, capping the score at the featured threshold.

Sep 25Friday

Ben's Bites

Ben built a 50-year device timeline in one morning with AI

Ben Tossell spent a morning building a site with Codex and Factory that shows 87 iconic devices from 1976 to 2026. All images were generated by Astra, using 40 messages and 7 subagents, and it has received 1,455 votes so far. The project started from a 'forgotten devices' idea, took inspiration from Cole's timeline scrubber demo, and had agents find reference photos, generate new images, and fill gaps after 2009. The post doesn't detail each device's model or generation cost.

AI HOT (Curated Pool)

Microsoft unveils new Copilot with Home, Code, and Autopilot capabilities

Microsoft repositions Copilot as 'the AI built for work' with three new modules: Home acts as a unified hub connecting people, teams, and projects; Code brings Copilot into coding, bug fixing, and code review; Autopilot embeds AI into business workflows to automate tasks like approvals and data entry. The post doesn't specify a launch date or pricing beyond 'today we're introducing.' I'd discount the hype a bit—big-org product announcements often paint a vision first, and real rollout cadence depends on follow-up updates.

Why it matters: Microsoft is giving Copilot a clear product redefinition with three concrete modules, not a concept paper. But the post lacks launch dates and pricing, so real-world impact is still unclear — score sits right at the featured threshold.

The Verge · AI

Microsoft redesigns Copilot as a super app bundling chat, coding, and agents

Microsoft officially unveiled its redesigned Copilot app today, merging chat, coding, and AI agents into a single interface. Home combines Copilot Chat and Cowork, while Code and Autopilot get their own tabs. Scout, the personal assistant shown at Build, is now rebranded as Autopilot. A Today dashboard feature is also planned. The post doesn't specify rollout dates or availability.

Why it matters: Microsoft's Copilot redesign merges chat, coding, and agents into one app, with a bold claim of Office-level influence. The structure is cleaner, but 'super app' feels like marketing, and no breakthrough capability is shown yet. Score at 72, pending hands-on reviews.

Hacker News front page

Claude converts Casio synth patches to web in five minutes

While building a web-based Casio CZ-101 emulator (CZP-1), the author found a YouTube demo by oliveoil22 and wanted those sounds. He downloaded the .syx patch file from the video description, fed it to Claude, and got back a JSON bank loadable into CZP-1 in five minutes. CZP-1 supports loading patches via JSON files or shareable URLs, with all data stored locally in the browser. The author notes this kind of conversion was unthinkable before LLMs.

AI Chat-Group Daily (群聊日报)

Daily Digest: Opus 5.5 generates promo videos from one prompt, Jev valuation jumps 50x in nine days

The most actionable find: Opus 5.5 with Remotion can produce a full product promo video from a single prompt, verified by multiple group members with costs shared. Roughly 50M tokens yielded a 53-second video, though pure AI output still feels flat—human editing input noticeably improved storytelling. On the industry side, Jev's valuation rocketed from $200M to $10B+ in nine days, but its moat is paper-thin with seven open-source alternatives emerging in a week. OpenAI is reportedly preparing a $500/month Pro Max tier that buys compute priority rather than quota. The group also debunked a viral post criticizing vector similarity—it attacks a meaning-identity strawman, not the topic-relevance problem RAG actually solves.

Hacker News front page

DHH at Rails World 2026: Hey is leaving Rails for Rust and native apps, built entirely by LLMs

DHH opened Rails World 2026 by declaring himself retired from professional programming and now a 'maker.' He says English is the best programming language and hand-written code is no longer economically productive. 37signals is using LLMs to rewrite Hey into six native apps with a Rust backend—Rust is hideous for humans but great for LLMs. He wrote 150k lines of code in August; Ruby dropped to 3% of his output. Rails is reframed as a framework for 'web apps of necessity,' with convention-over-configuration rebranded as token efficiency. The author questions how products differentiated by UI/UX survive if everything becomes CLI-driven by agents. DHH offered Rails devs pep-talk confidence but no actual roadmap.

Why it matters: DHH's Rails World 2026 keynote barely touched Rails itself, instead delivering provocative claims backed by concrete numbers and product decisions. The post is a second-hand reaction rather than the full keynote transcript, and actual Rails roadmap details are thin—hence not p...

Computing Life · Share · Yage

Four report cards in one week: Step 5 Preview benchmarks, Jev's confidence blind spot, and what California's AI order actually requires

StepFun's Step 5 Preview posted four types of scores, but its self-reported Terminal-Bench result isn't on the public leaderboard yet, and weight license terms remain undisclosed. An independent dev tested TypeSafe's Jev classifier on 111 public hard questions: Jev's average confidence was 0.690 when wrong and 0.684 when right—confidence doesn't separate correct from incorrect. California's AI executive order only directs state agencies to draft legislative proposals; the only active law binding companies is 2025's SB 53.

Simon Willison

Note on 24th September 2026

Simon Willison 表示,与编码智能体协作越久,越确信它们让软件工程变得更难。借助智能体可以完成惊人的工作,但释放其全部潜力需要极高的纪律性和知识储备。

Hacker News front page

Vibe Coding Production Kit: a production workflow for AI coding agents

This GitHub repo offers a full workflow from idea to production for teams using AI coding agents like Copilot. It covers specs, architecture, testing, security, code review, and CI/CD, claiming to be battle-tested. The post doesn't include benchmarks or user stories, so you'll have to try it yourself.

Hacker News front page

LaunchVideo turns a URL or prompt into an explainer video with Opus 5.5 and a headless renderer

LaunchVideo generates a ~30-second product explainer from a URL or a text prompt. Opus 5.5 writes the HTML/CSS/animation script, and a serverless agent renders it frame by frame in a headless Chromium microVM — no video generation model is used. Each video costs roughly 100k tokens and takes about four minutes, outputting 1080p 30fps MP4 with a virtual clock for deterministic frames. The page shows five unedited examples including NVIDIA and Linear. The whole product is one TypeScript agent file plus three tools, fully open-source and one-click deployable to your own OpenComputer account. The post doesn't mention pricing or whether models other than Opus 5.5 are supported.

Why it matters: A clever packaging of Opus 5.5's coding ability into a 'URL-to-launch-video' tool, with a clearly explained pipeline and visible examples. But the product is still lightweight—more a sharp demo than an industry-shaking release. H and K both hit, R is weak, landing right at the...

Simon Willison

commit-rewriter 0.2

Simon Willison 发布 commit-rewriter 0.2。该工具与 git 相关,具体功能与更新细节原文未作说明。

Hacker News front page

AI labs need to start funding historical research

The author tested GPT-6 Sol and Opus 5.5 on two historical problems: decrypting a 1941 Enigma message—where the model independently located supplementary records from the German Federal Archives—and tracing a Latin alchemical passage by Isaac Newton back to a previously unidentified French source. He argues frontier models can now deliver verifiable results on codebreaking, cross-language text tracing, and linking findings across niche subfields, a leap from last year's assistant-level performance. The post does not specify a collaboration framework or funding figures, but points to digitized, falsifiable historical problems as the sweet spot.

Why it matters: The author demonstrates frontier models' real capability in codebreaking and cross-lingual text tracing with two verifiable cases. But the topic is academic history, which limits resonance with AI industry readers, so the score sits right at the featured threshold.

AI HOT (Curated Pool)

GitHub Security Lab launches LLM-powered fuzzing agent for C/C++ projects

GitHub Security Lab released Taskflow Agent, an LLM-driven tool that automates the full fuzzing pipeline for C/C++ projects. It writes test cases, compiles, runs the fuzzer, and generates reports on crashes. The post doesn't specify which LLM is used or how many real-world bugs it has found.

AI HOT (Curated Pool)

Anthropic launches Claude Opus 5.5, optimized for cost in long-context coding sessions

Anthropic released Claude Opus 5.5, explicitly targeting cost reduction for coding sessions that run long and use heavy context. The post body only contains the title and site navigation; it does not disclose pricing, benchmarks, or context-window specs. The one confirmed takeaway is the cost-optimization angle for extended coding workflows—everything else is still missing from the article.

Why it matters: Anthropic model launch is a signal, but the body is just a title and nav bar — all key facts are missing. H and R hit, K doesn't. Barely clears the featured threshold (≥2 of 3), but thin content caps the score at 72, the featured floor.

Sep 24Thursday

Latent Space

AI made thinking cheap in science, but doing is still expensive

Adrian Sanborn splits AI biotech into Foundries and Navigators. Foundries like Xaira and Insitro industrialize experiments to lower the cost of doing science. Navigators spend the surplus of cheap thinking on faster analysis, dashboards, and decision-making without needing proprietary models. The post argues Navigator gains are invisible but available to every company, and early-stage startups adopt them fastest. At Endura Therapeutics, adapting analysis code to a protocol change dropped from a week to an afternoon, letting science 'move fast and break things.'

Why it matters: Original framework with concrete examples, but it's an opinion piece rather than hard news, landing at the lower end of featured per policy.

TechCrunch · AI

Lovable's annualized revenue hits $600M as vibe coding goes enterprise

Lovable co-founder Fabian Hedin announced at HumanX that annualized revenue has passed $600M, up from $500M three months ago. Growth is driven by enterprise adoption: people at two-thirds of Fortune 500 companies now use it, with Microsoft, Nvidia, and Deutsche Telekom named as customers. Apps built on the platform collectively draw nearly 1 billion monthly views. The post doesn't disclose profit or valuation. I'd discount the annualized figure a bit—it's last month's revenue times 12, not actual booked revenue.

Why it matters: Lovable crossing $600M ARR with named enterprise logos is a concrete signal in the vibe coding space. But the $600M is a monthly run-rate extrapolation, not audited annual revenue, so it doesn't hit 85+.

AI Chat-Group Daily (群聊日报)

Opus 5.5 effort blind test: high mode costs 30% more tokens but catches real bugs tests miss

A double-blind test on real PRs shows Opus 5.5 high mode costs ~30% more tokens and 1.33× time vs medium, but wins 16 vs 7 in blind review by catching real bugs tests missed. Claude Code Cloud Sessions goes GA with $100 Pro / $250 Max trial credits. HLE-Diamond benchmark updated: GPT-6 Astra leads at 60.6%, Gemini 3.8 Flash surprises at 34.3% beating GPT-6 Sol. Muse phone calls were partly handled by human contractors; Meta rolled back the test. The newsletter's generation tool is now open source.

Hacker News front page

Stanford and NVIDIA introduce Contrastive Language Models, up to 9× faster than Jev for decision-making

CLM encodes states and actions separately and scores pairs via cosine similarity instead of generating tokens. CLM-8B matches Jev on computer-use, gaming, and tool-calling while cutting latency by up to 9×. With light fine-tuning it hits 81.6% on DeepSWE and 87.6% on Terminal Bench 2.1, running 4–6× faster than Jev. Only the 20M-parameter projection head is trained; the frozen LLM backbone keeps pre-training to about one hour on a single RTX 4090. The post does not disclose whether weights are open or if sizes beyond 8B are planned.

Why it matters: CLM proposes a decision-making architecture orthogonal to autoregressive generation, cutting latency 9× while matching Jev on agent benchmarks — a rare paradigm-level exploration. The Notion-page format and academic author lineup mean the path to production is still unclear, c...

AI HOT (Curated Pool)

Claude Opus 5.5 tops Code Arena WebDev with 1818 points

Anthropic's Claude Opus 5.5 (Max) scored 1818 on Arena's Code Arena WebDev leaderboard, taking first place. It leads GPT-6 Astra (Max) by 26 points and beats Opus 5 (Max)'s 1692 by 126 points. The post doesn't include evaluation details beyond the scores and rankings.

Why it matters: Claude Opus 5.5 tops Code Arena WebDev with concrete scores and gaps — directly useful for Claude-heavy devs. But the post doesn't disclose methodology, task scope, or evaluation conditions, so the information density only clears the featured threshold, not p1.

AI HOT (Curated Pool)

Claude Opus 5.5 tops Coding Agent Index, but per-task cost rises to $13.04

Artificial Analysis tested Claude Opus 5.5 under Claude Code max effort and it scored 66 on the Coding Agent Index, up from Opus 5's 60. All three subtests improved: Terminal-Bench 4.0 63.1%, DeepSWE v1.1 68.4%, SWE-Atlas-QnA 66.4%. The trade-off: per-task cost jumped from $3 to $13.04. The post doesn't break down how max effort drove the cost increase.

Why it matters: Claude Opus 5.5 tops the Coding Agent Index with a 6-point jump to 66, but $13.04 per task is the hard number. Anthropic substantive update + independent third-party benchmark + concrete data — all three HKR axes hit. Not scoring higher because this is a single benchmark, not ...

Hacker News front page

1Password's FLAWED paper on AI patching criticized for thin citations and factual errors

Suha Sabi Hussain publicly criticized 1Password's FLAWED paper from Off-by-1 Labs. The paper claims frontier models often produce flawed vulnerability patches, but Hussain notes it cites only 19 sources—mostly corporate blogs and XKCD—while omitting directly relevant prior work like Meta's AutoPatchBench and an NDSS paper. The paper also contains mislabeled diagrams and arithmetic errors. Hussain argues that 1Password adopted the tone of rigorous research without the corresponding rigor, and that this work overshadowed higher-quality research from less-resourced groups like EleutherAI. She calls for a retraction or correction and suggests partnering with academic researchers.

Why it matters: The author, a security researcher, provides concrete evidence (missing citations to Meta's AutoPatchBench and an NDSS paper) against 1Password's FLAWED paper — not empty criticism. But it's a personal blog rebuttal, not primary research or a product launch, so importance sits ...

Computing Life · Share · Yage

Anthropic used Claude to optimize 36 biomolecular modeling packages, achieving up to 4.1× speedup in four weeks

Two Anthropic researchers with biomodeling expertise but no GPU kernel background spent under four weeks with Claude refactoring 36 open-source biomolecular packages. They built FlashPairformer, a custom GPU kernel that fuses scattered triangle-attention ops into high-throughput streaming, then applied per-model caching and CUDA graph replay. Benchmarked on H100 against a hand-tuned expert baseline, the bitwise-identical exact mode averages 1.6× speedup; the fast mode, which allows noise within the model's own stochastic range, averages 4.1×; the memory-saving big mode averages 3.4×. Exact and fast modes can push memory up to 3×. DockQ acceptable rates stayed at 54–55% across modes, with no systematic accuracy loss. The report draws clear lines: big mode ran a 10,761-token complex at TM-score 0.92–0.997, but on 31k–70k-residue viral capsids the outputs collapsed into dense balls (TM-score 0.08–0.14). The authors attribute this to the model's 768-token training-crop limit, not the optimizations. In protein design, a single Claude instance driving optimized models on one H200 for 24 hours hit a median ipSAE of 0.785, up from 0.749 in the earlier multi-agent campaign, but none of the designs have been wet-lab tested. Code is open-sourced under Apache-2.0 with no ongoing maintenance.

Why it matters: Anthropic researchers used Claude to refactor 30+ biomolecular model codebases in under four weeks, shipping FlashPairformer kernels and reproducible optimizations. Concrete technical details, open-source code, measured results — not a fluff piece. Points off: this is a yage.a...

AI HOT (Curated Pool)

Fireworks launches Ember-1, matching Kimi K3 quality with 40% fewer tokens

Fireworks Research released Ember-1, a model built on Kimi K3 that cuts reasoning tokens by 35–50% while keeping accuracy. Across 7 benchmarks and live A/B tests with two customers, quality held. The team ran 50+ training experiments and found K3 spends over 90% of tokens on internal reasoning, much of it unnecessary. Ember-1 preserves useful self-correction and skips unproductive loops. Savings compound in multi-turn agent tasks where prior reasoning is re-read each turn. The model is live on Fireworks' platform as the first in their own model series.

Why it matters: Fireworks distilled Kimi K3 into Ember-1, cutting reasoning tokens by 35-50% with no accuracy drop, backed by 50+ training runs and live customer A/B tests. Score stays below 85 because this is an optimization of an existing model rather than a new capability release, and Fire...

Hacker News front page

Anthropic made claude.ai 3x faster in two weeks, with Claude itself finding bottlenecks, shipping fixes, and watching deploys

Anthropic ran a two-week sprint in August that made four core journeys on claude.ai and the desktop app about 3x faster. Cold-load time to a typeable page dropped from 3.1s to 0.55s, starting a new Claude Code session from 0.8s to 0.3s, and loading a Claude Cowork cloud session from 2.6s to 0.73s. The team ran everything from a single Slack channel where Claude Tag (beta, running a research model close to Opus 5.5) analyzed Datadog data, built benchmarks, proposed and shipped improvements, and watched every deploy — humans set goals, made tradeoffs, and approved changes. Over 3,000 changes were merged with zero customer-facing incidents or rollbacks. Optimizations included baking a static composer into HTML, precompiling a V8 code cache, keeping the composer mounted across conversations, prefetching sessions on hover, and cutting sidebar re-renders by 90%. The team also built deterministic lab benchmarks (Valgrind instruction counts, React commit counts, V8 call counts) so Claude could validate optimizations without waiting for production deploys.

Why it matters: Official Anthropic engineering blog with concrete latency numbers and the Claude Tag hill-climbing approach — useful for Claude users and engineers. But it's a performance optimization, not a new capability launch, so it lands at the 78 featured threshold rather than higher.

AI HOT (Curated Pool)

Claude team shares how they used Claude to make claude.ai 3× faster in two weeks

The Claude team made claude.ai 3× faster in two weeks and published their method. They used Claude itself to measure latency, find bottlenecks, and suggest fixes—prompts included. The post links to a blog; before/after metrics aren't in the snippet.

Why it matters: Anthropic team published a hands-on case study and prompts for using Claude to 3x their own product speed. Hits all three HKR axes. Deduction: the post doesn't give before/after latency numbers—you have to click through to the blog for the actual seconds saved—so it stays belo...

AI HOT (Curated Pool)

Antigravity SDK now supports local models for fully offline agents

Google added local model support to the Antigravity SDK, starting with Gemma 4 26B A4B via LiteRT. Agents can now run fully offline, keeping code and requests on-device. A hybrid demo uses Gemini 3.8 Flash as a cloud planner (95 tokens) while local Gemma 4 26B instances handle the audit-and-patch work—97.2% of tokens stay local. Another example shows the agent building a live CLI resource monitor from a single prompt. The post recommends >24GB VRAM or unified memory.

Why it matters: Google added local model support to the Antigravity SDK, starting with Gemma 4 26B. The hybrid mode—cloud planner at 95 tokens, local executor—comes with concrete cost numbers, not just a concept. Directly useful for devs building on-device agents. Not an 85 because it's locke...

AI HOT (Curated Pool)

GPT-6 Sol (Max) ranks 4th in WebDev arena at $8/M tokens

Arena released the real voting results for GPT-6 Sol (Max). It scored 1689 in Code Arena: WebDev, ranking 4th. Price is $8/M tokens (mixed input/output). The post doesn't spell out test setup, comparison models, or latency.

Sep 23Wednesday

AI HOT (Curated Pool)

Anthropic engineer shares 6-step prep for AI-driven code modernization

An Anthropic field engineer shares a practical guide for modernizing legacy code with AI. The key: don't start by having AI rewrite code. Instead, follow six preparation steps: map dependencies, write tests, pick a small pilot, choose the right model (e.g., Claude Code), and set up human review. The post doesn't include specific case studies or cost figures, but the steps are concrete enough for teams unsure where to start.

Latent Space

Claude Opus 5.5 launches with Fable 5.1-level performance at 40% lower cost, plus a rare focus on writing quality

Anthropic released Claude Opus 5.5, the first model in the new 5.5 family. It matches Claude Fable 5.1 on most tasks, costs 40% less to run than Opus 5, and is about 30% faster. The launch unusually highlights writing improvements: the model puts key info up front and follows user style rules. Artificial Analysis notes that token usage on frontier tasks jumped ~80%, so per-task cost remains around $6—similar to Opus 5. OpenAI shipped GPT-6 Sol and Luna an hour later at 50% lower prices than GPT-5.6, but Opus 5.5's launch post hit 17M views and dominated the day. Anthropic's system card also reports multi-agent scaling with up to 100 parallel agents for the first time. Latent Space tested both and switched to Opus 5.5 as the default model immediately, calling the writing quality a night-and-day difference over Sol 6.

Why it matters: Anthropic drops the first model in a new flagship family, claiming Fable 5.1 parity at 40% lower cost, with writing improvements front and center — a directly actionable upgrade signal for heavy Claude users. Held below 90 because we only have the official claim and Latent Spa...

AI Chat-Group Daily (群聊日报)

Anthropic Opus 5.5 and OpenAI Sol/Luna drop same day; community breaks down effort cost-efficiency and migration pitfalls

Anthropic 毫无预兆地放出 Opus 5.5,在终端操作和编程任务上跑分领先,但 max 档输出 token 量是 GPT-6 Astra 的三倍多。群友分析发现 high 档是性价比甜区:比 medium 多花 36% 的钱,智能指数涨 3 分,再往上边际成本陡增。两小时后 OpenAI 上线 Sol 和 Luna,Luna 输入价格打到每百...

Why it matters: Anthropic Opus 5.5 launched without warning, OpenAI followed with Sol and Luna two hours later — three model resets in one day. The daily digest provides real-user effort-tier cost/performance breakdowns and prompt-migration war stories, high signal density. Deduction: this is...

AI HOT (Curated Pool)

OpenAI releases GPT-6 Sol and Luna with 50% cheaper API pricing and benchmarks

OpenAI added two models to the GPT-6 family: Sol for complex coding and professional tasks, Luna for fast high-volume work. API pricing is cut by 50% vs GPT-5.6 promo rates—Luna's output price actually dropped 58%. Sol beats Claude Opus 5 on AutomationBench and Agents' Last Exam at roughly one-tenth the cost per task. Both are live in the API today; no weights are released.

Why it matters: OpenAI drops two new GPT-6 variants with a 50% API price cut — an industry-shaking move. Sol's Aura score and Luna's $0.5 output price are concrete, though the post doesn't include the full benchmark table. Still, this is a must-cover story.

Simon Willison

SF October 14th: A Birds of a Feather Session on Agentic Engineering

Simon Willison 与 Jesse Vincent 将于 10 月 14 日(周三)在旧金山举办一场面向 coding agent 构建者的晚间交流活动,主题为 Agentic Engineering。活动采用非正式的 show-and-tell 形式,鼓励参与者分享尚未公开的尝试、奇怪实验和未完成项目,无需正式演讲,也不是产品推销。

Computing Life · Share · Yage

Same tool toggle: Nemotron-3 550B gained, Mistral-Medium-3.5 crashed

A new paper breaks down coding agent harnesses into three independent toggles and measures each one. The most striking result: switching from dedicated file tools to a pure CLI made Nemotron-3 550B's SWE-Bench Verified score jump 3.6 pp while cutting per-task cost from $2.33 to $1.11, but Mistral-Medium-3.5-128B dropped from 68.60% to 45.40%. Trajectory analysis shows 550B composing dense shell one-liners, while Mistral failed to locate files in 32.80% of tasks and submitted no edits. On Terminal-Bench 2.1, both models improved under CLI mode. Planning boosted the 30B model from 13.60% to 25.20% but only saved ~30% cost for larger models without accuracy gains. Context management mainly prevents window overflow; at 128k the gap shrinks to 2.7 pp, and complex read-back mechanisms were almost never invoked. The takeaway: no universal best harness design—it depends on the model's CLI fluency and the task type.

Why it matters: A controlled experiment that isolates three harness design switches and shows Nemotron-3 and Mistral-Medium-3.5 reacting in opposite directions, with concrete numbers and engineering takeaways. Not an 85 because it's a single preprint without cross-source cluster yet, but HKR ...