Opus 5.5 effort blind test: high mode costs 30% more tokens but catches real bugs tests miss
2026-09-23 群聊日报
A double-blind test on real PRs shows Opus 5.5 high mode costs ~30% more tokens and 1.33× time vs medium, but wins 16 vs 7 in blind review by catching real bugs tests missed. Claude Code Cloud Sessions goes GA with $100 Pro / $250 Max trial credits. HLE-Diamond benchmark updated: GPT-6 Astra leads at 60.6%, Gemini 3.8 Flash surprises at 34.3% beating GPT-6 Sol. Muse phone calls were partly handled by human contractors; Meta rolled back the test. The newsletter's generation tool is now open source.