This caught my eye because the numbers are unusually concrete. One user ran DSH + DeepSeek V4 Flash Vision Exp for a full week of overnight sessions: 294 turns, 11,318 steps, 5.1B input tokens, nearly 10M output tokens. Average first-token latency was 2.3 seconds, cache hit rate 99.8%, total cost ¥362.84. That's roughly $50 for a week of heavy agent work.
The comparison that stuck: Sol insisted on Rust + Tauri and burned hours just setting up the toolchain. V4 Flash checked the local environment, found Go, and recommended Go + Wails—straight into iteration. Multiple users described Sol as an overthinker that produces bloated, hard-to-trust output. One called it "a seasoned office politician."
I'd discount this a bit. These are chat-group vibes, not controlled evals. One person noted V4 Flash still lags on coding benchmarks, another said Sol is steadier for information synthesis. The real variable might be the harness—DSH vs Codex—not the model itself. Comparing them head-to-head is messy.
The "subscription gym paradox" theory is fun: subscription-based harnesses have incentives to quietly throttle heavy users, while pay-per-token models don't care how much you burn. Plausible, but the post provides zero evidence of actual throttling. Treat it as a hypothesis.
Same day, Anthropic launched Fable 5.1 with 75% cheaper cache reads, but Fable 5 scored below Opus 5—users called it underwhelming. Qwen 3.8-Max-0902 also dropped, with coding and agent benchmarks close to Opus 5. Astra hit Critical cybersecurity tier, scored perfect on ExploitBench, and found two zero-days. That one deserves its own follow-up.