Skip to content
AI Chat-Group Daily (群聊日报)

Chat Group Weekly Vol.4: Benchmark scores soar but Chinese writing gets worse; does mastering AI actually get you promoted; Fable 5 refuses to summarize chat logs

Issue 4 of Yage's AI chat group weekly covers three topics through blind tests and group complaints. First, Opus 4.7 and 4.8 score higher on benchmarks but Chinese writing quality has clearly regressed—output reads like it's not real Chinese and documentation is nearly unusable. The author argues this isn't models getting dumber but reinforcement learning creating lopsided specialists: math and coding with clear right/wrong answers get optimized aggressively, while writing ability that relies on taste gets sacrificed because it's not in the reward function. Kimi's team admitted in a Reddit AMA that maintaining writing taste across versions is a challenge requiring dedicated monitoring. Second, Fable 5's safety guardrails are overly sensitive—it refused to summarize chat logs for three straight days and burned $5 because the discussion mentioned an article about hackers using sensitive keywords to evade LLM analysis. Anthropic was also caught secretly degrading Claude's performance when used to train competing models, which critics called "secret sabotage"; they later apologized. Third, a group member shared an article asking: does 10x productivity with AI actually lead to promotion? The answer is no. One professor decided to stop recruiting students after First Proof benchmarks showed $1,000 worth of AI could match a PhD student's five-year output.

Why it matters: This community newsletter uses first-hand blind-test data to flag Opus 4.7/4.8's Chinese writing regression, with concrete evidence rather than empty opinion. But it's a personal blog observation, not an official announcement or reproducible study, so authority is limited—henc...

Read the original ↗Export Markdown