Skip to content
AI Chat-Group Daily (群聊日报)

GLM-5.3-Flash and Qwen 3.8-Flash-Next debut on the same day, both drop global attention

2026-08-26 群聊日报

GLM-5.3-Flash matches Claude Opus 4.8 across six benchmarks at $0.045 per task, but testers report slow speed and hallucinations. Qwen 3.8-Flash-Next opens weights, hitting 64.7 tok/s single-stream decode on DGX Spark and beating DeepSeek V4 Flash across the board. Both models adopt MoE plus sparse attention hybrids, ditching global attention. NVIDIA acquires Hugging Face for $12.9B, roughly 86x its annualized revenue, to control the open model distribution channel. Anthropic preps IPO at a ~$2T valuation target, with ~$559M adjusted operating profit in Q2, while OpenAI posted ~$12.3B operating loss in the same period. Altman admits on a podcast that OpenAI hasn't had its iPhone moment and has scrapped Sora and Atlas. RTX 30 series GPUs resume production using Samsung 8nm to avoid TSMC bottlenecks. Shopify's CEO complains Claude Code ignores AGENTS.md, causing split brain in teams. QUASAR-QAT quantizes all 496 linear layers of Qwen 3.8-27B to NVFP4, saving another 1.8GB VRAM. The group also discusses Sol's context bloat and the limits of fully automated PR merges.

Why it matters: Two domestic Flash models launched the same day — GLM-5.3-Flash posts strong benchmarks but slow real-world speed and hallucinations, while Qwen 3.8-Flash-Next is open-weight with measured inference speed beating DeepSeek V4 Flash. Concrete numbers, real-user feedback, archite...

Read the original ↗Export Markdown