Qwen has released a 35B-total, 3B-active Apache 2.0 MoE, but right now this reads more like a distribution move than a capability conclusion. The Reddit post lists agentic coding, multimodal reasoning, and thinking/non-thinking modes, yet it omits the three numbers that decide whether practitioners should care: benchmarks, context length, and latency. Without those, “on par with models 10x its active size” is marketing copy, not evidence.
My take is that Qwen is still pushing for default-open-model status, not trying to prove it has surpassed the closed frontier. The 35B/3B recipe already tells you the target: practical deployment envelopes, not giant clusters. This is aimed at teams that want something deployable, commercially usable, and flexible enough for code plus multimodal workloads. The Apache 2.0 license matters more than the Reddit post makes explicit. Over the last year, plenty of “open” models still created legal friction for commercial teams. Qwen at least removes that excuse.
There’s also a clear market pattern here. The open ecosystem has already learned from the DeepSeek-style MoE wave that total parameters are a weak proxy for real usability. Active parameters, KV-cache growth, context length, quantized memory footprint, and tokens per second matter more once a model leaves the benchmark slide and lands on an actual GPU box. If Qwen wants the “agentic coding” claim to stick, it should publish reproducible results on SWE-bench, Aider, LiveCodeBench, or at least a clearly scoped internal eval. If the multimodal claim is serious, I’d expect MMMU, MathVista, maybe VideoMME, plus hardware conditions. None of that is in the post.
I also have some doubts about the product packaging. “Multimodal perception and reasoning” sounds clean in a launch graphic, but in deployment those capabilities often pull in different directions. A model can be strong in text reasoning and still become awkward once vision is attached: latency rises, memory use changes, and tool-use stability gets worse. The thinking/non-thinking split raises the same question. That design is popular because it lets one model chase both responsiveness and long-chain reasoning scores, but the key issue is whether this is genuine controllable inference on one base model or just prompt-layer mode management. The post doesn’t say.
So I’m interested, but not persuaded yet. Qwen needs to publish three things before this becomes a serious practitioner story: public benchmark tables, real inference speed on named hardware, and context/quantization guidance. If those numbers land well, this can become one of the stronger local-first open releases of the year. If they don’t, this was mainly a well-timed attention grab.