Simon ran a 20.9GB quantized Qwen3.6-35B-A3B on a MacBook Pro M5, and it beat Claude Opus 4.7 in two SVG drawing prompts. The interesting part is not “Qwen is smarter than Opus.” The interesting part is that this test keeps exposing a split people still blur together: structured code-like rendering is not the same thing as general model quality.
I mostly agree with Simon’s own read. Qwen won here because SVG generation rewards a narrow cluster of skills: local structure control, shape consistency, willingness to commit to decorative details, and clean token-by-token execution inside a markup format. A bicycle frame is a constraint satisfaction problem expressed as code. Claude Opus 4.7 messing that up tells you something real about its behavior on this task. It does not tell you that a laptop Qwen is broadly better than Anthropic’s flagship across reasoning, tool use, long-horizon agents, or enterprise reliability.
That context matters because local models have been closing fastest on exactly these “visible artifact” tasks over the last year: SVG, small front-end apps, component rewrites, JSON-heavy transforms, short code edits, constrained generation. Qwen and DeepSeek have been especially strong there. Closed models still tend to win on messy multi-step work and production stability, but if the task can be compressed into a short context with an objectively inspectable output, mid-sized open or open-weight models are now regularly good enough. Sometimes they’re more pleasing because they commit harder to style.
I do have some pushback on how far to take this result. First, the article shows examples and links transcripts, but it does not summarize the full experimental controls in the post itself: sampling settings, temperature, whether seeds were fixed, whether any retries were discarded, whether the “best of N” effect crept in. For a blog post, that’s fine. For a capability claim, it’s thin. Human-scored one-off generations are great for intuition and terrible for ranking.
Second, Opus 4.7 is carrying the burden of “flagship model” expectations. People see a premium proprietary model and assume it should dominate every small task. I don’t buy that assumption. Frontier closed models are increasingly optimized as compromise machines: safety tuning, refusal behavior, long-context coherence, tool calling, enterprise controls, latency tiers, and cost discipline all pull in different directions. A tightly constrained SVG illustration prompt may simply sit lower in Anthropic’s optimization stack than users assume.
The backup test matters, though. Simon swapped “pelican riding a bicycle” for “flamingo riding a unicycle,” and Qwen still won in his judgment. That weakens the joke theory that labs are secretly training on his exact pelican benchmark. It does not eliminate a broader explanation: Qwen may just be very well tuned for SVG-ish and front-end-ish generations. I haven’t verified their latest training mix, and the article does not disclose it, so I wouldn’t invent a story about deliberate benchmark targeting. But I would absolutely read this as evidence of model-family bias toward structured visual code outputs.
There’s also a deployment angle people should care about more than the pelican itself. A 20.9GB GGUF running through LM Studio on a laptop is not a lab-only setup. That is close enough to commodity for a lot of individual developers and small teams. If your product needs rough icons, diagrams, mascots, landing-page sketches, or editable vector placeholders, the workflow implication is pretty direct: generate drafts locally, keep expensive API calls for planning, review, orchestration, and high-stakes tasks. That is a more useful conclusion than debating who “won” the meme benchmark.
So my read is simple. Qwen did not prove that a 35B local model has overtaken Opus in general intelligence. It did prove, again, that a growing slice of creative output generation has become cheap, local, and inspectable. If Anthropic keeps missing on deterministic short-form artifact tasks like this, users will start asking an uncomfortable pricing question: what exact capability premium are they still paying for?