Skip to content

Alibaba's Qwen family: open releases and iterations, from flagship models to small on-device ones.

Latest picks

161–180 of 205

May 5Tuesday

r/LocalLLaMA

Heretic 1.3 Released: Reproducible Models, Integrated Benchmarks, Lower Peak VRAM

Heretic 1.3 adds reproducible runs, integrated benchmarks, lower peak VRAM, and broader model support. The project claims 20,000 GitHub stars and 13 million model downloads. Reproduce directories capture PyTorch, GPU, driver, and accelerator details; benchmarks use lm-evaluation-harness for MMLU, EQ-Bench, GSM8K, and HellaSwag. The post names Qwen3.5 and Gemma 4 support, but does not disclose VRAM reduction figures.

Why it matters: HKR-K/R pass: 20k stars, 13M downloads, reproducibility metadata, and eval harness are concrete. HKR-H fails and VRAM reduction lacks numbers, so this sits at the featured threshold.

r/LocalLLaMA

Prompt injection benchmark: delimiter and strict prompt took Gemma 4 from 21% to 100% defense rate

A Reddit user posted a prompt-injection benchmark covering 15 models, 7 attack types, and 6,100+ cases. The setup wraps untrusted documents in long random delimiters; Gemma 4 E4B rose from 21.6% to 100% defense. The key detail is the reproducible metric: blocked/(blocked+failed).

Why it matters: HKR-H/K/R all pass: Gemma 4’s defense-rate jump is clickable, the test setup is concrete, and prompt injection matters to builders. Single Reddit benchmark keeps it in the 78–84 band.

r/LocalLLaMA

MTPLX: 2.24x Faster TPS Native MTP Inference Engine for Apple Silicon

MTPLX raises Qwen3.6-27B on a MacBook Pro M5 Max from 28 to 63 tok/s. The test used 4-bit MLX, temperature 0.6, top_p 0.95, top_k 20, with D3 as the best depth. The key detail is native MTP heads: no external drafter and no second-model memory.

Why it matters: HKR-H/K/R all pass: a 2.24x speed hook, concrete test conditions, and a local-inference cost nerve. Reddit single-post sourcing and narrow Apple Silicon scope keep it in low featured, not P1.

r/LocalLLaMA

Benching Local Qwen as a Codex Validator, Co-agent, and Challenger

robert896r1 tested Qwen3.6 27B GGUF beside Codex as a coding validator and released a reproducible eval suite. The runs covered Bartowski, Unsloth, 65k/128k context, and q8/f16 KV cache; three 128k profiles tied for best, with no measured q8 KV accuracy loss in this suite. The useful signal is the sidecar eval: missed directives, overbuilding, UI judgment, and long-context misses, not a universal leaderboard.

Why it matters: HKR-H/K/R all pass: a reproducible sidecar eval with concrete Qwen/Codex conditions beats a normal Reddit tip. Source authority and event scale keep it in the 72–77 band, not a same-day must-write.

May 4Monday

r/LocalLLaMA

Deep research report with Hermes Agent and qwen3.6-35b-a3b Q6_K

A Reddit user used Hermes Agent and qwen3.6-35b-a3b Q6_K to produce a 21-page research report. The run took 6 loops and over 5 hours on an RTX 4060, at about 28 tokens/s. The repo includes prompts, scripts, intermediate artifacts, and the final report.

Why it matters: HKR-H/K/R all pass: this is a local-agent experiment with hardware, runtime, speed, and artifacts. Reddit source limits reach, so it stays in the 72–77 featured-threshold band.

r/LocalLLaMA

Mistral Medium 3.5 128B and Qwen 3.5 122B A10B on 4x RTX 3080 20GB

A Reddit user benchmarked Mistral Medium 3.5 128B and Qwen 3.5 122B A10B on 4x RTX 3080 20GB. llama.cpp tensor split raised Mistral tg128 from 10.37 to 21.59 t/s, but Qwen MoE fell from 60.08 to 53.49 t/s. vLLM served Qwen GPTQ-Int4 at 187.04 tok/s; the key signal is MoE sensitivity to parallel strategy.

Why it matters: HKR-H/K/R all pass: the 4×RTX 3080 setup is a strong hook, and the post gives concrete llama.cpp/vLLM throughput deltas. Reddit single-run sourcing keeps it in the 72–77 band.

r/LocalLLaMA

Pushing a 5-Year-Old 6GB VRAM Laptop to Its Limits: Qwen3.6-35B-A3B

Reddit user abhinand05 ran Qwen3.6-35B-A3B on a 5-year-old Asus ROG Zephyrus G14, reaching about 23 t/s plugged in and 10+ t/s unplugged. The setup uses RTX 2060 Max-Q 6GB, 24GB DDR4, Ryzen 7, plus llama-server configs for 64k and 128k context. The key detail is the mix of CPU MoE, KV-cache quantization, and ngram speculative decoding.

Why it matters: HKR-H/K/R all pass: the old-laptop angle is clicky, the post gives speeds and configs, and local-LLM cost resonates. It remains a single Reddit run, not a broader release.

May 3Sunday

r/LocalLLaMA

Local LLM Benchmark for Backend Generation via Function Calling: GLM vs Qwen vs DeepSeek

AutoBe posted a controlled backend-generation benchmark and says qwen3.5-35b-a3b matches gpt-5.4 on DB/API design. One shopping-mall run uses 200–300M tokens, costing $1,000–$1,500 per model at GPT 5.5 pricing. The key caveat is n=4 projects and self-scoring harness bias.

Why it matters: HKR-H/K/R all pass, but Reddit sourcing, n=4 projects, and self-eval harness bias keep it at the low featured band. Concrete cost and test constraints carry the score.

r/LocalLLaMA

Paper on Hummingbird+: low-cost FPGAs for LLM inference

A Hummingbird+ paper claims low-cost FPGAs run Qwen3-30B-A3B Q4 at 18 t/s generation. The title lists 24GB memory and an expected $150 mass-production cost; the post does not disclose FPGA model, power, or test conditions.

Why it matters: HKR-H/K/R all pass: the hook is a $150 FPGA running a 30B Q4 model, with speed, memory, and cost stated. Power, FPGA SKU, and test conditions are missing, so this lands at 79, not P1.

r/LocalLLaMA

Local image generation on Mac: 10 models compared

A Reddit user tested 10 image models on an M1 Max with 64GB RAM. Qwen-Image Lightning’s 8-step distillation beat the full model at 10 minutes versus 93. Flux dev led local photorealism but showed English-centric bias; Gemini handled kanji and context better but is cloud-only.

Why it matters: Named first-person test with concrete numbers: HKR-H from a 10-model Mac comparison, HKR-K from timing and quality deltas, HKR-R from local-vs-cloud tradeoffs. Single Reddit sample keeps it below must-write.

May 2Saturday

r/LocalLLaMA

Qwen 3.6 wins benchmarks, but Gemma 4 looks stronger in local vision tests

A Reddit user compared Qwen 3.6 and Gemma 4 locally on vLLM FP8 across 27B/31B vision models. Qwen burned 8,000+ tokens on hard GeoGuessr cases, while Gemma often used 1,500; Qwen also needed 2 FPS video preprocessing. The practitioner detail: vLLM and Llama.cpp can default Gemma visual tokens to 280, while 1,120+ improved fine-detail accuracy.

Why it matters: HKR-H/K/R all pass: the post has a sharp benchmark-vs-reality hook and concrete local vLLM/FP8 settings. A single Reddit test limits authority, so it sits just above the featured threshold.

r/LocalLLaMA

Qwen3.6-27B hits 72 tok/s on RTX 3090 with native vLLM on Windows

Reddit user One_Slip1455 released a native Windows vLLM launcher for Qwen3.6-27B, reaching 72 tok/s on an RTX 3090. It reports 64.5 tok/s at ~25k tokens, 53.4 tok/s at 127k ctx on one GPU, and 160k ctx with PP=2 on 2×3090. The key detail is no WSL or Docker, an OpenAI-compatible endpoint, and an INT4 quant path.

Why it matters: HKR-H/K/R all pass: native Windows on an RTX 3090 is the hook, the post gives tok/s and ctx figures, and it hits local-inference cost concerns. Reddit single-source limits it to the lower featured band.

May 1Friday

r/LocalLLaMA

PFlash: 10x prefill speedup over llama.cpp at 128K on an RTX 3090

PFlash cuts Qwen3.6-27B Q4_K_M 128K TTFT to 24.8s on an RTX 3090, versus 248.4s cold for llama.cpp. It uses a Qwen3-0.6B drafter to score token importance, keeps 5% of spans, and runs C++/CUDA without Python, Triton, or PyTorch. The quality caveat is clear: only NIAH single-needle passes from 32K to 128K; RULER and multi-needle results are not disclosed.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit claim with quality evidence limited to single-needle NIAH 32K–128K. RULER and multi-needle results are not disclosed, so it stays at featured threshold.

r/LocalLLaMA

Follow-up: Qwen3.6-27B on 1× RTX 3090 reaches ~218K context and stable tool calls

A Reddit user ran Qwen3.6-27B on one RTX 3090, reporting ~218K context at 50/66 TPS. After fixing Genesis PN12 patch anchor drift, ~25K-token tool outputs stopped OOMing; 198K plus vision reached 51/68 TPS. Single-prompt single-GPU runs still hit a second memory cliff near 50–60K.

Why it matters: HKR-H/K/R all pass: the single-3090 context claim is catchy, the post gives measured TPS and OOM conditions, and local-inference cost pressure resonates. Reddit source keeps it in the low featured band.

r/LocalLLaMA

Long-context coding on RTX 5080 16GB: Qwen3.6-35B-A3B holds 30 t/s at 128K

A Reddit user tested a local coding-agent setup on RTX 5080 16GB; the title says Qwen3.6-35B-A3B reaches 30 t/s at 128K. The post lists Ryzen 9700X, 96GB DDR5, Windows 11, and CUDA 12.9.1 as required. Qwen3.6-27B dense hit only 3.2 t/s at 128K, so the key path is KV quantization plus MoE offload.

Why it matters: HKR-H/K/R all pass: 30 t/s at 128K on a 16GB RTX 5080 is a strong hook, with hardware/CUDA details and a dense baseline. Single Reddit run lacks multi-source reproduction, so featured not P1.

Apr 30Thursday

MIT Technology Review · AI

Goodfire releases Silico, a mechanistic interpretability tool for debugging LLMs

Goodfire released Silico, letting engineers inspect and adjust LLM parameters during training. It maps neurons and pathways; one Qwen 3 neuron triggered trolley-problem-style outputs. Pricing is case-by-case, and the post does not disclose rates.

Why it matters: HKR-H/K/R all pass: Silico offers a concrete interpretability-debugging mechanism. It stays at 76 because this is a startup product preview with no pricing or adoption scale disclosed.

r/LocalLLaMA

Actual comparison between locally run Qwen-3.6-27B and proprietary models

The author compared 5 model setups on an autoresearch-loop task; only Qwen-3.6-27B via OpenRouter nearly solved it. The local q4_k_m run took about 8 hours and used 39k/45k tokens; full-quality Qwen used 4.4M tokens and cost $0.939. The useful signal is failure quality: both Qwen runs needed small fixes, while Gemma, Codex-Spark, and Claude Haiku 4.5 missed tests or key logic.

Why it matters: HKR-H/K/R all pass: the post has a concrete agent-test surprise, token and cost data, and local-vs-proprietary tension. Single Reddit run limits source authority, so it stays in the lower featured band.

r/LocalLLaMA

Notes on what actually breaks when you run a coding agent on small local models

A Reddit user tested small local and free-tier cloud models for weeks on multi-file coding tasks. Sub-7B structured output was unreliable; failures included markdown fences, wrong-file edits, and read/write misclassification, with post-processing and validation as fixes.

Why it matters: HKR-H/K/R pass: the post names real local coding-agent failure points, a sub-7B threshold, four failure classes, and mitigations. Reddit single-post scope keeps it below release-tier news, so 75.

r/LocalLLaMA

Qwen-Scope: Official Sparse Autoencoders (SAEs) for Qwen 3.5 models

Qwen Team released Qwen-Scope, SAEs for Qwen 3.5 models from 2B to 35B MoE. It maps residual-stream features across all layers, including Feature #6159 for Chinese activation. The key point is feature-level debugging and steering; the license discourages removing safety filters.

Why it matters: HKR-H/K/R all pass: official Qwen SAEs are novel, concrete, and useful for interpretability work. This is not a new model release, so it stays in the 78–84 recommendation band.

QbitAI · WeChat

NUS and collaborators propose ViF to curb visual hallucination snowballing in multi-agent systems

NUS LV-Lab and collaborators proposed ViF, accepted to ICLR 2026. Across 8 benchmarks, 4 MAS structures, and 10 VLMs, it reports 2.4%–3.8% average gains. ViF replaces text-only passing with visual relay tokens and layered attention redistribution, cutting HS by over 30% on average and nearly 40% in ring topology.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the mechanism and eval grid are disclosed, and hallucination control matters to agent builders. Scope stays research-heavy, so it sits at the featured threshold, not same-day must-write.