Skip to content

Alibaba's Qwen family: open releases and iterations, from flagship models to small on-device ones.

Latest picks

101–120 of 205

May 24Sunday

Synced · WeChat

ICML 2026: First Parallel Thinking Framework for Vision-Language Models

Visual Para-Thinker introduces a parallel thinking framework for vision-language models, using Pa-Attention and LPRoPE to isolate four visual reasoning paths and training on 163,000 question-answer pairs.

Why it matters: HKR-H/K/R pass: the ICML 2026 paper offers a concrete parallel-thinking mechanism, four isolated paths, and 163K training pairs. It remains a single research release without broad replication or product impact, so it fits 78–84.

r/LocalLLaMA

It's OK to Quantize the KV Cache; Model Quant Matters More in Qwen3.6 27B KLD Tests

Reddit user hopbel tested Qwen3.6 27B with approximate KLD on wikitext-2 at 16k context, using Q5_K_M as the proxy baseline; Q5_K_S weights with q4_0 KV cache scored 0.016304, while Q4_K_XL with f16 KV cache scored 0.026067, so weight quant tier dominated KV-cache quant in this setup.

Why it matters: HKR-H/K/R all pass, backed by first-person test numbers. Source is a single Reddit post, the metric is approximated KLD, and the claim is narrow, so it sits at the featured threshold.

May 23Saturday

r/LocalLLaMA

club-rdna16: Practical 16GB AMD/Radeon local LLM testing repo

club-rdna16 publishes a practical 16GB Radeon local LLM testing repo, with an RX 6900 XT running llama.cpp on ROCm/HIP and Qwen3.6 35B-A3B reaching a stable 131k context using q8 KV cache.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the body only discloses test conditions, not speed, VRAM curves, or reproducible logs. It clears the featured floor as a practical local-LLM repo.

r/LocalLLaMA

How small can the orchestration model in an agent be? Separating it from code generation

HomoAgens1 runs a local ReAct orchestration loop on Qwen3.6-35B-A3B, with about 3B active parameters, a 12GB GPU, 30 expert offload, and 40 tokens/s prompt generation; smaller dense models fail first on tool-call discipline, inventing arguments or repeating bad calls, while reasoning is not identified as the first break point.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment rather than a formal release. The VRAM, speed, and failure-mode details put it at the 72 featured threshold.

r/LocalLLaMA

Experts first llama.cpp

comanderxv published a llama.cpp fork that caches MoE experts in 12GB VRAM; on an RTX 2060 with Qwen3.6-35B-A3B, throughput rose from 19/22 tk/s to 26 tk/s at about a 62% expert-cache hit rate.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE speedup on a 12GB RTX 2060, with concrete caching and hit-rate data. Scope stays niche to local inference, so it lands at the featured threshold rather than must-write.

May 22Friday

AI HOT (Curated Pool)

Alibaba Qianwen App, PC, and Web Add Qwen3.7-Max

Alibaba added Qwen3.7-Max to the Qianwen app, PC client, and web client, with free access after updating the app to version 6.9.7 or later, and the official test reports a 35-hour autonomous kernel optimization run with more than 1,000 tool calls.

Why it matters: HKR-H/K/R all pass: Alibaba ships Qwen3.7-Max across three Qianwen clients, with v6.9.7+ free access and a 35-hour, 1,000+ tool-call claim. Benchmarks, context window, and API pricing are not disclosed, so it stays below 90.

May 21Thursday

Xinzhiyuan · WeChat

USTC Papers Study Lifelong Learning for LMMs via Multimodal Knowledge Injection

USTC researchers released MMEVOKE and KORE: MMEVOKE contains 9,422 samples across 159 subcategories, while KORE uses knowledge-tree augmentation and null-space constrained fine-tuning to reduce catastrophic forgetting during multimodal knowledge injection.

Why it matters: HKR-K and HKR-R are solid: the post gives dataset size plus a concrete fine-tuning mechanism. It stays in the low featured band because this is paper-level knowledge injection without production evidence or full reproducibility details.

May 20Wednesday

AI HOT (Curated Pool)

Qwen3.7: Agent Frontier

Qwen Studio released Qwen3.7 with chatbots, image and video understanding, and image generation. It also covers document processing, web search integration, tool calling, and artifact generation. The RSS snippet frames it as an agent-focused model, but the post does not disclose context length. It also omits benchmark scores, pricing, API limits, release schedule, and reproducible evaluation conditions.

Why it matters: HKR-H/K/R all pass: this is a Qwen flagship-model update with concrete capability coverage. Lack of benchmarks, pricing, and context-window details keeps it at the low end of the 85–94 band.

r/LocalLLaMA

Nemotron-Labs-Diffusion from NVIDIA

NVIDIA released the Nemotron-Labs-Diffusion 3B, 8B, and 14B dense model family with AR decoding, diffusion parallel decoding, and self-speculation; the 8B model reaches 850 tok/s on GB200 at concurrency 1, compared with 253 tok/s for AR and 360 tok/s for Eagle3.

Why it matters: HKR-H/K/R all pass: NVIDIA diffusion LLMs, concrete sizes/mechanisms, and an 850 tok/s GB200 claim. Single-source Reddit sourcing keeps it in the 78–84 band, not P1.

r/LocalLLaMA

KV cache quantization benchmarks: TurboQuant is overrated, q5 deserves attention, q8 may waste VRAM

Anbeeld benchmarked KV cache quantization for Qwen 3.6 27B on one RTX 3090 at 64k and 128k context, reporting q4_0 tail KLD 32% worse than q5_0 and turbo4 running 17% slower than q4_0 with little memory saving.

Why it matters: HKR-H/K/R all pass, with a first-person benchmark and concrete deltas. Scope is narrow: one RTX 3090, one model, and a Reddit source, so it stays near the featured threshold.

r/LocalLLaMA

Floor for local meeting summarization on a 6GB GPU: Qwen3.5 0.8B works in 57s, Granite 4 350M hallucinates

The author tested VoiceFlow 1.6.0 on an RTX 3060 Laptop 6GB, where Qwen3.5 0.8B summarized a 4-minute meeting in 57 seconds with 16K context, while Granite 4 350M returned summaries in 0.6-2.8 seconds but fabricated Binance and Star Trek content.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the test reports hardware/context/timing, and local meeting summarization hits privacy and cost nerves. Single Reddit experiment limits authority, so 73 featured.

May 19Tuesday

r/LocalLLaMA

Sapient Intelligence releases HRM-Text 1B: 40B tokens, ~$1k pretrain

Sapient Intelligence released HRM-Text 1B, a 1B-parameter model trained from scratch on 16 GPUs for 1.9 days with 40B tokens and a reported ~$1,000 budget; its self-reported chart shows MATH 56.2 and DROP 82.2, while independent evaluation remains pending.

Why it matters: HKR-H/K/R all pass: low-cost pretraining plus a smaller model beating a larger one is clickable, with concrete training and benchmark numbers. Independent eval is unfinished, so this stays at 78, not 85.

QbitAI · WeChat

JD and CAS IIE Publish Three Papers Defining Self-Taught RLVR

JD and CAS IIE released three Self-Taught RLVR papers covering RLSD, NPO, and CoPD; RLSD reports that 200 training steps on Qwen3-VL-8B-Instruct exceed GRPO at 400 steps across 8 benchmarks.

Why it matters: HKR-H/K/R pass: self-taught RLVR is a clear hook; RLSD reports 8 benchmarks and a 200-vs-400-step GRPO comparison; it hits reasoning fine-tuning cost. Not a top-lab model launch and replication heat is undisclosed, so it stays low featured.

AI HOT (Curated Pool)

Qwen3.7 Preview lands on Arena; Alibaba rises to fifth in vision ranking

Alibaba says Qwen3.7-Plus-Preview has landed on Arena and that Alibaba now ranks fifth in vision; the post does not disclose benchmark scores, the number of competing models, or a release timeline for the Qwen3.7 series.

Why it matters: HKR-H/K/R pass: Qwen3.7-Plus-Preview appears on Arena with a #5 vision rank. Score stays in the low featured band because the vendor post omits scores, model count, access, and timeline.

AI HOT (Curated Pool)

Qwen 3.7 Preview

The title identifies Qwen 3.7 Preview, but the post body is empty; it does not disclose parameter size, context window, release timing, access method, pricing, model card details, or benchmark results.

Why it matters: HKR-H/R pass because an official Qwen 3.7 preview matters for flagship-model competition, but HKR-K fails: no size, context, access path, or benchmarks are disclosed, so it stays at the featured threshold.

r/LocalLLaMA

llama.cpp MTP support landed: Qwen3.6 27B reaches 2.44× on Strix Halo

llama.cpp merged MTP speculative decoding in PR #22673; Qwen3.6 27B Q8_0 rose from 7.4 to 18.1 tok/s on Strix Halo, while a dual RTX 3090 Q8_0 setup rose from 25.7 to 55.9 tok/s.

Why it matters: HKR-H/K/R all pass: llama.cpp adds MTP speculative decoding with Qwen3.6 27B speedups on Strix Halo and RTX 3090. The scope is local inference, not a broad model release, so 78 fits featured.

May 18Monday

r/LocalLLaMA

Qwen 3.6 27B on 24GB VRAM: backend comparisons, quant choice, and settings

The author tested Qwen 3.6 27B on an RTX 3090 24GB and kept ik_llama.cpp with Qwen3.6-27B-MTP-IQ4_KS.gguf; at 156k context with q8_0 KV and MTP, a ~5.9k-token prompt plus 1024-token output reached about 1261 tok/s prefill and 72.9 tok/s decode, while vLLM lacked a clean single-card long-context run.

Why it matters: HKR-H/K/R all pass: this is a first-person local inference benchmark with concrete VRAM, quant, context, and speed numbers. Its reach is narrower than a model release, so it sits at the featured threshold.

r/LocalLLaMA

I trained TIME: short context-triggered thinking on Qwen instead of overthinking

An independent author trained TIME with QLoRA on Qwen3 4B/8B/14B/32B to trigger short mid-response reasoning when context changes; the post says datasets, notebooks, scripts, curriculum, and TIMEBench are public, with 24GB VRAM enough for training up to 14B.

Why it matters: HKR-H/K/R all pass: the post has a clear tuning hook, concrete reproducible details, and strong local-LLM resonance. Reddit single-post sourcing keeps it in the 72-77 featured band, below lab-level releases.

r/LocalLLaMA

LLMs on Android: Snapdragon 8 Elite MoE Experience

A Reddit user tested MoE LLMs on an Honor Magic 7 Pro with Snapdragon 8 Elite and 24GB RAM; under Q4 quantization, LFM2-24b-a2b reached about 24 tokens/s while Gemma reached about 11 tokens/s, and CPU inference was still faster than NPU or GPU in the reported setup.

Why it matters: HKR-H/K/R all pass: a named Reddit test gives hardware, quantization, and token/s figures. Single-device anecdote and weak source authority keep it at the low featured band.

May 17Sunday

r/LocalLLaMA

MiroThinker-1.7 Open-Weight Deep Research Agent Based on Qwen3 MoE

MiroMindAI released the MiroThinker-1.7-deepresearch and mini APIs, with the mini version using 30B total parameters and 3B active parameters, weights on HuggingFace, and context management based on sliding window K=5 plus episode restarts.

Why it matters: HKR-H/K/R all pass, but the source is a Reddit thread and the lab is not top-tier. Open weights, MoE sizing, and context-management details clear featured, not same-day must-write.