Skip to content

Alibaba's Qwen family: open releases and iterations, from flagship models to small on-device ones.

Latest picks

121–140 of 205

May 17Sunday

r/LocalLLaMA

85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B

Abliterlitics compared five Qwen3.6-27B abliteration variants against the base model using 85 GPU-hours of benchmarks, HarmBench, KL divergence, and weight forensics; Huihui had the smallest benchmark deltas, Heretic had the lowest KL divergence, and all five variants reached near-complete safety removal.

Why it matters: HKR-H/K/R all pass: the post gives an 85-GPU-hour comparison across five abliteration methods on Qwen3.6-27B. Niche open-model safety work, not a lab release, so it stays at the featured threshold.

r/LocalLLaMA

Same Models Tested Across Strix Halo, RTX 3090, and RTX 5070

C_Coffie published 55 local inference benchmark runs across Strix Halo, RTX 3090, RTX 5070, five backends, and 0.35B to 35B-A3B models; RTX 5070 beats RTX 3090 on models fitting 12GiB, while RTX 3090 leads in the 14–31B band that exceeds 12GiB but fits 24GiB.

Why it matters: Hits HKR-H/K/R with a named first-person benchmark: 55 runs and concrete GPU crossover points. Source is a single Reddit post, so it stays in the low featured band.

May 16Saturday

r/LocalLLaMA

Qwen3.6-35B-A3B and 9B land on the public Terminal-Bench 2.0 leaderboard

little-coder × Qwen3.6-35B-A3B scored 24.6% ±3.2 on Terminal-Bench 2.0, above Gemini 2.5 Pro on Gemini CLI at 19.6% and Qwen3-Coder-480B on Terminus 2 at 23.9%.

Why it matters: HKR-H/K/R all pass, but this is a Reddit post with leaderboard numbers only; test setup and reproducibility details are not disclosed. Strong code-agent benchmark signal, not a 78+ release story.

QbitAI · WeChat

Zhejiang University and Microsoft use 3,000 text prompts to improve video 3D consistency with World-R1

Zhejiang University and Microsoft introduced World-R1, training Wan 2.1 with about 3,000 text-only prompts, Flow-GRPO, and a four-part reward; the 1.3B version improves PSNR over the baseline by 10.23 dB.

Why it matters: HKR-H/K/R all pass: the hook is unusual, and the post gives 3,000 text samples, Flow-GRPO, and a +10.23 dB PSNR gain. Strong multimodal research, but not a foundation-model launch, so 78.

r/LocalLLaMA

Orthrus-Qwen3-8B: Up to 7.8× tokens/forward on Qwen3-8B with frozen backbone

Orthrus-Qwen3-8B adds a trainable diffusion attention head to a frozen Qwen3-8B backbone, reaches up to 7.8× tokens per forward and about 6× wall-clock speedup on MATH-500, while training 16% of parameters with under 1B tokens over 24 hours on 8×H200 GPUs.

Why it matters: HKR-H/K/R all pass: 7.8× speed is a strong hook, the post gives parameter, hardware, and benchmark conditions, and inference cost is a live practitioner pain. Single-source Reddit research keeps it in the lower good-quality band.

May 15Friday

r/LocalLLaMA

Used over a million tokens in three sessions to test Qwen 3.6 35B MTP

A Reddit user tested Qwen3.6-35B-A3B MTP across three million-token-scale sessions, using 300k context and KV Q8_0, and reported about 1.5x the tok/sec of earlier tests.

Why it matters: HKR-H/K/R all pass: the million-token test is clickable, 300k context and KV Q8_0 add testable detail, and local speed maps to cost. Source is one Reddit post, so it stays below the high-importance band.

r/LocalLLaMA

I Let a Small Model Train on Its Own Mistakes; It Reached 80% on HumanEval and Beat GPT-3.5 on Math

The author fine-tuned Qwen 2.5 7B base on self-mined mistake-correction pairs, raising HumanEval from 25/164 to 112/164; Qwen 2.5 14B used 100 pairs and a 95-minute H100 run costing $3.50.

Why it matters: HKR-H/K/R pass: the hook is strong and the post gives samples, H100 time, cost, and HumanEval deltas. Kept at 78 because it is a single Reddit post and the 80% claim differs from 112/164.

r/LocalLLaMA

The RTX 5000 PRO 48GB arrived and is better than expected

A Reddit user built a $5,600 RTX 5000 PRO 48GB PC and ran Qwen3.6-27B-FP8 with full-precision cache; they report up to 80 tok/s in TG, about 50–60 tok/s on very large prompts, 4,400 tok/s in prompt processing, and 200k tokens fitting in BF16 KV cache.

Why it matters: HKR-H/K/R all pass: a first-person local-inference test gives price and speed numbers, not vendor copy. Single Reddit source limits reach, so it lands in the featured-threshold band.

May 14Thursday

r/LocalLLaMA

Automated AI researcher running locally with llama.cpp

Hugging Face’s ml-intern added local-model support through llama.cpp and ollama; the post says Qwen3.6-35B-A3B can orchestrate CPU/GPU sandboxes and Hub jobs to run an end-to-end SFT workflow.

Why it matters: HKR-H/K/R all pass, but this is a Reddit-sourced open-source tool update, not a major model release. Local sandbox and Hub-job orchestration for SFT put it just above the featured threshold.

r/LocalLLaMA

Open-source one-prompt-to-cinematic-reel pipeline on one GPU with FLUX.2 and Wan2.2-I2V

The developer open-sourced StudioMI300, an 8-stage sequential pipeline that turns one English sentence into a 720p MP4 on a single AMD Instinct MI300X, cutting end-to-end time from 25.9 minutes to 10.4 minutes per clip.

Why it matters: HKR-H/K/R all pass: the post has a concrete one-GPU video pipeline, runtime numbers, and a local-build cost/control hook. Reddit single-source status and no third-party replication keep it below the 78+ band.

Synced · WeChat

ACL 2026: Alibaba DAMO I²B-LPO Improves RLVR Exploration

Alibaba DAMO Academy introduced I²B-LPO, an RLVR post-training framework that branches rollouts at high-entropy nodes and filters them with an information-bottleneck self-reward, reporting up to 5.3% accuracy gains and 7.4% semantic-diversity gains on math benchmarks using Qwen2.5-7B and Qwen3-14B.

Why it matters: HKR-H/K/R all pass: the ACL 2026 DAMO paper has a clear RLVR exploration hook, concrete I²B-LPO mechanics, and benchmark gains. It is still a training-method paper, not a major model or product release, so 78 fits the lower good-quality band.

AI HOT (Curated Pool)

UnslothAI Releases Qwen3.6 MTP GGUF Models With Over 1.4x Faster Inference

Daniel Han released experimental Qwen3.6 MTP GGUF models, with the 27B model reaching 140 tokens/s on one GPU and the 35B-A3B version reaching 220 tokens/s, using two draft tokens for speculative decoding.

Why it matters: HKR-H/K/R pass via concrete single-GPU speed claims and local-inference relevance. Score stays in low featured because the post is a single X source and does not disclose GPU, quantization settings, or repro steps.

r/LocalLLaMA

2x RTX 3090 setup for local Qwen 3.6 27B inference

A Reddit user ran Qwen 3.6 27B on a dual RTX 3090 Ubuntu setup, reporting 48GB VRAM, a 262k context window, no NVLink, about 4000 pp/s prompt processing, and 113 tk/s generation.

Why it matters: All HKR axes pass, and this is a first-person local-inference run with concrete numbers. Source is a single Reddit post with limited reproducibility detail, so it sits at the low featured threshold.

r/LocalLLaMA

24+ tok/s from ~30B MoE models on an old GTX 1080

User mdda ran Qwen 3.6 35B-A3B on an i7-6700, GTX 1080, and 32GB RAM machine at about 24 tok/s with 128k context; the setup uses llama.cpp MoE offloading plus TurboQuant/RotorQuant KV cache quantization, with PCIe 3.0 x16 saturated and GPU utilization at about 40–50%.

Why it matters: Single Reddit source limits authority, but the GTX 1080 + Qwen 3.6 35B-A3B + 128k + 24 tok/s setup gives a concrete local-inference result. HKR-H/K/R all pass; this is a practical featured item, not a major model or product launch.

May 13Wednesday

Hacker News front page

Show HN: Needle Distills Gemini Tool Calling into a 26M Model

Cactus open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, with MIT-licensed weights released on Hugging Face.

Why it matters: HKR-H/K/R all pass: the tiny Gemini-style tool-calling angle is clickable, with concrete speed and license claims. Source is still Show HN/GitHub self-reporting, not an independent benchmark or major lab release, so it stays below the 78–84 band.

r/LocalLLaMA

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Cactus Compute open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, using an attention-and-gating architecture with no MLPs.

Why it matters: HKR-H/K/R all pass: a 26M tool-calling model has a strong hook and concrete speed/design claims. Single Reddit source and a less-known team keep it in the lower 78–84 band.

May 12Tuesday

r/LocalLLaMA

Local LLM Autocomplete and Agentic Coding on a Single 16GB GPU + 64GB RAM

Reddit user grumd runs Qwen2.5-Coder-7B Q6 for autocomplete and Qwen3.6-35B-A3B Q8 for agentic coding on one RTX 5080 with RAM offloading; the post reports about 145k context, 56GB RAM used with other apps open, and Qwen3.6-35B-A3B speed of tg128 at 35.29 tokens/s.

Why it matters: HKR-H/K/R all pass: a named first-person local coding experiment with concrete model, quantization, context, and throughput data. Source is a single Reddit post without replication or comparisons, so it stays in the low featured band.

r/LocalLLaMA

Prompt caching for RL training: 7.5x speedup on long-prompt, short-response workloads

The author proposes prompt caching for RL training. On Qwen3.5-4B, it reports a 7.5x speedup with 16k-token prompts and 64-token outputs, and the G=8 example with 1000-token prompts and 100-token responses reduces 8800 processed tokens to 1800 unique tokens.

Why it matters: HKR-H/K/R all pass: the angle is novel, and the post gives 16k/64 plus G=8 token-dedup numbers. Kept at 78 because this is a single Reddit post without independent replication or a paper/code artifact disclosed.

May 11Monday

r/LocalLLaMA

ExLlamaV3 Major Updates

ExLlamaV3 added DFlash in v0.0.31, raising Coding throughput from 59.21 t/s to 177.67 t/s; v0.0.32 optimized five models, with Trinity-Nano gaining 72.4% on 6000 Pro², while v0.0.33 adds DFlash model quantization plus bug fixes and efficiency work.

Why it matters: HKR-H/K/R all pass, but the blast radius is mostly LocalLLaMA and ExLlama users. This fits a mid-weight open-source inference update, not a same-day industry-wide story.

Synced · WeChat

ICML 2026: PRISM Brings Efficient Test-Time Scaling to dLLMs

PRISM raises LLaDA-8B-Instruct on GSM8K from 67.58% to 85.30% by combining hierarchical trajectory search, partial remasking, and self-verified feedback, reducing dLLM test-time scaling cost from O(NT) toward O(N+KT) under a final candidate width K.

Why it matters: HKR-H/K/R all pass: the hook rejects brute-force scaling, the post gives GSM8K and complexity numbers, and it speaks to inference cost. Still an ICML framework paper, not a mainstream product release, so it sits in 78–84.