Skip to content

Alibaba's Qwen family: open releases and iterations, from flagship models to small on-device ones.

Latest picks

41–60 of 205

Jul 8Wednesday

AI HOT (Curated Pool)

vLLM's transformers backend now matches or beats hand-written native speed

HuggingFace announced that vLLM's transformers modeling backend now matches or beats hand-written native implementations in throughput. Benchmarks on Qwen3 4B, 32B, and 235B MoE models all hit or exceed native speed. Model authors can now get vLLM's optimizations for free with a single --model-impl transformers flag, no porting needed. The post doesn't disclose latency numbers, only throughput charts.

Why it matters: A solid engineering improvement with concrete benchmarks across three Qwen3 scales — directly useful for model deployers. But the audience is narrow; most AI pros don't touch inference backends, so resonance is weak, keeping it right at the featured threshold.

Jun 30Tuesday

Hacker News front page

Qwen 3.6 27B is the sweet spot for local development

Piotr Migdał tested Qwen 3.6 27B and calls it the first local model that works as a general intelligence. Running 8-bit quantized on a Macbook Max M5 128GB with llama.cpp and multi-token prediction, it hits 32 tok/s using 42GB RAM. It handled constrained writing, generated a hexagonal minesweeper npm package in one shot, and built a reactive landing page. The post includes full llama.cpp setup commands and recommends against Ollama on ethical grounds.

Why it matters: A first-person experiment with real numbers, not a press release. Qwen 3.6 is already a hot topic, and this piece adds practical local-deployment details. Score capped at 72 because it's a personal review, not an official launch or major product update.

Jun 24Wednesday

AI HOT (Curated Pool)

Qwen-AgentWorld open-sourced: an agent that predicts before it acts

Qwen released Qwen-AgentWorld, a native language world model covering seven domains: MCP, Search, Terminal, SWE, Web, OS, and Android. Trained on over 10 million real interaction trajectories through CPT→SFT→RL, it scored 58.71 on AgentWorldBench, edging out GPT-5.4 (58.25) and Claude Opus 4.8. As a decoupled environment simulator, it hit 50.3% F1 on WideSearch via Sim RL, beating real-environment RL at 45.6%. When used as an agent foundation model with LWM warm-up, it transfers to seven benchmarks—three of which never appeared in training. Both model and benchmark are open-sourced.

Why it matters: Qwen dropped an agent model with a clear methodology and benchmark — not concept hype. The 7-environment coverage and 10M+ training traces make it substantive, but it just went open-source and the community hasn't reproduced it yet, so the score stays below 85.

Hacker News front page

Qwen-AgentWorld: Language World Models That Simulate Environments for General Agents

Qwen team released Qwen-AgentWorld, a language world model that predicts environment dynamics for general agents. It covers 7 domains and uses long chain-of-thought reasoning to forecast next states. Two model sizes are available: 35B-A3B and 397B-A17B, trained on over 10 million real-world interaction trajectories via a three-stage pipeline—CPT injects world modeling from state transitions, SFT activates next-state prediction, and RL sharpens fidelity with hybrid rubric-and-rule rewards. The team also built AgentWorldBench from real interactions of 5 frontier models across 9 benchmarks. Qwen-AgentWorld significantly outperforms existing frontier models. It works in two modes: as a decoupled simulator enabling scalable RL across thousands of environments, surpassing real-environment-only training; and as a unified agent foundation model where world-model training serves as effective warm-up, boosting performance on 7 agentic benchmarks. Code is open-sourced.

Why it matters: Qwen team trains an LLM-based world simulator on 10M interaction traces across 7 environments. Novel approach with concrete scale and benchmarks, relevant for agent builders. Score capped below 85 because it's a paper, not a product release — real-world agent task gains aren't...

Jun 18Thursday

Hacker News front page

Local Qwen isn't a worse Opus, it's a different tool

Alex Ellis ran local models on an RTX 6000 Pro and recouped the cost in 2–3 months. Qwen 27B scores only 12% below Claude Opus 4.8 on SWE-Bench, but for Go distributed systems, quantized models hit infinite loops and hallucinations—he still won't trust them unsupervised. The post doesn't disclose token speed or latency figures.

Why it matters: Alex Ellis ran a real-world comparison of local Qwen vs Claude Opus on his own company's codebase, with concrete numbers and failure cases — not a vague opinion piece. Downside: no token speed or latency data disclosed, and the conclusion is anecdotal rather than systematic. B...

Jun 16Tuesday

r/LocalLLaMA

HalBench tests 29 open models on sycophancy and hallucination; Qwen 3.6 and Gemma 4 punch far above their weight

HalBench is an open benchmark that gives models a false premise and measures whether they push back or play along. v2.3 covers 33 models, 29 of them open. Only Sonnet 4.6 (65.1%) and Grok 4.3 (50.9%) clear 50% pushback. The best open model is Qwen 3.6 (~27B dense) at 36.6%, beating GPT-5.4 and Gemini 3.1 Pro. Gemma 4 26B follows at 29.2%. Model size barely predicts performance; phi-4 sits dead last at 2.3%. Dataset, scoring code, and Space are all open.

Why it matters: A community benchmark with concrete numbers and rankings, where Qwen 3.6 outperforms GPT-5.4 on refusal rate, is real signal. Not p1 because it's a self-built eval without peer review yet — treating it as a strong recommendation.

AI HOT (Curated Pool)

Local coding stack: Qwen 3.6 35B-A3B delivers 5x speedup for free

Tomasz Tunguz analyzed a 500+ comment Hacker News thread to map the local coding stack. Qwen 3.6 35B-A3B leads model mentions at 33%, with the 27B variant at 20%, followed by DeepSeek Pro and Gemma4 31B. All use MoE architectures that run on consumer hardware. For agents, Pi leads at 49% and OpenCode at 45%, both lightweight harnesses for local inference. One commenter compared local Qwen to a junior dev needing guidance versus Claude Opus as a senior who thinks with you on architecture—15x vs 5x speedup. But zero cost, full offline capability, and privacy make the tradeoff worthwhile for many. SWE-bench Verified scores back this up: Qwen3.6 27B hits 77.2%, the 35B-A3B MoE variant hits 73.4%, close to Claude Sonnet 4.6 at 79.6%.

Why it matters: Tunguz mined real local coding stack configs from 500+ HN comments: Qwen 3.6 35B-A3B at 33%, Pi at 49%, with MoE enabling consumer GPU inference. Concrete data with comparisons, not vendor fluff. Docked because it's secondhand curation rather than firsthand benchmarking, and t...

AI HOT (Curated Pool)

Qwen-RobotNav: One model, five navigation domains, and a tool-call primitive for agentic systems

Qwen released Qwen-RobotNav, a single set of weights built on Qwen3-VL and trained on 15.6M samples that handles instruction following, object search, tracking, driving, and embodied QA. It exposes visual context as tunable inference-time parameters—token budget, temporal decay, per-camera weights—so an upper-level planner (Qwen3.7-Plus) can reconfigure it per call without retraining. On EXPRESS-Bench it beats the prior best by 15.4% while using 77% fewer navigation steps. Zero-shot deployment on a Unitree Go2 with a single low-res camera works in unseen outdoor environments.

Why it matters: Qwen ships a robot navigation model built on Qwen3-VL with 15.6M samples across five tasks. The core pitch is a parameterized visual memory interface configurable at inference time—frame count, attention weights, no retraining needed. Paper and GitHub are available, but no rea...

AI HOT (Curated Pool)

Qwen-RobotWorld: A world model that unifies 20+ robot embodiments via natural language

Qwen released an embodied world model that treats natural language as a universal action interface, covering 20+ robot embodiments and 500+ action categories without per-robot control APIs. It uses Qwen2.5-VL as the action encoder, trained jointly on 8.6M video-text pairs across manipulation, autonomous driving, and indoor navigation, and claims top results on 4 benchmarks. The model generates 2–4 geometrically consistent views and supports human-to-robot transfer across 14 morphologies. I'd hold for real-world latency numbers—the post doesn't name the benchmarks or disclose inference speed.

Why it matters: Qwen drops RobotWorld, an embodied world model using natural language as a universal action interface, trained on 8.6M video-text pairs across three domains. The scale and cross-embodiment approach are substantive. Not scoring higher because it's a blog + paper release with no...

Jun 12Friday

AI HOT (Curated Pool)

inclusionAI releases VISTA-4B, a vision-language model for GUI element grounding

inclusionAI open-sourced VISTA-4B on Hugging Face, a 4B-parameter vision-language model built on Qwen3.5. It focuses on GUI grounding: given a screenshot and a text instruction, the model pinpoints the target button or region. The model card lists gui-grounding and reinforcement-learning tags, indicating RL was used to improve localization accuracy. Code examples cover Transformers, vLLM, and SGLang, under an Apache 2.0 license. The post doesn't disclose benchmark scores, training data size, or inference latency—I'd hold off on performance claims until those numbers surface.

Why it matters: A 4B GUI grounding model is a practical direction and RL training is a real technical signal, but the model card has zero benchmarks, no training data disclosure, and no comparison to OmniParser or UI-TARS. Too many gaps to push higher.

Jun 10Wednesday

r/LocalLLaMA

Fine-tuned Qwen2.5-7B to 96% of Claude Haiku using about $3 of API calls

A Reddit user fine-tuned Qwen2.5-7B with 1,040 DV-DPO preference pairs, costing about $3 at Claude Haiku rates, and reported 96% composite performance versus Claude Haiku on a domain task, with 11-second latency on a 4-bit T4 setup.

Why it matters: HKR-H/K/R all pass: the hook is cheap fine-tuning, with concrete DV-DPO and latency numbers. Kept at 74 because it is a single Reddit post; task scope and eval details are thin.

AI HOT (Curated Pool)

Google Gemini 3.5 Live Translate enters public preview with 70+ languages

Google released Gemini 3.5 Live Translate in public preview through the Gemini API, offering low-latency speech-to-speech translation across 70+ languages and 2,000 language pairs.

Why it matters: HKR-H/K/R all pass: Google’s speech-to-speech translation API has a clear developer hook and concrete scale numbers. Single X-source detail and missing price, latency benchmarks, and regions keep it at 78.

r/LocalLLaMA

ICML paper on predictable hallucination gate and ntkMirror open-weight implementation

An ICML 2026 paper presents an ISR=1 answer-abstain gate for evidence-grounded QA, and ntkMirror implements it for local open-weight models with multiple evidence orderings, reporting 0.0–0.7% hallucination at about 24% abstention in the held-out audit.

Why it matters: HKR-H/K/R all pass: an ICML paper with an open implementation, a concrete ISR=1 gate, and measured abstention-vs-hallucination tradeoff. Scope stays within evidence QA/RAG reliability, so it sits below must-write level.

Jun 9Tuesday

AI HOT (Curated Pool)

Cohere Releases North Mini Code, an Open Coding Model for Developers

Cohere released North Mini Code, a 30B-parameter MoE coding model with 3B active parameters, under Apache 2.0; it supports 64K/128K context lengths and reaches 80.2% pass@10 on SWE-Bench Verified.

Why it matters: HKR-H comes from a compact MoE code model with a strong SWE-Bench claim; HKR-K has params, license, context, and benchmark. Cohere is notable but not a frontier-lab launch, so this fits the 78–84 open-source code-model band.

AI HOT (Curated Pool)

Qwen3.7-Max Delivers Mobile and Web Apps from Scratch Using One Document

Qwen3.7-Max delivered mobile and web applications from a roughly 150,000-character product research document without design files or backend code; each client took about 4 hours, used staged constraint injection and error feedback, and the web app passed typecheck, build, and 34 reachable routes.

Why it matters: HKR-H/K/R all pass: the coding-agent claim is clickable, quantified, and emotionally relevant to developers. The summary lacks eval setup, failure rate, and human-intervention detail, so it stays in the 78–84 band.

r/LocalLLaMA

2X tk/s on 1× MI50: Qwen3.6-27B inference rises from 19.4 to 38.1 tk/s

bigattichouse raised Qwen3.6-27B throughput on a single MI50 from 19.4 to 38.1 tk/s by running same-model parallel computations for Q8-or-lower quantization, exploiting unused compute lanes instead of adding a smaller speculative decoding model.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person experiment with numbers and a hypothesis, not a validated release. No code or broader replication is disclosed, so it stays at the featured threshold.

r/LocalLLaMA

Levi: Run AlphaEvolve on Your Local Qwen 30B

LEVI runs an AlphaEvolve-like search system with Qwen3-30B-A3B and reports tests on ADRS, IFBench, and HotpotQA, claiming up to 35x lower cost overall and up to 12x fewer evals under the same single-model, same-budget comparison.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with model, benchmarks, and cost ratios only; code maturity and reproducibility details are not disclosed. Scores as a strong open-source agent/inference item, not a major release.

Jun 8Monday

r/LocalLLaMA

Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax

Luce Spark runs Qwen3.6 35B-A3B at 13.3 GiB peak VRAM on an RTX 3090 by keeping hot experts on GPU, swapping cold experts through a bounded async cache, and using one fused graph for decode at about 100 tok/s.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE on a 16 GB GPU, with 13.3 GiB peak use and ~100 tok/s. Reddit-source and no third-party replication keep it at 78.

r/LocalLLaMA

DFlash Speculative Decoding and KV Cache Compression on RTX 5090 Show 3.26x Speedup

The author tested Qwen3.6-27B on an RTX 5090 with DFlash plus KV cache compression, reaching up to 3.26x speedup; q4_0/turbo4 delivered 3.18x speedup with only +0.02% PPL on WikiText-2.

Why it matters: HKR-H/K/R all pass: RTX 5090 testing, DFlash speculative decoding, KV cache compression, 3.26x speedup, and PPL delta are concrete. Single Reddit source keeps it near the featured floor.

r/LocalLLaMA

Weird to get near-linear scaling by adding another GPU?

A Reddit user benchmarked qwen3.6-27b-autoround-int4 on 1x3090 versus 2x3090. Narrative decode rose from 53 TPS to 94 TPS, and code decode rose from 62 TPS to 120 TPS, under no NVLink, 8x/8x PCIe, P2P automatically enabled, tensor parallelism set to 2, and different KV-cache settings.

Why it matters: HKR-H/K/R all pass: the result is counterintuitive, includes concrete TPS and TP conditions, and speaks to local-inference cost. Single Reddit test lacks multi-model replication and full setup details, so it stays near the featured threshold.