Skip to content

Alibaba's Qwen family: open releases and iterations, from flagship models to small on-device ones.

Latest picks

1–20 of 205

Sep 24Thursday

Computing Life · Share · Yage

Qwen-Image-2.1: A version rollback that packs text rendering, editing, and native RGBA into one open-weight model

Qwen released Qwen-Image-2.1 on Sep 20, a 7B open-weight image model that unifies text-to-image, local editing, and native RGBA output in a single pipeline. The version number rolled back from 3.0 to 2.1 reflects a 2026 split: 3.0 is a closed-source commercial API, while 2.1 continues the open research branch. A built-in RGBA VAE outputs PNGs with transparency, skipping external matting. The interface supports up to 10 reference images and three mask types at native 2K. The license shifted from Apache 2.0 to a research-only agreement; commercial use requires a separate license. Community tests show ~25s per megapixel image on RTX 5070/5080 at 25 steps, ~15.6GB VRAM with Q8 quantization. Text rendering remains a strength, but multi-subject consistency shows facial generalization on well-known public figures—official demos don't guarantee universal performance.

Why it matters: Qwen open-sourced a model that combines image generation, editing, and native transparency output into one pipeline — a clear engineering increment, not a reskin. The backward version jump is inherently clickable, and it resonates with both designers and developers. Not scorin...

Sep 23Wednesday

AI HOT (Curated Pool)

Qwen releases Qwen-Audio-3.1 family: ASR, TTS, Realtime upgrades plus new TTS-Next and ASR-Next

Qwen dropped five audio models covering ASR, TTS, real-time conversation, and creative generation. The Realtime model supports interruption and slows down with empathetic responses when it detects low mood. Pricing is slashed: TTS ~70% off, Realtime ~85% off, ASR up to 95% off. The post doesn't disclose benchmark scores or latency numbers.

Why it matters: Full Qwen audio stack refresh with two new product lines (TTS-Next, ASR-Next), not just a version bump. The ~70% TTS price cut and Realtime emotion-aware interruption are verifiable details. Held below 85 because the post doesn't disclose ASR-Next and TTS-Next capability bound...

Sep 22Tuesday

Hacker News front page

Frontier AI on Your Own Hardware

Tim Dettmers's dlab is open-sourcing a full stack this week to run frontier AI on local hardware. An agent auto-optimized Metal kernels to run Qwen 3.6 35B-A3B at 1.5 bits per weight, hitting 450 tokens/s on a Mac. The core argument: the unit of research is no longer the paper but a coherent ecosystem. Full details are still under wraps, but the release includes an autonomous research agent, efficient test-time scaling, and auto-compaction that beats Claude Code on token savings.

Why it matters: Tim Dettmers is a key figure in quantization, and this isn't a single paper but a full toolchain release with concrete numbers (1.5 bits, 450 tok/s) and a reproducible path. The deduction: it's a blog announcement — actual usability and compatibility won't be clear until the o...

Sep 20Sunday

AI HOT (Curated Pool)

Qwen open-sources Qwen-Image-2.1: a 7B model unifying generation and editing with native transparency

Qwen released Qwen-Image-2.1, a 7B model that merges text-to-image generation and image editing into one lightweight system. It natively handles transparent images—generating them from prompts, editing layers, and extracting subjects from photos as RGBA assets. Editing supports up to 10 reference images, local edits, and identity preservation. A mixed-granularity attention design with KV cache reuse cuts inference cost for multi-image tasks. The model is open-sourced on GitHub, Hugging Face, and ModelScope.

Why it matters: Qwen open-sources a 7B unified image model with native transparency — a real differentiator, not a benchmark flex. Editing supports up to 10 reference images, which is practically useful. Score held back because the post doesn't disclose inference latency or VRAM requirements,...

Sep 18Friday

AI HOT (Curated Pool)

Qwen launches Qwen3.8-LiveTranslate real-time interpretation model with 2.3s latency and speaker separation

Qwen3.8-LiveTranslate cuts simultaneous interpretation latency from 2.8s to 2.3s by interleaving audio and text into a single stream. It supports 60 input languages, real-time speaker separation with voice cloning, and synchronized bilingual output. On the Omnilingua-MSpeaker benchmark it outperforms current mainstream systems in faithfulness, fluency, and conciseness. API is available via Alibaba Cloud DashScope.

Why it matters: Qwen ships a real-time interpretation model with 2.3s latency, speaker diarization, and bilingual subtitles — three new capabilities that push simultaneous interpretation beyond translation into scene understanding. Score stays below 85 because only the official blog is availa...

AI HOT (Curated Pool)

Qwen launches Qwen3.8-Omni-Flash, a native omnimodal model built for audio-visual agent workflows

Qwen3.8-Omni-Flash is a native omnimodal model that shifts focus from audio-visual understanding to task planning, tool use, and delivery in real-world workflows. It supports a 1M-token context window, with average scores across 29 evals up over 25% vs Qwen3.5-Omni-Plus. API pricing for audio input dropped over 98%, and audio-visual input over 93%. It gained 36.5 points on WildClawBench-MM and 22.3 on AgenticVBench; AliMeeting DER fell from 88.11 to 3.35. Qwen claims overall audio performance exceeds Gemini 3.8 Flash, with audio-visual performance close to it. Qwen-Live Harness is open-sourced for real-time interaction, and Qwen-MM-Plugins now supports tool use and workflows for long-form audio and video.

Why it matters: Qwen pushes omnimodal models from understanding to task delivery, backed by concrete benchmarks and pricing. Score stays at 82 rather than higher because it's a launch-day post with no third-party validation or cross-source cluster yet.

Sep 17Thursday

Hacker News front page

Training a 4B model to produce 81% faster query plans than Postgres

Rohan Bansal post-trained a Qwen 4B model to beat Postgres's default query plans. After SFT distillation from 500 GPT-6 Astra trajectories and a custom GRPO variant for RL, the model achieved 44.7% latency reduction and 81% geometric mean speedup across 113 join-heavy queries. Training ran on a rented 2×H100 node with four Postgres containers on his desk for measurement. The post doesn't disclose total training time or per-inference latency.

Why it matters: A 4B model trained via RL beats Postgres default plans by 81% on the Join Order Benchmark. The method is practically interesting, but only 113 queries were tested—generalization is unproven, capping the score at 78.

Sep 7Monday

r/LocalLLaMA

Qwen 3.8-27B NVFP4 beats Q5_K_M and nears BF16 after tweaking temp and min_p

A user compared Qwen 3.8-27B NVFP4 (NInfer) against Q5_K_M (llama.cpp) on a 5090. With default sampling, Q5 led on IFBench strict (76% vs 74%), but NVFP4's loose score was 78%, meaning its failures were mostly minor formatting drift. After lowering temp to 0.9 and setting min_p to 0.05, NVFP4 hit 80% strict while Q5 stayed at 76%. The official BF16 baseline is 79.5% on 300 samples; NVFP4's 80% on 50 samples has variance but the trend held across reruns. NVFP4 also ran nearly 3x faster and used far less VRAM. The takeaway: default sampling hides NVFP4's real quality—tighten it and you get near-BF16 instruction following.

Why it matters: First-person benchmark with concrete numbers and a reproducible recipe — not armchair theory. Directly useful for local LLM users. Score capped because it's a single-GPU, single-model case relying on the niche NInfer engine, limiting generalizability.

Sep 6Sunday

Computing Life · Share · Yage

tok/s is the most deceptive performance number in agentic scenarios

The author tested Cerebras-hosted Qwen 27B (claimed 1,500 tok/s) against a local instance (measured 102 tok/s) in real coding workflows. After accounting for prefill, actual effective throughput differed by only 3.5×, not 15×. Worse, the cloud session's context exploded from 11K to 99K in 65 seconds, hitting 865K total tokens against a 450K TPM cap—killed by a 429 error in three minutes with a $1.57 bill. The same task locally took 11 minutes and cost $0.017 in electricity, running stably for hours. The takeaway: ignore advertised tok/s for agentic workloads; measure end-to-end effective throughput and weigh context lifespan, cost, and rate limits. The query tool is open-sourced.

Why it matters: The author uses 30 days of real agentic workflow data to pull Cerebras' claimed 1500 tok/s down to an effective 357 tok/s — only 3.5x faster than local. Concrete numbers and methodology, not hand-waving. Not scored higher because the cloud sample is small (28 generations) and ...

Sep 4Friday

r/LocalLLaMA

Qwen3.8-27b called the first local model users can 'blindly trust'

A Reddit user reports that Qwen3.8-27b ran 8+ hours of continuous agentic work without a single mistake, making it the first local model they trust like a frontier model. Another user confirmed 20-hour sessions with sub-agents and commit gates, and said the INT8 quant even solved a coding problem that DeepSeek V4 Flash couldn't fix. The post doesn't disclose specific task types or failure rates, but the community feedback points to noticeably better reliability in long-chain agent workflows. Take it as personal experience, not a systematic eval.

Why it matters: Two independent users report Qwen3.8-27b's stability in multi-hour agent tasks, one with a direct comparison to DeepSeek V4 Flash. But the post doesn't specify task types or failure criteria — this is community word-of-mouth, not a reproducible eval. Score 72 at the featured t...

Computing Life · Share · Yage

AgentFlow trains a 7B decision node in the loop, gaining 17.2 points over swapping in GPT-4o

Stanford's AgentFlow paper shows that in the same agent orchestration, swapping a frozen Qwen2.5-7B decision node for GPT-4o adds only 5.8 points on average across six benchmarks. Training that same 7B node with real tool feedback adds 17.2 points. Only the Planner's selection policy is updated; the system skeleton stays fixed. The model learned to prefer Wikipedia over Google for medical queries, and tool-calling errors dropped by up to 28.4%. The post also lists four gates for real-world adoption: high-frequency tasks, automatic success verification, bottlenecks truly in decision logic, and a resettable environment. The cost story is incomplete—the paper discloses 8×A100 but not total training time or the cumulative bill for the GPT-4o judge.

Why it matters: AgentFlow from Stanford answers a concrete bottleneck question for agent builders: swapping in GPT-4o only adds 5.8 points, but training the 7B decision node on real execution feedback adds 17.2. Has numbers, mechanism, and engineering reproducibility—directly actionable signa...

Sep 2Wednesday

Hacker News front page

Slotstream runs the 104GB Qwen3.8-Flash-Next on a 48GB Mac at ~12 tok/s

carloslfu open-sourced Slotstream, an MLX + Swift tool that runs the 125B-parameter MoE model Qwen3.8-Flash-Next (104GB at 4-bit) on a Mac with only 48GB RAM. It streams expert modules from SSD on demand instead of loading everything into memory. Speed is ~12 tok/s, and it exposes an Ollama-compatible API. The post doesn't disclose time-to-first-token or the SSD model used, so real-world feel is still an open question.

Why it matters: Streams MoE experts from SSD on demand via MLX + Swift, letting a 104GB Qwen 125B model hit ~12 tok/s on a 48GB Mac. Clean engineering with an Ollama-compatible API that lowers the trial barrier. Docked a few points because it's a solo project with no community validation or m...

Sep 1Tuesday

Hacker News front page

Hugging Face Summer 2026: Chinese labs ship the biggest open models, but small models drive real usage

Hugging Face's biannual report covers Jan–Aug 2026. Chinese labs released the largest open models almost every month, ranging from 754B to 2.78T parameters, while US labs mostly stayed under 130B except for NVIDIA's Nemotron 3 Ultra (561B) and Thinking Machines Lab's Inkling. Attention doesn't equal adoption: 85.6% of models have under 200 lifetime downloads, and 1.5% of repos account for 99.2% of downloads. Qwen is now the community's go-to base model, small models remain the practical layer, and agents are emerging as the new user of models.

Why it matters: Hugging Face's biannual ecosystem report with concrete numbers and a US-China comparison framework hits all three HKR axes. Deduction because it's a survey, not a primary release, and the body only gives an excerpt — full data requires clicking through.

Aug 29Saturday

r/LocalLLaMA

Qwen 3.8 27B hits 50 tok/s with 100k context on a 16GB GPU

A user squeezed Qwen 3.8 27B into a 16GB RTX 4070 Ti SUPER, hitting 47–50 tok/s with a 100k-token context window. The trick is beellama.cpp's asymmetric kvarn KV cache quantization—kvarn5 for K, kvarn4 for V—which saved ~6% VRAM and pushed context from 88k to 100k. The model is jrell's IQ4_XS hybrid quant, purpose-built for MTP and long contexts on 16GB cards. MTP speculative decoding with 2 draft tokens and a 1,024-token high-precision tail helped keep quality up. Worth noting: this is a single-user tuning report, not a benchmark, so your mileage may vary.

Why it matters: A community tip with concrete parameters and real-world numbers, directly useful for 16GB GPU owners. Not scored higher because it's a user-shared trick rather than an official release, and beellama.cpp is a niche engine with a smaller audience.

Aug 26Wednesday

AI HOT (Curated Pool)

Qwen3.8-Flash-Next open-sourced: 125B total, 6B activated, previewing Qwen4 architecture

Qwen released Qwen3.8-Flash-Next weights as an early preview of the Qwen4 architecture. The model has 125B total parameters, activates only 6B per token, and carries an extra 51B N-gram embedding table that can be offloaded to host memory. Four architectural changes: attention uses Gated DeltaNet plus Qwen Sparse Attention for long-sequence compression and sparse block selection; residuals become four-branch gated residuals; embeddings add N-gram lookup for cheap capacity scaling; the optimizer switches to Muon. Training cost is roughly 1/9 of Qwen3.7-Plus, yet it scores higher on coding and office benchmarks. API pricing is $0.16 per million input tokens and $0.47 per million output tokens. Native context is 262K, extendable to 1M with YaRN. Take the scores with a grain of salt—they come from Qwen's own tech report; wait for community reproduction.

Why it matters: Early Qwen4 architecture preview: 125B total params, 6B activated, attention layers replaced with Gated DeltaNet plus sparse attention. Concrete new mechanisms. A flagship Chinese model architecture release with direct relevance for inference and open-source work. Score not hi...

Aug 25Tuesday

Hacker News front page

Qwen 3.8-Flash-Next open-release tomorrow: 125B total, 6B active MoE model

Qwen teased Qwen3.8-Flash-Next on ModelScope, a multimodal MoE model built on the next-gen Qwen4 architecture with 125B total and ~6B active parameters. The early release is meant to preview Qwen4's design for the community. It drops 2026-08-26 15:00 UTC, with an FP8 variant alongside. The post doesn't disclose benchmarks, inference speed, or specific multimodal capabilities—I'll hold judgment until the model card lands.

Why it matters: Qwen is previewing the Qwen4 architecture with a 125B-total / 6B-active MoE design — real new information with high attention in the Chinese open-source community. The deduction is because it's not open-sourced until tomorrow, and no benchmarks or inference speed data are avai...

Aug 21Friday

Hacker News front page

Nari Labs pushes Qwen3-TTS to sub-50 ms time-to-first-audio at 10 RPS on a single H100

Nari Labs open-sourced a Qwen3-TTS 1.7B CustomVoice serving implementation that hits sub-50 ms p95 time-to-first-audio at 10 RPS on a single H100 SXM with zero underruns. They benchmarked against vLLM-Omni, SGLang-Omni, VoxServe, and M*—default p95 latencies ranged from 277 to 1,160 ms at 1 RPS. At full utilization the system costs roughly $2 per 1M characters, compared to $100 for ElevenLabs V3 and $49 for Cartesia Sonic 3.5. Key optimizations include dynamic leading-silence trimming (~80 ms saved) and tuned codec-frame accumulation. Code and benchmarks are public; the post does not disclose underrun details at higher concurrency or long-form performance.

Why it matters: Nari Labs open-sourced a deployment recipe for Qwen3-TTS 1.7B that hits sub-50 ms p95 time-to-first-audio at 10 concurrent requests on a single H100—an order-of-magnitude improvement over vLLM-Omni and others. The post includes concrete benchmarks and reproducible optimization...

Aug 20Thursday

AI HOT (Curated Pool)

Alibaba releases Qwen-UI-Agent, a GUI agent foundation model that operates across mobile, desktop, web, and deep search

Alibaba's Qwen team released Qwen-UI-Agent, a model built to understand and operate mobile, desktop, and web interfaces. It scores 92.2% on the real-device benchmark MobileWorld-Real and 97.5% on AndroidDaily. On desktop, it hits 79.5% on OSWorld-Verified, beating GPT-5.5 and Gemini 3.1 Pro. Training used over 100 real phones and 150+ apps, and the model handles tasks longer than 100 steps. It pauses for user confirmation on payments or privacy actions. The tech report, project page, and GitHub repo are all public.

Why it matters: Alibaba Qwen dropped a GUI-focused agent model with real-device scores beating GPT-5.5 and Gemini 3.1 Pro across mobile and desktop benchmarks — a flagship-level agent capability update from a domestic lab. Score capped at 82 because the body only provides headline and summary...

Aug 19Wednesday

Hacker News front page

modelmap: paste a HuggingFace model ID, get an interactive architecture map

modelmap.cc turns any HuggingFace model ID into an interactive architecture diagram—no weights downloaded. It instantiates the model on a meta device to get the structure, then runs a traced fake forward pass to infer tensor shapes. The homepage shows trending models like Qwen3.8-27B, DeepSeek-V4-Pro-0813, and Kimi-K3, plus classic reference architectures such as GPT-2, BERT, and DeepSeek-V3.1. Public repos work out of the box; gated ones work after adding a token. The post doesn't disclose whether it's open source, backend costs, or concurrency limits.

Why it matters: A practical Show HN tool that maps model architectures without downloading weights, featuring DeepSeek-V4-Pro and Kimi-K3 on the landing page. Hits H and K but lacks the discussion hook R requires — more of a bookmark than a conversation piece. Scores 72 at the featured thresh...

Aug 18Tuesday

Hacker News front page

Qwen3.8 27B scores 52 on Artificial Analysis, ranking #1 among open-weight models

Alibaba's Qwen3.8 27B, released August 2026, tops the Artificial Analysis Intelligence Index with a score of 52 across 135 models. The index aggregates 9 evals covering agentic tasks, coding, scientific reasoning, and knowledge. The model is very verbose—160M output tokens, nearly 4× the median. API pricing shows $0; the post doesn't clarify whether this is a free tier or missing data. Weights are on Hugging Face under Apache 2.0, with text+image input and a 256k-token context window.

Why it matters: Qwen3.8 27B hits #1 on the Artificial Analysis Intelligence Index with a score of 52, the highest among open-weight models. Solid data with concrete numbers and a deployment caveat, hitting all three HKR axes. Not scoring higher because this is a third-party benchmark rather t...