Skip to content

Alibaba's Qwen family: open releases and iterations, from flagship models to small on-device ones.

Latest picks

21–40 of 205

Aug 17Monday

Hacker News front page

Qwen3.8 27B at 256K context on a 24GB GPU hits 50 tok/s with MTP

The author runs Qwen3.8 27B at its full 256K context on a single 24GB RTX PRO 4000 SFF, averaging 50.44 tok/s. The gain comes from a custom NVFP4 quant that protects sensitive layers, embedded MTP speculative decoding, and tuned CUDA kernels—not from any single component. Target-only decoding hits 21.19 tok/s; MTP pushes it to 59.46 tok/s. At a nearly full 256K cache, throughput drops to 12.61 tok/s without OOM. The post doesn't disclose total cost, but it's a detailed engineering log, not a plug-and-play recipe.

Why it matters: A solid hands-on local inference post with reproducible numbers and a clear technical path. Hits all three HKR axes, but it's a personal experiment, not an official release or industry event, so it lands at the featured threshold of 78.

Financial Times · Technology

The next China shock will come from open-source AI

An FT op-ed argues that China's open-source LLMs are repeating the playbook of its manufacturing boom—turning tech into a commodity at ultra-low cost and eroding Western pricing power. It names DeepSeek and Alibaba's Qwen series as key examples, noting their open-source strategy builds ecosystems fast while US firms stay closed-source and capex-heavy. The post doesn't cite specific market share or enterprise adoption figures, so treat this as a directional argument.

Why it matters: FT op-ed frames Chinese open-source AI as a replay of the manufacturing shock — a catchy angle, but the body lacks hard data, making it more of a directional warning. H and R hit, K is missing evidence, landing right at the featured threshold.

AI HOT (Curated Pool)

Qwen 3.8 27B is excellent, but defaults to wildly overthinking things

Simon Willison tested Alibaba's Qwen 3.8 27B and found the default xhigh reasoning effort causes absurd overthinking. A simple circle prompt triggered minutes of animated SVG generation; a pelican-on-a-bike SVG burned 22,276 reasoning tokens over 21 minutes. Turning reasoning off cut the same task to just over two minutes. He recommends starting with low or no reasoning. The model also nailed bounding-box detection on a pelican photo with near-perfect accuracy.

Why it matters: Simon Willison's hands-on test of Qwen 3.8 27B reveals severe overthinking from default reasoning settings, with concrete token and time comparisons. A data-backed first-person experiment directly useful for local deployment users. Not above 80 because the core finding is a co...

Aug 16Sunday

Hacker News front page

A leaderboard tracking 30 model cards to see which benchmarks frontier labs actually report

This project scanned 30 model cards from 11 orgs and counted how often 79 benchmarks are mentioned—it measures vendor attention, not benchmark quality. MATH-500 and Arena-Hard are near ceiling, losing discriminative power. DeepSeek's own models gained 40.6 points on AIME and 25.4 on LiveCodeBench in 26 days. Six benchmarks, including BrowseComp and SWE-bench Pro, are reported by at least 4 orgs but have no readable scores. The newer APEX-Agents already appears in 3 independent cards, though scores couldn't be read either.

Why it matters: Scans 30 model cards from 11 orgs, measuring vendor attention rather than benchmark quality — a useful lens. Concrete numbers like MATH-500 near-saturation and DeepSeek's 40.6-point AIME jump in 26 days will spark discussion. Docked because it's a personal project with limited...

Aug 13Thursday

AI HOT (Curated Pool)

Alibaba open-sources Qwen3.8-2.4T-A95B: 2.4T MoE, 95B active, native 256K context

Alibaba's Qwen team open-sourced its first Qwen-Max-level weights. Qwen3.8-2.4T-A95B has 2.4T total parameters with 95B active per token, native 262K context expandable to 1.01M tokens. It uses a 512-expert MoE, routing 10 experts plus one shared expert per token, and includes multi-token prediction training. The model targets coding, office tasks, research, and long-horizon agent workflows. Benchmarks against Opus 4.8, Fable 5, and GPT 5.6 Sol show mixed results, with top scores on PaperBench and IFBench among listed models. Post-training combines combinatorial environment scaling, a unified reward system, and an online data balancer to reduce gradient variance. The post does not disclose the open-source license or inference hardware requirements.

Why it matters: Alibaba's first full open release of a cloud-grade flagship — 2.4T total params, 95B activated, native 256K context — puts it in the top tier. Hits all three HKR axes and triggers the domestic flagship model positive signal. Held back from 90+ because we only have the announce...

Aug 8Saturday

AI HOT (Curated Pool)

Apple support doc confirms Qwen integration for Apple Intelligence on Mac in China

An Apple support doc updated on Aug 8 confirms that Apple Intelligence on Mac will work with Alibaba's Qwen model in China. It requires macOS 26.6 or later, a mainland China Mac, and a mainland China Apple Account. Users enable the Qwen extension in System Settings and sign in with a Qwen account. The extension covers Writing Tools and Siri: it can compose and rewrite text in Notes and Mail, and Siri can hand off requests like writing a poem or summarizing a document to Qwen after asking the user. Alibaba previously stated Qwen would integrate into Apple Intelligence across iOS, iPadOS, macOS, and visionOS; this doc is the first concrete sign of rollout.

Why it matters: Apple's official support doc confirms Qwen integration for China-region Macs — a substantive partnership between a top domestic model and a major hardware ecosystem. Concrete details on requirements and scope. Not scored higher because it's a doc update without real-world perf...

Aug 6Thursday

AI HOT (Curated Pool)

Alibaba Cloud launches Qwen-Image-3.0 with high-res image generation starting at $0.03

Qwen-Image-3.0 is pitched as production-ready: 4.5K-token prompts, 100%+ text accuracy with no broken logos, and native support for 12 languages. High-res generation starts at $0.03. The post only provides a headline and links—no model architecture, inference speed, or benchmarks are disclosed, so I'd hold off on the accuracy claim until third-party tests appear.

Why it matters: Qwen's first dedicated image gen model, priced at $0.03 with a 4,500-token prompt ceiling and 12-language support — real differentiators. Score held back because the source is a single tweet with no architecture details, inference speed, or third-party benchmarks; the '>100% t...

Aug 5Wednesday

Hacker News front page

Qwen Image 3.0 Pro targets production use with 4.5k-token layouts, 10px text, and realistic detail

Qwen Image 3.0 Pro handles up to 4,500 input tokens and generates dense layouts—newspapers, storyboards, menus—in one pass. It reliably renders text down to 10px across 12 languages and 20+ fonts, and reproduces micro-expressions, pores, and hair strands at near-photographic quality. Output pricing is $0.04 per 1K image and $0.075 per 2K image, but the rate limit is just 1 request per minute, so high-throughput use cases are off the table for now.

Why it matters: Qwen ships a new image model with concrete specs — 4.5k token input, nested-image layouts, 10px text rendering — not marketing fluff. $0.04 per 1K images is competitive, but the 1 request/minute rate limit bottlenecks batch use, capping the score.

Aug 4Tuesday

AI Chat-Group Daily (群聊日报)

Qwen 3.8 Max matches Fable 5 at 2.4T params, open weights next week

Qwen 3.8 Max launched with Terminal Bench 2.1 score 86.6 and PaperBench 93.0, beating Fable 5's 88.8. A 500-yuan token plan burned out in one day; the model lands between Luna and Terra, with price as the main draw. Open weights for both Qwen 3.8 Max and Qwen 3.8-27B drop next week. DS V4 Flash hit 8T tokens consumed in a single day, topping weekly charts—the group sees tokens becoming a commodity. On tools: an M5Stick voice dongle turns a keychain into an agent remote, LoopX keeps agent state across 200+ hours, and reverse-skill injects reverse-engineering toolchain knowledge into coding agents. The wildest methodology story: an agent autonomously downloaded a local Qwen mid-translation task and auto-installed Whisper when its API key ran out of funds.

Why it matters: Qwen 3.8 Max official release, 2.4T params matching Fable 5 with open weights coming next week — a major domestic flagship model update. The chat digest provides concrete benchmarks and real-world impressions, high information density. Deduction because the source is a group c...

AI HOT (Curated Pool)

Swiftlet runs 80B Qwen on Mac with 4.3 GB RAM, 35B on iPhone

Swiftlet rewrites MoE inference in Swift + Metal, keeping only a small dense core in memory and streaming expert weights from storage on demand. An 80B Qwen3-Next runs on Mac with 4.3 GB RAM, and a 35B model runs on iPhone. The post doesn't disclose latency or tokens per second, so I'd hold off on real-time expectations.

Why it matters: Swiftlet rewrites MoE inference in Swift + Metal, letting an 80B model run on a Mac with only 4.3 GB and a 35B model on an iPhone — the engineering path is concrete and the numbers are striking, hitting all three HKR axes. The post doesn't give latency or tokens/sec, so real-t...

Aug 3Monday

AI HOT (Curated Pool)

Qwen3.8-Max: 2.4T-parameter open-source model sets a new bar for coding and cowork

Qwen released Qwen3.8-Max, a 2.4T-parameter model (95B active) with open weights coming next week. It handled three long-horizon tasks without human help: a 16-day autonomous coding run that built a self-evolving CLI harness from scratch (265 commits, 127 PRs); a ~5-day research reproduction where it wrote 7,600 lines of code, ran 33 GPU training rounds, matched all six findings of a paper, then invented a method that beat the paper's own AIME24 score by +2.7 points; and a 24-hour contest entry that outperformed 526 human teams on Alibaba Cloud's Tianchi platform. These are self-reported results—community replication after the weight release will be the real test.

Why it matters: Qwen's first open-weight Max-class model at 2.4T total / 95B active params, demonstrated via three zero-human-intervention long-horizon tasks (16-day autonomous coding with 265 commits, 5-day paper reproduction with 7,600 lines of code) instead of benchmark tables. A Chinese f...

Jul 25Saturday

r/LocalLLaMA

Laguna S 2.1 solves a hard algorithm problem after 60k+ thinking tokens

A user tested Laguna S 2.1 on a Union-Find data rearrangement problem that took them days to solve, requiring a Julia implementation with zero dynamic allocation. Qwen 3.5-122B and 3.6-27B both failed. Laguna produced 60k+ thinking tokens and eventually wrote passing code, though it relied on packing two integers into a 64-bit value. Multiple commenters report reasoning loops when context exceeds ~50k tokens; forcing yarn-attn-factor to 1.0 helps, but tool calling remains unreliable.

Why it matters: First-person experiment with concrete problem, model comparison, and failure/success details — not a generic review. But it's a single Reddit post, not an official release or cross-source event, so authority is limited. Scored at featured threshold 72.

Jul 24Friday

Hacker News front page

Hetzner experiments with an LLM inference API

Hetzner launched an experimental, OpenAI-compatible inference API with a single model: Qwen 3.6 35B FP8. The author measured 153 ms median time-to-first-token and 224 tokens/sec output speed—fast, but the model failed simple arithmetic. There is no billing or SLA yet; Hetzner says it wants to learn about demand, scaling, and load. The more interesting angle is Hetzner's potential to turn spare GPU capacity into a low-margin inference commodity, given its cost-efficient hardware operations. The post does not confirm which GPUs power the service; Hetzner's public GPU lineup includes RTX 4000 SFF (20 GB) and RTX PRO 6000 (96 GB), leaving the hardware question open.

Why it matters: Hetzner dipping into inference is a signal: a European cloud provider moving beyond bare metal into model hosting. The author's benchmarks are solid — latency and throughput look decent — but the model choice (Qwen 3.6 35B FP8) and the arithmetic fail show this is very early-s...

Jul 23Thursday

AI HOT (Curated Pool)

Alibaba Qwen releases Qwen-Audio-3.0-TTS, tops TTS leaderboard

Alibaba Qwen dropped two TTS variants: Flash for real-time interaction and Plus for high-quality generation. The model supports fine-grained inline tags like 【whisper】 and 【angry】, natural-language style control, 16 languages, and up to 3 minutes of audio per generation. It currently ranks #1 on the Artificial Analysis TTS leaderboard. The post doesn't disclose parameter counts, latency figures, or pricing.

Why it matters: Alibaba Qwen drops a TTS model with clear Flash/Plus tiering and tops the Artificial Analysis leaderboard — a solid signal. Score held at 78 because the post doesn't disclose param count, latency numbers, or API pricing, which are the numbers builders actually need. Domestic f...

Jul 22Wednesday

AI HOT (Curated Pool)

Open models recap: Kimi K3, Qwen 3.8, distillation, and the US-China gap

Nathan Lambert and Florian Brand discuss recent open model releases. Kimi K3 dropped last week; weights are promised for July 27, but API errors are widespread—Lambert's $200 plan has been stable so far. They see big fine-tuning potential in K3, though it requires a full B300 node just to load weights. Qwen announced its next major model will be open-weight, and Xi Jinping's WAIC speech explicitly backed open source as a strategy, signaling acceleration from Chinese labs. The hosts push back on the 'how many months behind closed models' framing—benchmark gaps vary wildly, and on agentic coding tasks a few months' lag matters a lot. Distillation debates also get a critical look; they argue most takes miss the nuance.

Why it matters: Nathan Lambert's podcast recap covers concrete open-model updates: Kimi K3 weights dropping 7/27, Qwen's next flagship going open-weight, and WAIC speech signals. Downside: it's a roundup transcript, not a primary release, and some topics (distillation, open-closed gap) are on...

Jul 15Wednesday

TechCrunch · AI

Apple Intelligence approved for launch in China with Alibaba’s Qwen AI

China's Cyberspace Administration approved Apple Intelligence for launch, backed by a deal to integrate Alibaba's Qwen model into iOS, iPadOS, macOS, and visionOS. Alibaba confirmed Qwen will power text and image understanding and generation, but gave no timeline. Apple previously explored deals with Baidu, DeepSeek, and ByteDance but hit adaptation issues. The approval matters for Apple's Greater China business, which hit $20.5B in Q2. Alibaba US shares rose over 6% on the news.

Why it matters: Apple Intelligence getting CAC approval with Alibaba's Qwen is a concrete collision of US-China AI deployment rules. All three HKR axes hit: the backstory of three failed negotiations is a hook (H), the approval and failure reasons are new info (K), and it directly matters to ...

AI HOT (Curated Pool)

Alibaba's Qwen to power Apple Intelligence for users in China

Apple picked Alibaba's Qwen to run Apple Intelligence in China, covering iOS, iPadOS, macOS, and visionOS for text and image understanding plus content generation. China's cyberspace regulator just published seven mobile generative AI filings, including Apple Intelligence, Huawei's Xiaoyi, and OPPO's AndesGPT. Joe Tsai confirmed Apple talked to multiple Chinese firms before choosing Alibaba.

Why it matters: Apple picking Alibaba's Qwen for China AI is a market signal — not Baidu, not ByteDance. The CAC filing also surfaces Huawei and OPPO in the same batch, adding density. Score held below 85 because the post doesn't disclose model version or launch timeline; we're working off th...

AI HOT (Curated Pool)

China Apple Intelligence gets regulatory filing, Alibaba Qwen confirmed as partner

China's cyberspace regulator filed Apple's 'Apple Intelligence' model on July 8 for iPhone use, clearing a key hurdle for the China launch. Alibaba confirmed Qwen will be integrated into Apple Intelligence across iOS, iPadOS, macOS, and visionOS, giving users in-app access to Qwen's text and image understanding plus content generation. Alibaba co-founder Joe Tsai had already disclosed the partnership at a Dubai summit in February 2025, noting Apple talked to multiple Chinese firms before picking Alibaba. The post does not disclose a launch date or supported device list.

Why it matters: China Apple Intelligence filing + Alibaba confirming Qwen integration — two pieces of info that together mark a clear milestone. The filing removes the biggest regulatory uncertainty; Alibaba replacing Baidu as the partner is a surprise but makes sense. Score isn't higher beca...

Computing Life · Share · Yage

AI Trains AI: What a Public Self-Improvement Experiment Actually Closed the Loop On

Dan Austin open-sourced a full AI-trains-AI loop. An outer Qwen3.6 agent designs post-training recipes; an inner Qwen3-0.6B or 1.7B model runs real GPU training, and hidden eval scores feed back as reward to update the outer policy. Over 54 steps, the agent first learned to reduce invalid submissions, then shifted 1.7B model usage from 42% to 95% and began tuning temperature, optimizer, and other hyperparameters. Trained small-model scores rose from noise level into the 0.22–0.48 range, with limited transfer to a held-out triage task. A postmortem also revealed an evaluator bug: the old tool-use detector looked for `.function.name` instead of `.name`, so the 0.4-weighted score never fired—yet the reward curve still climbed. The fix required a full restart. The experiment shows outer RL can reshape a training agent's behavior, but tasks, rewards, and budgets are still human-designed.

Why it matters: Dan Austin open-sourced a full AI-trains-AI experiment: an outer Qwen3.6 agent designs post-training recipes, inner models actually train on GPU, and after 54 steps the agent learned to pick models and tune hyperparams, pushing 1.7B usage from 42% to 95%. All three HKR axes hi...

Hacker News front page

PrismML releases Bonsai 27B, the first 27B-class model that runs on a phone

PrismML compressed Qwen3.6 27B to 3.9 GB, fitting it on an iPhone 17 Pro. The ternary variant (5.9 GB) retains 95% of the full-precision baseline; the 1-bit variant (3.9 GB) retains 90%. Math and coding scores barely drop, tool calling holds up, but vision tasks degrade more noticeably. Both variants are multimodal, support 262K-token context and speculative decoding, and are released under Apache 2.0. PrismML argues this lets agentic workflows run locally, eliminating per-step API costs and keeping user data on-device.

Why it matters: PrismML compressed Qwen3.6 27B to 3.9 GB running on an iPhone 17 Pro — ternary version retains 95% capability, 1-bit retains 90%, with math and coding scores nearly intact. This is a real on-device milestone, not a paper concept. Points off for significant vision degradation, ...