Skip to content
AI HOT (Curated Pool)

Qwen launches Qwen3.8-Omni-Flash, a native omnimodal model built for audio-visual agent workflows

Qwen 发布原生全模态模型 Qwen3.8-Omni-Flash,主打音视频智能体任务交付

Qwen3.8-Omni-Flash is a native omnimodal model that shifts focus from audio-visual understanding to task planning, tool use, and delivery in real-world workflows. It supports a 1M-token context window, with average scores across 29 evals up over 25% vs Qwen3.5-Omni-Plus. API pricing for audio input dropped over 98%, and audio-visual input over 93%. It gained 36.5 points on WildClawBench-MM and 22.3 on AgenticVBench; AliMeeting DER fell from 88.11 to 3.35. Qwen claims overall audio performance exceeds Gemini 3.8 Flash, with audio-visual performance close to it. Qwen-Live Harness is open-sourced for real-time interaction, and Qwen-MM-Plugins now supports tool use and workflows for long-form audio and video.

Why it matters: Qwen pushes omnimodal models from understanding to task delivery, backed by concrete benchmarks and pricing. Score stays at 82 rather than higher because it's a launch-day post with no third-party validation or cross-source cluster yet.

Read the original ↗Export Markdown