Skip to content
AI HOT (Curated Pool)

JD open-sources JoyAI-VL-Interaction, a full-stack interactive vision-language model

京东全栈开源JoyAI-VL-Interaction,从"一问一答"走向"边看边说"

JD released JoyAI-VL-Interaction, the first fully open-source interactive model with native vLLM-Omni support. It watches a live video stream, decides when something matters, and speaks up in real time—complex tasks can be handed off to a backend agent. In a 58-person blind test, it beat Doubao's video call assistant 77.6% of the time and Gemini's 87.9%, hitting 100% on surveillance alerts. The release includes weights, the interaction dataset, training recipes, and a deployable system that takes camera or live-stream input with voice interaction and long-term memory, targeting security, elder care, and live commentary.

Why it matters: JD.com open-sourced a real-time video interaction model that decides when to speak, with strong blind-test numbers and 100% win rate in surveillance. Not scoring higher because JD isn't a tier-1 model lab yet and community adoption is unproven, but the open-source commitment a...

Read the original ↗Export Markdown