inclusionAI releases VISTA-4B, a vision-language model for GUI element grounding
inclusionAI 发布 VISTA-4B GUI 定位视觉语言模型
inclusionAI open-sourced VISTA-4B on Hugging Face, a 4B-parameter vision-language model built on Qwen3.5. It focuses on GUI grounding: given a screenshot and a text instruction, the model pinpoints the target button or region. The model card lists gui-grounding and reinforcement-learning tags, indicating RL was used to improve localization accuracy. Code examples cover Transformers, vLLM, and SGLang, under an Apache 2.0 license. The post doesn't disclose benchmark scores, training data size, or inference latency—I'd hold off on performance claims until those numbers surface.
Why it matters: A 4B GUI grounding model is a practical direction and RL training is a real technical signal, but the model card has zero benchmarks, no training data disclosure, and no comparison to OmniParser or UI-TARS. Too many gaps to push higher.