Skip to content
Trending storyPast story

Apple open-sources LensVLM-9B: compress long context as images, expand only relevant pages

1 report1 sourceupdated 6 days ago

What happened

Summary

苹果在 Hugging Face 上放出了一个 9B 参数的视觉语言模型 LensVLM-9B。它的思路是把长文档先压缩成图片,需要时再按需展开相关页面,而不是一股脑把整本书塞进上下文窗口。目前只有模型卡,训练数据、跑分、压缩比这些关键信息都没披露,我会先打个折,等更多细节出来再看。

Coverage

Follow the reports to see the story from different sides.

Sep 24
  1. Hacker News front page
    Apple open-sources LensVLM-9B: compress long context as images, expand only relevant pages

    Apple released LensVLM-9B on Hugging Face, a 9B-parameter vision-language model. The core idea: compress long documents into images, then expand only the relevant pages on demand instead of stuffing everything into the context window. The model card is the only source right now—training data, benchmarks, and compression ratios aren't disclosed, so I'd hold off until more details land.