Apple open-sources LensVLM-9B: compress long context as images, expand only relevant pages
LensVLM: Compressing long context as images, expanding only relevant pages
Apple released LensVLM-9B on Hugging Face, a 9B-parameter vision-language model. The core idea: compress long documents into images, then expand only the relevant pages on demand instead of stuffing everything into the context window. The model card is the only source right now—training data, benchmarks, and compression ratios aren't disclosed, so I'd hold off until more details land.