Skip to content
Synced · WeChat

Remember more, answer faster, use less: HERMES speeds real-time streaming video understanding by 10x

记得住、答得快、用得省:HERMES 让流式视频理解实时响应提速10倍

Fudan University, Shanghai Academy of AI for Science, and NUS proposed HERMES, a training-free framework that turns KV cache into hierarchical memory for streaming video understanding and cuts TTFT by up to 10x. The post lists three mechanisms: hierarchical cache management, cross-layer memory smoothing, and position re-indexing; it reports 68% fewer video tokens with comparable or better results, and Qwen2.5-VL-7B on StreamingBench rising from 73.31% to 79.44%. What matters for practitioners: it answers without external retrieval, with TTFT around 27/29/28 ms at 16/64/256 frames.

Why it matters: Strong HKR-H/K/R: the 10x speed claim is a real hook, and the article includes concrete mechanisms and numbers, including 68% fewer video tokens and 27-29 ms TTFT. It stays below major product-news bands because this is an academic research release, not a market-moving launch.

Read the original ↗Export Markdown