Stanford and NVIDIA introduce Contrastive Language Models, up to 9× faster than Jev for decision-making
What happened
CLM 的思路和传统语言模型不一样:它不生成文字,而是把“当前状态”和“可选动作”分别编码成向量,用余弦相似度打分,选最匹配的动作。这相当于让模型直接做选择题,省掉了逐字生成的时间。CLM-8B 在电脑操作、游戏和工具调用上效果跟 Jev 持平,但延迟最高能降 9 倍——候选动作越多、动作能跨状态复用时,提速越明显。在 DeepSWE 和 Termin...
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pagePickStanford and NVIDIA introduce Contrastive Language Models, up to 9× faster than Jev for decision-making
CLM encodes states and actions separately and scores pairs via cosine similarity instead of generating tokens. CLM-8B matches Jev on computer-use, gaming, and tool-calling while cutting latency by up to 9×. With light fine-tuning it hits 81.6% on DeepSWE and 87.6% on Terminal Bench 2.1, running 4–6× faster than Jev. Only the 20M-parameter projection head is trained; the frozen LLM backbone keeps pre-training to about one hour on a single RTX 4090. The post does not disclose whether weights are open or if sizes beyond 8B are planned.