Skip to content
AI HOT (Curated Pool)

UEmbed: One decoder-only model that outputs both sparse and dense multimodal embeddings

UEmbed:统一稀疏与稠密的多模态嵌入模型

UEmbed is a decoder-only multimodal embedding model that produces both sparse lexical vectors and dense semantic vectors in a single causal forward pass. It appends N learnable special tokens and partitions the vocabulary into N disjoint subsets; each token predicts sparse weights over its assigned subset, and the N subsets are concatenated into the full sparse vector. The authors release UEmbed at 2B, 4B, and 9B scales, all trained on public data. UEmbed-9B hits 71.8 (dense) and 71.0 (sparse) on MMEB-v2, outperforming RzenEmbed, and stays competitive with strong baselines on BEIR. The paper also demonstrates utility across effectiveness, efficiency, and agentic applications. The post doesn't disclose inference latency or memory footprint, so real-world cost is still an open question.

Why it matters: A single model that outputs both sparse and dense vectors is a real engineering improvement for search and RAG. The 9B version shows benchmark wins, but without deployment cases or product plans, it stays at 'worth recommending' rather than 'must-write.'

Read the original ↗Export Markdown