Skip to content
Google DeepMind

Google DeepMind releases EmbeddingGemma 2 for on-device multimodal embeddings

EmbeddingGemma 2: an open, lightweight multimodal embedding model

Google DeepMind released EmbeddingGemma 2, built on the Gemma 4 architecture, which maps text, code, images, video and audio into one shared embedding space. The model has 740 million parameters, ships under the Apache 2.0 license, and supports an 8K token context window for fully offline cross-modal search and RAG. Text-only tasks need just 270M parameters, and output vectors can be cut from 768 dimensions to 128.

Why it matters: EmbeddingGemma 2 runs cross-modal retrieval on local devices, so text, image, audio and video search never has to send data to the cloud.

Read the original ↗Export Markdown