Google has released EmbeddingGemma 2, an open-weight multimodal embedding model it says is designed to bring semantic search across text, code, images, video and audio to laptops, phones and other edge devices — without sending data to a server.
Built on the Gemma 4 architecture and released under the Apache 2.0 licence, the model is deliberately small: sub-1B parameters, with a modular design that lets developers load only the encoders they need. The text and code backbone is 270M parameters; adding vision brings it to 440M, adding audio to 570M, and the full multimodal configuration totals 740M. Crucially, every configuration projects into the same 768-dimensional vector space, so a query embedded by the text-only setup can be matched directly against documents embedded by the full model.
That shared space is what makes cross-modal retrieval practical: a single text query can be compared with an image, a video frame or an audio clip by semantic similarity. Embeddings are produced through the sentence-transformers library, and Google says the same checkpoint underpins all four setups — meaning a text-only index can later be extended with image or audio embeddings without recomputing what has already been stored.
Google also leaned on Matryoshka Representation Learning, which lets vectors be truncated from 768 dimensions down to 512, 256 or 128 while retaining much of their quality. At 256 dimensions, Google says text and code keep most of their original quality and image, video and speech retrieval retain roughly 95 percent, at a third of the storage. At 128 dimensions storage falls by 6x — around 250 MB for a million vectors instead of 1.5 GB — with text and code retaining about 90 percent of quality, though multimedia retrieval drops to roughly 75 percent.
Google says EmbeddingGemma 2 scores 14 percent higher than the original EmbeddingGemma on the MTEB code benchmark and keeps its predecessor's multilingual text accuracy while adding image, video and audio retrieval. The company is pitching it for on-device uses such as search-as-you-type media retrieval, finding moments inside video, zero-shot intent routing and local codebase indexing for coding agents. Weights are available on Hugging Face, and Google points developers to MediaPipe Tasks and LiteRT for CPU, GPU and NPU acceleration, alongside vLLM, Ollama, LM Studio, MLX and SGLang for local serving.




