On October 6, 2026, Google DeepMind released EmbeddingGemma 2, the company's first native multimodal open-source embedding model, designed for local on-device operation. The model has 740 million parameters and unifies five modalities—text, code, images, video frames, and audio—into the same 768-dimensional vector space. It is publicly released under the Apache 2.0 license and is available for download and use on Hugging Face and Kaggle starting today.
EmbeddingGemma 2 is built on the Gemma 4 base and uses an independent modular encoder architecture. The total of 740 million parameters is made up of three independent parts: a 270 million-parameter text model (including a 130 million-parameter Transformer backbone and a 140 million-parameter embedding layer), a 170 million-parameter vision encoder, and a 300 million-parameter audio encoder. The three encoders are independent but share the same 768-dimensional output space.
During actual deployment, developers can load only the parts they need: pure text and code scenarios require loading only 270 million parameters, text plus vision requires 440 million, text plus audio requires 570 million, and full multimodality requires the complete 740 million. On Google Pixel 11 Pro, text mode requires only about 191 MB of active memory, while the full multimodal configuration requires about 567 MB.
This design compresses the previous workflow—which required chaining three independent pipelines: an image captioning model, a speech-to-text model, and a text embedding model—into a shared vector space. Developers only need to call one model, and queries and content are compared directly for cosine similarity in the same space.
How Much Content Can a Single 8192-Token Window Hold
The model's context window is 8192 tokens, four times that of the first-generation EmbeddingGemma. Different modalities have clearly defined fixed token consumption rates: each image consumes 280 tokens, each video frame consumes 140 tokens, and each second of audio consumes 25 tokens. By this conversion, a single input can hold up to 29 images, 58 video frames, or 5.5 minutes of audio. Inputs can be interleaved: text, images, and audio can coexist, with image positions marked by placeholders.
The expansion of context capacity allows operations such as "using a voice memo to retrieve a specific video clip" to be completed in a single inference pass, without being split into multiple steps.
The Underestimated Zero-Shot Routing Capability
EmbeddingGemma 2 can directly match user input against category labels and descriptions, without any training data or fine-tuning, completing zero-shot intent routing within millisecond-level latency. This enables it to act as an ultra-low-latency on-device decision engine, determining which intent category a user input belongs to and then triggering the corresponding local workflow, all without going through the cloud.
Benchmark Numbers: Where It Wins and Where It Falls Short
According to Google's official model card data, EmbeddingGemma 2's MTEB Code score for code retrieval improved from 68.76 in the first generation to 78.68, an increase of 9.92 points. On multilingual text, its MTEB Multilingual v2 score is 61.36, on par with the first-generation level. Scores for the other modalities are: MIEB lite 64.64, MMEB v2 image retrieval 57.28, visual document retrieval 67.84, video retrieval 50.67, MSEB sound-effect retrieval 69.54, and MAEB audio tasks 49.39.
The model reaches a leading level among multimodal embedders with under 1 billion parameters and surpasses specialized models more than twice its size on some task-specific benchmarks. Among competitors, Alibaba's Qwen3-VL-Embedding-8B reaches 77.8 on MMEB-V2, but has about 8 billion parameters, 11 times that of EmbeddingGemma 2.
Video retrieval's 50.67 is the lowest score among the current evaluations, and the two audio scores are widely spread (MSEB 69.54 versus MAEB 49.39), indicating a performance gap across different audio task types.
The Previous Generation's 20 Million Download Base
The first-generation EmbeddingGemma, released in 2025, has surpassed 20 million downloads, with primary use cases concentrated in on-device search tools and privacy-focused retrieval-augmented generation (RAG) pipelines. The first generation supported only text embeddings. This upgrade to multimodality expands on a download base that has already proven market demand.
EmbeddingGemma has already established some developer inertia in the on-device text embedding space, and EmbeddingGemma 2 provides multimodal capabilities directly in the same checkpoint, without changing interface habits, reducing migration costs.
The Next Few Weeks: ML Kit and NPU Acceleration
The Google AI Edge team disclosed that in the coming weeks it will offer EmbeddingGemma 2 as a service to Android developers through ML Kit, with NPU acceleration enabled, covering a broader range of device models. NPU acceleration means inference power consumption and latency will fall further, which is especially important for phone devices with limited battery capacity.
ML Kit is one of the most commonly used machine learning integration frameworks for Android developers, and it does not require developers to manually manage model deployment. Once EmbeddingGemma 2 lands as an ML Kit service, the target audience will expand from machine learning engineers to the broader Android app developer community.
Verdict
The core problem EmbeddingGemma 2 solves is making on-device multimodal retrieval engineering-feasible. Previously, implementing joint local search over text, images, and audio on a phone required maintaining multiple models and coordinating multiple inference pipelines, and the development complexity far exceeded what most app teams could handle. EmbeddingGemma 2 compresses this complexity into a 567 MB model file and an Apache 2.0 license.
EmbeddingGemma 1's 20 million downloads show that on-device embedding demand genuinely exists. EmbeddingGemma 2 extends that demand from text to five modalities. The relatively low video retrieval score and uneven audio task performance suggest that this unified vector space still comes at a cost when processing temporally intensive content.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接