{"article":{"slug":"embeddinggemma-2-an-open-lightweight-multimodal-embedding-model","title":"EmbeddingGemma 2: an open, lightweight multimodal embedding model","subtitle":null,"summary":"Google launches EmbeddingGemma 2, a 740M-parameter Apache 2.0 embedding model built on Gemma 4 that puts text, code, images, video and audio in one embedding space, with modular encoders, Matryoshka truncation, an 8K context window and on-device RAM use as low as about 191MB for text-only weights.","content_type":"announcement","language":"en","canonical_url":"https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/","author":{"name":"Sahil Dua and Henrique Schechter Vera","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Google","url":"https://blog.google/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":668,"reading_minutes":3,"published_at":"2026-10-06T16:00:00.000Z","added_at":"2026-10-06T23:15:20.736Z","updated_at":"2026-10-06T23:15:20.736Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/embeddinggemma-2-an-open-lightweight-multimodal-embedding-model","markdown_url":"https://listedarticles.com/articles/embeddinggemma-2-an-open-lightweight-multimodal-embedding-model.md","example":false,"citation":"Sahil Dua and Henrique Schechter Vera, Google. \"EmbeddingGemma 2: an open, lightweight multimodal embedding model.\" 6 Oct 2026. https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/"},"body_markdown":"# EmbeddingGemma 2: an open, lightweight multimodal embedding model\n\nWe introduced [EmbeddingGemma](https://developers.googleblog.com/en/introducing-embeddinggemma/) last year to provide a lightweight option for high-quality text embeddings, to help your apps organize, search, and connect information directly on consumer hardware. The developer community’s response blew past our expectations. With more than 20 million downloads, builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.\n\nToday, we’re launching EmbeddingGemma 2**,** expanding beyond text to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference. It can help find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model.\n\nBuilt from the same technology as Gemini Embedding models, EmbeddingGemma 2 is:\n\n- **Best-in-class for its size:** Achieves leading scores among sub-1B multimodal embedders for its size across benchmarks like MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision, and audio tasks.\n- **Modular by design:** Requires as little as 270M parameters for text-only workloads with optional vision (170M) and audio (300M) encoders for full multimodal support.\n- **Storage-efficient:** Using Matryoshka Representation Learning (MRL), developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions. This provides up to 6x storage reduction for local vector databases and memory usage.\n- **Optimized for on-device performance:** Runs efficiently within tight resource constraints. With quantization, on a Google Pixel 11 Pro, EmbeddingGemma 2 requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.\n- **Extended context ready:** Features an 8K token context window (4x larger than EmbeddingGemma 1), allowing it to process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations thereof directly on local hardware.\n\n## Achieving top-tier quality for code, vision, and audio\n\nEmbeddingGemma 2 matches the strong multilingual text performance of EmbeddingGemma while delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68), making it well-suited for local codebase indexing, semantic code search, and coding agent retrieval. Across image, video, documents, and audio, it sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size.\n\nFind full evaluation metrics and model information in the [EmbeddingGemma 2 model card](https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2).\n\n## Enabling semantic search, routing, and retrieval, fully on-device\n\nEmbeddingGemma 2 brings robust capabilities directly to edge hardware. Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline.\n\nWhen paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data. Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.\n\nTo learn how to build on-device search and RAG systems with LiteRT, read the [Google AI Edge blog post](http://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2).\n\n## Getting started with EmbeddingGemma 2\n\nWe worked closely with the following partners to ensure EmbeddingGemma 2 works immediately where you build:\n\n- **Download the models:** Find the model weights on[Hugging Face](https://huggingface.co/google/embeddinggemma-2) and[Kaggle](https://www.kaggle.com/models/google/embeddinggemma-2) , with Gemini Enterprise Agent Platform Model Garden availability coming soon. Visit[LiteRT Community on Hugging Face](https://huggingface.co/litert-community) for models optimized for on-device.\n- **On-device deployment:** Develop cross-platform apps with Google AI Edge[MediaPipe](https://developers.google.com/edge/mediapipe/solutions/decision/decision_maker) for turnkey embedding, retrieval & decision tasks or[LiteRT](https://developers.google.com/edge/litert-lm) for custom model integration. Build for the browser with transformers.js or[WebGPU](https://huggingface.co/spaces/webml-community/embeddinggemma-2-webgpu) .\n- **Use your favorite development tools** : Serve the model efficiently using transformers, sentence-transformers,[MLX](https://github.com/Blaizzy/mlx-vlm) , vLLM,[llama.cpp](https://huggingface.co/ggml-org/embeddinggemma-2-GGUF) , SGLang,[Ollama](http://ollama.com/library/embeddinggemma-2) , and LMStudio. Store your embedding vectors with[Qdrant](https://qdrant.tech/blog/embeddinggemma-2/) .\n- **Fine-tuning:** Follow guidance by[Unsloth](https://unsloth.ai/docs/models/embeddinggemma-2) for how to fine-tune EmbeddingGemma 2 for your use cases.\n\nExplore our [developer guide](https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/), [documentation](https://ai.google.dev/gemma/docs/embeddinggemma), and guides for [inference](https://ai.google.dev/gemma/docs/embeddinggemma/inference-embeddinggemma-with-sentence-transformers) and [fine-tuning](https://ai.google.dev/gemma/docs/embeddinggemma/fine-tuning-embeddinggemma-with-sentence-transformers).\n","body_html":"<h1 id=\"embeddinggemma-2-an-open-lightweight-multimodal-embedding-model\">EmbeddingGemma 2: an open, lightweight multimodal embedding model</h1>\n<p>We introduced <a href=\"https://developers.googleblog.com/en/introducing-embeddinggemma/\" rel=\"nofollow ugc noopener\">EmbeddingGemma</a> last year to provide a lightweight option for high-quality text embeddings, to help your apps organize, search, and connect information directly on consumer hardware. The developer community’s response blew past our expectations. With more than 20 million downloads, builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.</p>\n<p>Today, we’re launching EmbeddingGemma 2*<em>,*</em> expanding beyond text to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference. It can help find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model.</p>\n<p>Built from the same technology as Gemini Embedding models, EmbeddingGemma 2 is:</p>\n<ul><li><strong>Best-in-class for its size:</strong> Achieves leading scores among sub-1B multimodal embedders for its size across benchmarks like MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision, and audio tasks.</li><li><strong>Modular by design:</strong> Requires as little as 270M parameters for text-only workloads with optional vision (170M) and audio (300M) encoders for full multimodal support.</li><li><strong>Storage-efficient:</strong> Using Matryoshka Representation Learning (MRL), developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions. This provides up to 6x storage reduction for local vector databases and memory usage.</li><li><strong>Optimized for on-device performance:</strong> Runs efficiently within tight resource constraints. With quantization, on a Google Pixel 11 Pro, EmbeddingGemma 2 requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.</li><li><strong>Extended context ready:</strong> Features an 8K token context window (4x larger than EmbeddingGemma 1), allowing it to process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations thereof directly on local hardware.</li></ul>\n<h2 id=\"achieving-top-tier-quality-for-code-vision-and-audio\">Achieving top-tier quality for code, vision, and audio</h2>\n<p>EmbeddingGemma 2 matches the strong multilingual text performance of EmbeddingGemma while delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68), making it well-suited for local codebase indexing, semantic code search, and coding agent retrieval. Across image, video, documents, and audio, it sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size.</p>\n<p>Find full evaluation metrics and model information in the <a href=\"https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2\" rel=\"nofollow ugc noopener\">EmbeddingGemma 2 model card</a>.</p>\n<h2 id=\"enabling-semantic-search-routing-and-retrieval-fully-on-device\">Enabling semantic search, routing, and retrieval, fully on-device</h2>\n<p>EmbeddingGemma 2 brings robust capabilities directly to edge hardware. Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline.</p>\n<p>When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data. Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.</p>\n<p>To learn how to build on-device search and RAG systems with LiteRT, read the <a href=\"http://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2\" rel=\"nofollow ugc noopener\">Google AI Edge blog post</a>.</p>\n<h2 id=\"getting-started-with-embeddinggemma-2\">Getting started with EmbeddingGemma 2</h2>\n<p>We worked closely with the following partners to ensure EmbeddingGemma 2 works immediately where you build:</p>\n<ul><li><strong>Download the models:</strong> Find the model weights on<a href=\"https://huggingface.co/google/embeddinggemma-2\" rel=\"nofollow ugc noopener\">Hugging Face</a> and<a href=\"https://www.kaggle.com/models/google/embeddinggemma-2\" rel=\"nofollow ugc noopener\">Kaggle</a> , with Gemini Enterprise Agent Platform Model Garden availability coming soon. Visit<a href=\"https://huggingface.co/litert-community\" rel=\"nofollow ugc noopener\">LiteRT Community on Hugging Face</a> for models optimized for on-device.</li><li><strong>On-device deployment:</strong> Develop cross-platform apps with Google AI Edge<a href=\"https://developers.google.com/edge/mediapipe/solutions/decision/decision_maker\" rel=\"nofollow ugc noopener\">MediaPipe</a> for turnkey embedding, retrieval &amp; decision tasks or<a href=\"https://developers.google.com/edge/litert-lm\" rel=\"nofollow ugc noopener\">LiteRT</a> for custom model integration. Build for the browser with transformers.js or<a href=\"https://huggingface.co/spaces/webml-community/embeddinggemma-2-webgpu\" rel=\"nofollow ugc noopener\">WebGPU</a> .</li><li><strong>Use your favorite development tools</strong> : Serve the model efficiently using transformers, sentence-transformers,<a href=\"https://github.com/Blaizzy/mlx-vlm\" rel=\"nofollow ugc noopener\">MLX</a> , vLLM,<a href=\"https://huggingface.co/ggml-org/embeddinggemma-2-GGUF\" rel=\"nofollow ugc noopener\">llama.cpp</a> , SGLang,<a href=\"http://ollama.com/library/embeddinggemma-2\" rel=\"nofollow ugc noopener\">Ollama</a> , and LMStudio. Store your embedding vectors with<a href=\"https://qdrant.tech/blog/embeddinggemma-2/\" rel=\"nofollow ugc noopener\">Qdrant</a> .</li><li><strong>Fine-tuning:</strong> Follow guidance by<a href=\"https://unsloth.ai/docs/models/embeddinggemma-2\" rel=\"nofollow ugc noopener\">Unsloth</a> for how to fine-tune EmbeddingGemma 2 for your use cases.</li></ul>\n<p>Explore our <a href=\"https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/\" rel=\"nofollow ugc noopener\">developer guide</a>, <a href=\"https://ai.google.dev/gemma/docs/embeddinggemma\" rel=\"nofollow ugc noopener\">documentation</a>, and guides for <a href=\"https://ai.google.dev/gemma/docs/embeddinggemma/inference-embeddinggemma-with-sentence-transformers\" rel=\"nofollow ugc noopener\">inference</a> and <a href=\"https://ai.google.dev/gemma/docs/embeddinggemma/fine-tuning-embeddinggemma-with-sentence-transformers\" rel=\"nofollow ugc noopener\">fine-tuning</a>.</p>","headings":[{"level":1,"text":"EmbeddingGemma 2: an open, lightweight multimodal embedding model","id":"embeddinggemma-2-an-open-lightweight-multimodal-embedding-model"},{"level":2,"text":"Achieving top-tier quality for code, vision, and audio","id":"achieving-top-tier-quality-for-code-vision-and-audio"},{"level":2,"text":"Enabling semantic search, routing, and retrieval, fully on-device","id":"enabling-semantic-search-routing-and-retrieval-fully-on-device"},{"level":2,"text":"Getting started with EmbeddingGemma 2","id":"getting-started-with-embeddinggemma-2"}]}}