AI News Feed
Market watch
Large Language Models

Google releases EmbeddingGemma 2, a multimodal embedding model small enough for phones

Google released EmbeddingGemma 2, an open multimodal embedding model that puts text, images, audio and video in one space and runs on a smartphone. Weights are available under Apache 2.0.

The new model places images, audio and video in the same embedding space as text. According to SiliconANGLE, an app built on it could take a voice memo and locate the matching moment in a video without the data leaving the phone.

Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera wrote in the announcement that the response to the original "blew past our expectations." By their count, developers have downloaded it more than 20 million times.

Built on the Gemma 4 architecture Google released in April, the new version is more than twice the size of the original at 740 million parameters. Most of the growth sits in the vision and audio encoders, which apps that work only with text can leave off. The 270 million-parameter text core used about 191 megabytes of memory when Google tested a quantized build on a Google Pixel 11 Pro.

An app also has to store what the model produces. Each embedding is a list of 768 numbers, and every photo, clip or document it indexes adds one more entry to a local vector database. A training technique called Matryoshka Representation Learning lets developers cut those lists to as few as 128 numbers, reducing the space they take up by as much as six times. At 256 numbers, Google's developer guide says, image, video and speech retrieval keep about 95% of their full quality.

Code showed the biggest benchmark gain. EmbeddingGemma 2 scored 78.68 on the code section of the Massive Text Embedding Benchmark, almost 10 points above the first version, a result Google is pitching at developers who build retrieval for coding agents. Multilingual text scores barely moved.

The company also claims leading results among multimodal embedding models under 1 billion parameters, and said the model outperforms some specialist models more than twice its size on image, video and audio tasks.

Because EmbeddingGemma 2 shares a text tokenizer and an audio encoder with Gemma 4, an on-device retrieval-augmented generation setup running both needs less memory than two unrelated models would. Google's AI Edge Foresight meeting app for Mac already runs the pair together. In Google's AI Edge Gallery demo app, a Video Moments Finder feature locates a scene inside a video from a typed or spoken query.

Model weights are available now from Hugging Face and Google's Kaggle under an Apache 2.0 license that permits commercial use. Google said the model will reach the Model Garden in the Gemini Enterprise Agent Platform soon, and the weights already work with open-source serving tools such as vLLM, llama.cpp and Ollama.