AIARTICLE

EmbeddingGemma 2 brings multimodal semantic search to phones without heavy GPU

Google DeepMind launched EmbeddingGemma 2, an open model with 740 million parameters that combines text, image, video, and audio in a single vector space and runs with less than 600MB of RAM, without relying on the cloud.

What is EmbeddingGemma 2

Google DeepMind launched EmbeddingGemma 2 this Tuesday (10/6), a multimodal, open-weight embedding model that maps text, image, video frame, and audio into a single vector space. Unlike pipelines that chain an image-captioning model, an audio-transcription model, and a third text-embedding model, it does all three with a single encoder. For developers building local search and media retrieval, this cuts latency and memory use right at the point where these pipelines usually get stuck.

The announcement was published on the Google Developers Blog, jointly by the Google AI Edge and Android ML teams. The model has 740 million parameters and ships with modular encoders: loading text only consumes around 191MB of active RAM; the full multimodal version (text, image, and audio) sits at around 567MB on a Pixel 11 Pro. That's RAM, not disk storage: it can run in the same process as a regular app without blowing the device's memory budget.

One vector space for everything, no fine-tuning

The most technically interesting piece isn't similarity search itself, it's using EmbeddingGemma 2 as a zero-shot decision engine. Without any training data or fine-tuning, the model compares the user's input directly against classification labels and descriptions, returning intent routing in milliseconds. In practice, this replaces a classifier trained specifically for each app with a generic encoder that already understands the task from the category's own description.

Google demonstrates this with the new MediaPipe Decision Task, which runs in a chess game evaluating 500 options per round in under 100ms, entirely on-device. It's a toy example, but the usage pattern is serious: voice command routing, content triage, or deciding which flow of an app to trigger from a free-form sentence, without training anything beforehand.

The fastest way to see the model in action is through the two apps Google updated today. Google AI Edge Gallery, available for Android and iOS, got two new demos: Instant Media Search and Video Moments Finder.

  • Instant Media Search converts the query and the device's photos/videos into vectors, stores the vectors in a local SQLite database, and ranks them by cosine similarity. The search updates with every keystroke: typing "Katze" (cat, in German) already brings up photos of felines, and completing it to "Katze schläft auf Tastatur" (cat sleeping on the keyboard) reorders results and brings the right photo to the top, in real time, with no round-trip to the cloud.
  • Video Moments Finder indexes video frames and audio snippets locally and allows searches like "children laughing" or "dog catching the frisbee," highlighting the exact timestamp of the segment, without transcribing the audio or generating intermediate captions.

On Mac, Google launched Google AI Edge Foresight, an experimental meeting companion app that runs 100% locally: it enriches handwritten notes with details captured from the conversation in real time and lets you search in natural language across images, documents, transcripts, and notes, all processed on the Mac itself, with no cloud subscription and without the meeting's audio ever leaving the device.

For developers: ML Kit, MediaPipe, and LiteRT

For those building production Android apps, Google confirmed that EmbeddingGemma 2 is coming to ML Kit in the coming weeks, with NPU acceleration where available and automatic model updates, without the app having to bundle the weights and bloat the APK.

For those who need cross-platform support (iOS, macOS, Windows, Linux, Web), the path is MediaPipe Tasks, which got two new components:

  • Universal Embedder: takes raw image or text input and returns normalized 768-dimension vectors (or truncated to 128-512 via Matryoshka Representation Learning), handling resizing, tensor normalization, and multimodal tokenization with no extra code.
  • Semantic Retriever: the new interface for Google's RAG SDK, which indexes embeddings on-device and runs approximate nearest neighbor (ANN) search with single-digit millisecond returns.

Those who need fine-grained control over hardware acceleration go straight to LiteRT, the engine that already runs underneath ML Kit, MediaPipe, and the two demo apps. The model is distributed as a single .litertlm file that runs on CPU, GPU, or NPU without recompiling anything per platform.

The numbers Google released

The one concrete performance figure Google published is for visual embedding on a MacBook M5 Pro: 37.3 milliseconds per image using the GPU, equivalent to 26.9 images per second, with a maximum budget of 70 visual tokens per image. The post promises a table comparing CPU, GPU, and NPU on other devices, but the detailed per-device benchmark is in the model card on Hugging Face, not in the post itself.

Two optimizations explain this performance. The first is Quantization-Aware Training, which compresses the weights down to INT4 and INT8, bringing multimodal vector search to mid-range hardware. The second is output truncation via Matryoshka Representation Learning: the same 768-dimension vector can be cut in real time, reducing the space taken up by the local index by up to 8 times, at the cost of some search precision.

What this replaces, and where it still falls short

In architectural terms, EmbeddingGemma 2 replaces the common practice of chaining three specialized models with a single modular encoder that's already born in the same vector space. This eliminates the cost of keeping three models loaded at once, and avoids the mistake of comparing embeddings from different vector spaces, a classic source of silent bugs in semantic search.

For developers in Brazil, the most practical gain may not even be latency: it's not depending on a cloud embedding API call for every search, which matters both for cost and for scenarios with unstable connectivity. Running locally also simplifies the conversation around LGPD (Brazil's data protection law), since users' photos, audio, and transcripts never leave the device.

Where this doesn't solve everything: the model serves search and routing, not text generation or open-ended answers, so it remains a piece of a search or RAG system, not a replacement for a full LLM. Also, ML Kit integration with NPU acceleration is still "coming in the next few weeks" according to the post itself, so anyone who wants production on Android today depends on direct integration via LiteRT or MediaPipe Tasks, without the automatic model management layer promised later.

To start testing, the most direct path is downloading the pre-quantized .litertlm packages straight from the LiteRT community on Hugging Face, or following the guide on GitHub, which includes an interactive Colab and access to Google Cloud's Developer Device Platform, with more than 200 physical devices for running benchmarks before deciding whether the model performs well on the hardware your users actually have.

Translated from the Brazilian Portuguese original · Read the original

View profile →