Google has turned the task of "finding things locally" into a small model that fits in a phone. Its latest release, EmbeddingGemma2, is the first native multimodal open-source embedding model from Google, with a parameter size of 740M, which can map text, code, images, videos, and audio into the same semantic space all at once, making cross-modal retrieval run so smoothly on the edge for the first time.

The most obvious selling point is its small size and offline capability. The complete multimodal model takes up less than 600MB of memory, meaning it can run directly on phones, laptops, and browsers without uploading any content to the cloud — your photo album, recordings, and videos are searched right on your device, naturally protecting your privacy. The context window is set to 8K, four times that of the previous generation, allowing it to process longer materials before matching; language coverage has also been expanded to over 100 languages, removing barriers to searching across different languages.

In terms of performance, Google claims it significantly outperforms open-source competitors of the same scale in image, video, and code retrieval tasks. Even more thoughtful is its modular design: parts you don't need can be loaded as required, reducing the index size and saving both space and computing power. In real-world scenarios, it can perform specific tasks — find a particular photo or media in a local photo album, directly locate a segment in a long video, perform offline searches on scattered files, or search for functions and snippets in a local codebase by intent, without having to push the repository to a remote service for checking.

When a multimodal embedding model is small enough to fit in a phone and open-source enough for anyone to modify, the focus of AI retrieval is quietly shifting from "giving data to the cloud" back to "keeping the capabilities on the device." Google's move has opened a gap in edge-side private search.