The Google Developers Blog said Google DeepMind launched EmbeddingGemma 2 on 6 October 2026, a 740M-parameter model that maps text, images, video frames and audio into one vector space.
An embedding is a numerical description used to compare pieces of content. Google’s Instant Media Search demonstration converts both a query and local media into vectors, stores the media vectors in a local SQLite database and ranks matches by their similarity. The company says the model uses approximately 191MB of active RAM for text-only weights, or 567MB for the full multimodal model, on a Google Pixel 11 Pro.
Finding a particular photo in a collection could therefore begin with a description rather than a filename. In Google’s demonstration, the comparison happens on the device without an internet connection, and results update as the description is typed.
For developers, Google says MediaPipe Tasks support extends deployment across iOS, macOS, Windows, Linux and the web. It also says the model can match inputs against classification labels without training data or fine-tuning. In a chess example, Google reports that its MediaPipe Decision Task evaluates 500 options per turn in less than 100ms.