EmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings
Google introduced an open embedding model designed to represent text, code, images, video and audio in a shared space.
TL;DR
- Google announced EmbeddingGemma 2, an open multimodal embedding model.
- The model is designed to map text, code, images, video and audio into a shared embedding space.
- The release extends Google’s on-device AI portfolio, while the cited materials do not establish independent benchmark results.
Google introduced EmbeddingGemma 2 as an open model for multimodal embeddings, covering text, code, images, video and audio. [1]
The Next Web separately reported the release as a 740-million-parameter model intended for on-device use; that parameter count is attributed to its report. [2]
Why it matters
The release adds a multimodal retrieval component to Google’s edge AI offerings. Its practical reach will depend on deployment constraints and independent evaluation, which the launch materials do not settle.
Editor's note
The headline is copied verbatim from Google’s announcement. Parameter count and on-device framing are attributed to The Next Web; no performance claim is inferred.