Slop TVNewsLatest
Business

EmbeddingGemma 2 puts video search on a phone at 740M parameters

Google's open model unifies text, images, audio and video in one vector space, so a voice memo can find the matching shot without leaving the phone.

Illustration: EmbeddingGemma 2 puts video search on a phone at 740M parameters

Key takeaways

  • Google released EmbeddingGemma 2 on October 6, 2026, a 740-million-parameter open multimodal embedding model released under the Apache 2.0 licence.
  • It maps text, code, images, video frames and audio into one 768-dimension vector space; the first EmbeddingGemma, from September 2025, handled text only.
  • Google's model card reports a video retrieval score of 50.67 Hit@1 on MMEB v2 and an MTEB Code score of 78.68, up 9.92 points from the first version's 68.76.
  • On a Google Pixel 11 Pro with quantization, Google reports about 191MB of active RAM for the text-only weights and about 567MB for the full multimodal model.
  • Google's AI Edge Gallery app adds a Video Moments Finder that locates timestamps in a local video from a typed or spoken description.

Google released EmbeddingGemma 2 on October 6, a 740-million-parameter open model that maps text, code, images, video and audio into one vector space and runs on a phone, so a local app can find the shot a voice memo describes without a network call.

Sahil Dua and Henrique Schechter Vera, research engineers at Google DeepMind, wrote in the launch post that the first EmbeddingGemma "blew past our expectations," passing more than 20 million downloads since it shipped in September 2025 at 308 million parameters. That version held text only. Version 2 adds images, audio and video frames to the same space, built on the Gemma 4 architecture.

The model is modular. A 270-million-parameter text and code backbone does the base work, with a 170-million-parameter vision encoder and a 300-million-parameter audio encoder that developers can leave unloaded. Run text only after quantization, the weights need about 191MB of active RAM on a Google Pixel 11 Pro; the full multimodal model needs about 567MB, per the AI Edge post. The weights ship under Apache 2.0, which permits commercial use.

Running on a smartphone means LiteRT and MediaPipe on the device, a browser path through transformers.js and WebGPU, and, in the coming weeks, a service on Android through ML Kit with NPU acceleration where the hardware has it. Google measured vision encoding at 37.3 milliseconds per image, or 26.9 images a second, on a MacBook M5 Pro GPU in its AI Edge write-up. The model card lists an 8,192-token context window, four times the first version's, enough for about 5.5 minutes of audio, 29 images or 58 video frames in one pass, and a Matryoshka trick that shrinks each 768-number embedding to 256 or 128 dimensions, cutting vector storage up to six-fold.

Google's benchmark headline is code. On the code section of the Massive Text Embedding Benchmark, EmbeddingGemma 2 scored 78.68 against the first version's 68.76, a 9.92-point gain, while multilingual text barely moved, 61.36 against 61.15. The video number is the honest one: 50.67 Hit@1 on MMEB v2 video retrieval, the lowest of its retrieval scores. The figures are Google's own, reported with no independent validation, and SiliconANGLE notes the company claims leading results only among multimodal embedders under 1 billion parameters. The published figures have been repeated by Unite.AI and Jetstream, not independently reproduced.

It is an embedding model. It turns content into vectors so software can compare them; it does not generate a pixel. For a video tool that is still plumbing for jobs a creator does by hand today: searching a library by description rather than filename, finding the take where a line landed, matching an audio note to the shot it was about, flagging near-duplicate clips before they bloat a project. Google showed the pattern in its AI Edge Gallery app, which adds a Video Moments Finder that indexes frames and audio from a local clip and jumps to timestamps matching a typed or spoken query such as "kids laughing," plus an Instant Media Search that re-ranks images and video as you type.

The limits are real. One pass holds 58 frames, so indexing a long video means chunking it and storing many embeddings. At 128 dimensions, quality "degrades multimodal quality substantially," the model card warns, so the six-fold saving is a text-only trade. Holding most quality for image, video and speech retrieval means keeping vectors at 256 dimensions or wider, the developer guide states.

For someone who makes AI video, the change is smaller than a new generator and more useful than one. The cheapest way to organise a growing pile of clips, references and renders stops being a cloud service with a per-call bill and becomes a 740-million-parameter model sitting on the machine in your hand, working with no connection at all.

EmbeddingGemma 2 weights are on Hugging Face and Kaggle under Apache 2.0 now; ML Kit support for Android arrives in the coming weeks.

Sources

  1. blog.google - Google's launch post; 740M parameters, Apache 2.0, Pixel 11 Pro RAM, 8K context, 20M+ v1 downloads
  2. ai.google.dev - the model card; architecture, full benchmark table, MRL truncation, limitations
  3. developers.googleblog.com - developer guide; modular encoder footprints and the 256-dimension quality claim
  4. developers.googleblog.com - AI Edge post; Instant Media Search, Video Moments Finder, 37.3ms vision on M5 Pro GPU, ML Kit plans
  5. developers.googleblog.com - the September 2025 EmbeddingGemma 1 post; 308M text-only baseline and EdgeTPU latency
  6. huggingface.co - hosted weights and the vector-truncation table confirming the modality scores
  7. siliconangle.com - independent coverage; sub-1B comparison and encoder breakdown
  8. unite.ai - independent summary; token budgets per modality
  9. jetstream.blog - independent coverage; context and storage figures