All articles
LLM API Comparison & Cost/

EmbeddingGemma 2 Benchmarks and Features

Multimodal retrieval, MTEB benchmarks, float16 pitfalls, and hardware resource tradeoffs.

4 minutes read

EmbeddingGemma 2 Benchmarks and Features. A macro close-up of a high-end GPU's gold-plated silicon die reflecting iridescent light, partially obscured by a tangle of neon-blue fiber optic cables and the blurred, metallic fins of a server — GPUniq
Data as of 2 sources

Drafted with AI tools and edited by a human. The figures and conclusions were checked by a GPUniq editor.

TL;DR

EmbeddingGemma 2 is an open-source multimodal embedding model from Google DeepMind. It maps text, code, static images, video, and audio into a shared 768-dimensional vector space within an 8,192-token context window. Released under the Apache 2.0 license, it scores 78.68 NDCG@10 on MTEB code retrieval, up from EmbeddingGemma 1's 68.76 score. It supports Matryoshka vector truncation down to 128 dimensions, though visual and audio tasks lose quality at the smallest size. Deployments must run on bfloat16 or float32 precision because float16 causes silent numerical failure.

What Is EmbeddingGemma 2?

EmbeddingGemma 2 is an open-source, lightweight multimodal embedding model published under the Apache 2.0 license. EmbeddingGemma 2 projects text, source code, static images, video frames, and audio into a single 768-dimensional representation space.

Inputs map into a single shared embedding space. A text query vector compares directly against stored image, audio, or video vectors using standard distance metrics.

Input ModalityProcessing PipelineTarget Output
Text and Source CodeShared Encoder Pass768-Dim Vector Space
Images and Video FramesShared Encoder Pass768-Dim Vector Space
Audio ClipsShared Encoder Pass768-Dim Vector Space

The architecture is modular. Developers can load only the modality encoders required for their pipeline to reduce memory usage on constrained edge devices. The weights live on Hugging Face and Kaggle under the ID google/embeddinggemma-2 and run directly inside the SentenceTransformers library.

Benchmark Performance

EmbeddingGemma 2 improves on its predecessor across code retrieval and multilingual benchmarks.

BenchmarkEmbeddingGemma 2 ScoreEmbeddingGemma 1 ScoreTarget Metric
MTEB (code, v1)78.6868.76NDCG@10
MTEB (multilingual, v2)61.3661.15Mean(Task)
MSEB (Retrieval)69.54-Mean(Task)
Massive Image Embedding Benchmark (Lite)64.64-Mean-TaskType

According to Google DeepMind documentation, the model scores 78.68 NDCG@10 on the MTEB (code, v1) benchmark. It reaches 61.36 Mean(Task) on MTEB (multilingual, v2), covering over 100 languages. On multimodal evaluation tasks, it scores 69.54 Mean(Task) on MSEB Retrieval and 64.64 Mean-TaskType on the Massive Image Embedding Benchmark (Lite).

Supported Input Modalities

EmbeddingGemma 2 processes four main input categories:

  • Text and source code
  • Static images
  • Video
  • Audio clips

All modalities map directly into the same 768-dimensional space. Cross-modal retrieval works out of the box. A text string can match an image, or an audio clip can pull relevant text documents.

The model accepts mixed-modality payloads in single inference passes. You can supply an image combined with a text prompt. The encoder processes the combined payload into one vector for downstream retrieval.

Core Technical Limits

The model uses an 8,192-token context window. All modalities share this single budget during an inference call.

If you pass a single input type, the 8,192-token budget holds about 5.5 minutes of audio, 29 static images, or 58 video frames. Mixing modalities splits this allocation. If a prompt uses 4,000 text tokens, you have half the context window left for video frames or audio. Vision input size trades off latency against embedding quality.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

embeddings = model.encode([
    "def calculate_similarity(v1, v2): return np.dot(v1, v2)",
    "file:///path/to/diagram.png"
])

# Truncate vector from 768 down to 256 dimensions
truncated_embeddings = embeddings[:, :256]

Vector truncation introduces quality trade-offs. EmbeddingGemma 2 supports Matryoshka Representation Learning (MRL), letting you slice output vectors from 768 dimensions down to 512, 256, or 128 dimensions. Truncating to 128 dimensions causes strong quality degradation in multimodal retrieval, hitting image, video, and speech queries particularly hard.

Precision Limits and Float16 Failures

Do not run EmbeddingGemma 2 using float16 precision.

Precision ModeStatusOperational Outcome
float16UnsupportedSilent failures or NaNs
bfloat16SupportedSupported precision mode
float32SupportedSupported precision mode

Google DeepMind documentation explicitly states that float16 is unsupported. Running float16 causes silent failures or generates NaN values during embedding generation. Using float16 precision leads to silent failures or NaN values during execution.

Deployments must use bfloat16 or float32. If target hardware lacks native bfloat16 support, switch execution to standard float32 to maintain computational stability.

Hardware and Deployment Targets

EmbeddingGemma 2 runs locally on consumer CPUs, GPUs, and NPUs. Dedicated deployment options include Google AI Edge, ML Kit, MediaPipe, and LiteRT. Web applications can run the model directly inside browsers using transformers.js or WebGPU.

The model supports INT8 and INT4 quantization for low-latency, on-device setups. Pricing applies to Gemini Embedding 2 via Google Cloud services.

Methodology

Facts are compiled from the vendor's own announcement and documentation, linked below, at the time of writing; where the vendor hasn't published a number it is marked as such. GPUniq prices, when the model is already available here, come from our live catalog.

Sources

  1. 1.EmbeddingGemma 2 model card — Google AI for Developers (accessed )
  2. 2.EmbeddingGemma 2: an open, lightweight multimodal embedding model — Google DeepMind (accessed )

Want cheap GPUs for your next project?

Browse live GPU prices and rent the right card in seconds — H100, A100, RTX 4090, and 50+ more models.

More from the GPUniq blog